Skip to content

Model status and validation scope

Original models: Mol-JEPA, Rottach et al. · Boltz-2, Passaro et al. · STATE, Arc Institute.

Mol-JEPA, the selected STATE ST checkpoint, both STATE SE representations and Boltz-2 have passed numerical checks against their trained checkpoints within the scopes below. These checks assess model compatibility and output preservation; they do not rank biological prediction quality. Rewriting a Python model requires its input/output contract as well as the weight files listed on Hugging Face.

The active scope covers Boltz-2 and Nesso-1, the existing Mol-JEPA/STATE work, and the tested NovoMolGen 32M SMILES AtomWise checkpoint. X-Cell and OmniCell have been removed from the plan; stFormer is set aside because of the archive and backend work required. Bioptimus H-optimus-0 and H-optimus-1 both returned HTTP 403 GatedRepo in the supplied-token weight check, so they are deferred. The check saved no credential. The inactive rows retain their historical findings. list_targets() omits these inactive models; include_inactive=True retrieves the older survey explicitly.

Target Method and evidence Validation boundary
Boltz-2 Generic sharing stores identical immutable tensors once across the complete confidence/affinity pair; both original APIs remain callable. Joint registered storage 4.09 → 2.06 GB (49.55%). All 48 native tensor outputs byte-equal before/after sharing and fresh reload on CPU/MPS for one small protein–ligand fixture with standard sampling schedules. No speedup claim. See Boltz-2.
Nesso-1 Pack channelwise contractions into contiguous matrices; retain the original weights and complete native outputs. Lossless checkpoint packing is a separate storage option. Modest CPU gains on two prepared inputs; no reliable MPS gain. The 165.4 MB checkpoint packs to 140.8 MB, with unchanged resident weights. Full output checks and runnable commands are in Nesso-1.
Mol-JEPA Exact graph rewrites and explicit SMILES specialization retain predictions, CLS, latent embeddings and requested attentions. Pretrained CPU/MPS numerical tests passed; separate from a biological benchmark.
STATE ST, HVG Replogle K562 The real checkpoint has a frozen, entirely zero vocabulary table with 32,000 rows and width 328. A general exact rewrite stores one row and retains its full-shaped weight view and token-index API. 49,396,728 to 38,901,056 stored parameters, a 21.25% reduction. All predict_step outputs, including decoded counts and metadata, were bitwise identical on CPU and MPS float32 at 1, 7, 64 and 128 cells; direct token-ID probes also matched.
STATE SE-100M Two generic finite-domain tables replace the large protein lookup while the original encoder remains available for raw-vector inputs. 212,038,824 to 151,243,944 parameters (28.67%). All 105 CPU output comparisons passed, including 2048-gene inputs; maxabs 6.68e-6. Portable reload is bitwise. Seven synthetic AnnData variants also reproduce preprocessing and all 1,034 exported features bitwise. These precomputed tables remain CPU-only. A separate lossless original-table mode saves 6.31% of registered state and passes 377 byte comparisons per backend after CPU/MPS reload, but the tested MPS calls are about five times slower. No biological benchmark is claimed.
X-Cell Mini The proposed 55M model exposes an AnnData/perturbation API. A future compiler must preserve diffusion sampling, masks and all output genes. Both loading and prediction are stubs. The HF repository contains only an image, README and .gitattributes; there is no model to compress yet.
Bioptimus H-optimus Public timm-compatible model IDs are verified. H-optimus-0 uses a ViT-g/14 with 40 blocks, 24 heads, width 1536 and four registers. Its ordinary affine/norm blocks can be inventoried once accessible. Gated model access, no weights or runtime tests. Do not drop image tokens or registers under a same-output promise. H0-mini is an author-distilled model, not evidence for this package's non-distillation method.
OmniCell, BGIResearch Actual official weights were loaded safely and hashed. Shared scalar experts admit an affine sum; finite gene routing must retain the live gene table and load-balance outputs. The routing table is larger than its router. A near-tied gene switches experts despite tiny local errors, producing 0.2503 embedding error: rejected. Existing non-flash attention changes layout and scaling; whole-model macOS support remains unvalidated.
stFormer, csh3 The active model copies scFoundation position tables and continuous-value encoders. The defined Embedding→LayerNorm GeneEncoder is unused; equal trained tables have not been established. Source audit only. CUDA FlashAttention and the old dependency stack need a verified port. Released weights are in a 9.34 GB compressed Zenodo archive; no separately downloadable checkpoint was verified.
NovoMolGen 32M AtomWise Only this existing variant remains in scope, with pinned native Hugging Face float32 safetensors. First-layer token normalization and Q/K/V projections are a conditional finite-domain candidate. The input embedding must remain for residuals and hidden states. Removing projection weights requires a token-only contract. A 32M research prototype passes complete CPU/MPS checks bitwise, including generated and cached outputs, using shape-specific tables. Whole-model reductions are 1.676% CPU and 1.267% MPS; the broad native embedding-input API and production export remain unsupported. See NovoMolGen audit.

STATE ST saves 41,982,688 bytes of resident and raw tensor storage by deduplicating a table unused by its usual inputs_embeds forward pass. It removes no active matrix multiplications. The 21.25% parameter reduction therefore saves resident memory without implying faster inference or a 21.25% smaller packed file: a zero table already packs well. The original 471.7MB training checkpoint also contains optimizer state, so compression comparisons use the original inference weights.

The CPU STATE SE finite-domain artifact stores encoder(normalize(protein_embedding[gene])) for each allowed gene. The unnormalized gene_embedding_layer path needs a separate table, and CLS/dataset tokens need their own compiled constants. Counts are added after this encoder, so continuous count inputs remain variable. The AnnData adapter reuses the original dataset and collator with ordered gene names only; it needs no raw protein dictionary. Known gene IDs use the compiled tables; arbitrary raw vectors still use the original encoder. The original gene-name helper requires its external raw protein dictionary, and direct access to the removed raw table is outside this adapter. The smaller shared projection-plus-norm proposal failed complete-output checks and remains rejected. See STATE usage and finite-domain compilation.

A CPU check of the unmodified SE FlashTransformerEncoderLayer, instantiated with small random weights, retained all 8,544 parameters and produced bitwise-identical outputs after the general affine compiler. Its nonlinear and residual barriers prevented a rewrite. Despite the class name, it calls PyTorch SDPA; enabling the CUDA backend flag alone does not establish a CUDA dependency. Its SDPA dropout remains nonzero in eval unless overridden at construction, as the upstream inference loader does by requesting zero dropout.

The selected STATE ST release needs Transformers 4.52.3. Transformers 5 rejects hidden width 328 with 12 attention heads even though the checkpoint supplies head dimension 64, so STATE uses a separate .venv-state environment. Its source also eagerly imports a disabled legacy VCI decoder through an obsolete package name. The local loader defers that unavailable branch and executes the original ST/base/utility source unchanged. The portable architecture directory contains the source hashes, JSON hyperparameters and attribution.

The package registry is compressme.targets. Detailed reachable-path corrections and the actual OmniCell counterexample are in the OmniCell/stFormer audit. Raw source inventory is in benchmarks/cell_model_source_audit.json; the portable STATE artifact contains its source/config hashes and original licenses.

Primary sources: STATE source, ST checkpoint, SE checkpoint, X-Cell source, X-Cell files, Bioptimus release, Bioptimus HF, OmniCell source, OmniCell checkpoint location, stFormer source, stFormer release archive.