SwissIsoform v2

About SwissIsoform v2

Most genes are assumed to make one protein. Many actually start translation in more than one place, producing proteins with different beginnings. SwissIsoform catalogues those alternative proteins and weighs the evidence that each one is real — and that it does something different from the familiar form of the protein.

In the language of the field: alternative start codons (translation-initiation sites, TIS) are detected by ribosome profiling (Ribo-TISH). Each produces a protein that differs from the canonical one across a defined differential region, which is scored across six evidence categories — Conservation, Detection, Localization, Mutation landscape, Predicted structure and Structural characteristics (CDLMPS).

How to read an isoform card

Every gene in the catalogue gets a card, and they are all built the same way. Reading one takes about ten seconds:

  • Gene symbol and function chips — what the gene is, plus its curated UniProt keywords as clickable chips you can filter on.
  • The CDLMPS lineup — six letters, one per evidence category, each coloured by an LLM's verdict for that category on this gene's headline isoform: interesting, neutral, or not interesting. The lineup is the fastest read on the card — it tells you which kind of evidence is carrying this isoform.
  • The Best: isoform — which isoform the lineup describes, for genes that have more than one.
  • The interesting index — the lineup condensed into a single number (green +1, red −1) so the whole catalogue can be sorted by it.

Open a gene to see every isoform in full: the differential region and where it sits, all fifteen criteria with / / , the underlying numbers, and the model's written reasoning for each category.

Six evidence categories (CDLMPS)

Fifteen criteria are grouped into six categories, one per letter of CDLMPS. Each criterion returns ✓ (true), ✗ (false), or – (not evaluable) — a criterion whose underlying data could not be computed reports "not evaluable" rather than counting against the isoform. On every isoform, an LLM reads each category and calls it interesting, neutral, or not interesting; those verdicts colour the CDLMPS lineup you see on each gene card and drive the "interesting index" (see below).

C Conservation

C1 Primate conservation
How well is the isoform-unique region conserved at the amino-acid level across primates (mean percent identity)? High AA identity in close relatives argues the alternative protein is real and translated, not a sequencing or annotation artifact. (Reading-frame intactness is shown as context.)
C2 Mammalian conservation
The same amino-acid identity test across mammals. Conservation over deeper evolutionary distance is stronger evidence the ORF is under selection to be translated. (Reading-frame intactness is shown as context.)
C3 Coding Selection
phyloP / phastCons measure per-base evolutionary constraint. A high absolute phyloP over the differential region means the sequence itself is under strong purifying selection — evidence it is coding and does something. (Enrichment over the shared core is shown as context only.)

D Detection

D1 Expression Breadth
Is the alternative start used in more than one cell line? Reproducible initiation across independent samples argues against a one-off ribosome artifact.
D2 Start-Site Usage
How efficiently ribosomes initiate at this start (TIS) relative to background, per cell line. Strong, reproducible initiation supports a genuine translation event.
D3 Peptide Evidence
Were tryptic peptides unique to the differential region detected by mass spec? Direct peptide evidence is the strongest proof the isoform protein exists.

L Localization

L1 Compartment
Do the predicted localization features (DeepLoc prediction, sorting signals, or membrane association) differ between the canonical and the isoform? A protein with changed localization features acts in a different cellular context.
L2 Sorting Signals
Do N-terminal targeting signals (SignalP secretion, TargetP mitochondrial/chloroplast) differ between canonical and isoform? N-terminal changes most directly add or remove targeting peptides.

M Mutation Landscape

M1 Germline Variants
Does healthy human germline variation (gnomAD) avoid this region (depletion ratio < 1×), and is it intrinsically constrained (ESM-C constraint delta > 0)? Depletion of population variation plus high sequence constraint mean the region resists change — it is functionally important. Scored only where the unique region is canonical coding sequence, i.e. on truncations. gnomAD is a tolerance catalogue, not a disease one; disease/cancer variants (ClinVar / COSMIC) live in M2.
M2 Clinical Variants
Are disease (ClinVar / COSMIC) variants enriched per nucleotide in the differential region versus the shared core (disease enrichment ratio ≥ 1×)? Disease variants concentrating in the unique region tie it to phenotype.

P Predicted Structure

P1 Fold Confidence
Does the gained/lost region fold into ordered structure (pLDDT) and shift the protein's biophysical character (charge, hydropathy, disorder)? Structured, biophysically distinct regions are more likely to be functional.
P2 Core Fold Perturbation
The shared region is the stretch of protein identical in both the isoform and the canonical (the canonical body for an extension; the post-truncation body for a truncation). Folded in both contexts it normally comes out nearly identical, so its Cα backbone RMSD ≈ 0. A high shared-region RMSD (superposed on the shared residues only, from the ESMFold2 structures) means the extension or truncation reorganizes how the retained region folds — a rare, high-interest functional signal. TM-score is a length-normalized companion; the check only fires when both structures are confidently folded, and uORF/altORF isoforms (no shared region) are not evaluable.
P3 Secondary Structure
Does the differential region contain actual secondary structure — a helix or strand — rather than coil? ESMFold2 predicts coordinates but no secondary structure, so elements are assigned from those coordinates (P-SEA, Cα geometry). Because they are derived from a prediction, each element carries its own mean pLDDT: a geometrically clean helix running through a disordered stretch is geometry fitted to a guess, so BOTH length and confidence are required to score. The direction differs by ORF type — an extension GAINS the element (a candidate functional addition), a truncation LOSES one, and there the element is read off the canonical structure because the removed segment exists only in it. This says nothing about whether the element is integrated with the rest of the fold; the contact and PAE evidence in this same category answers that.

S Structural Characteristics

S1 Domain change
Does the isoform gain or lose a real InterPro functional domain in the differential region (disorder/structural-only signatures excluded)? Gaining or losing a domain changes function directly.
S2 Biophysics
Does the whole-protein biophysical character shift between the canonical and the isoform — hydropathy (GRAVY), charged fraction, or intrinsic disorder? A meaningful shift on any of these means the alternative region changes the protein's physicochemical behaviour (solubility, membrane affinity, condensate propensity), independent of whether it folds.
S3 SAE features
Do sparse-autoencoder (SAE) interpretability features of the ESM-C protein language model differ between the isoform and the canonical — features gained or lost in the model's learned feature space? A presence check: it fires when any interpretable feature is active in one protein but not the other. The score is a binary gained-or-lost check.

The differential region & reading frame

The differential region is the part of the isoform that is not shared with the canonical protein. Two cases, and the distinction (diff_space) governs which tools are valid:

  • Truncation (diff_space = canonical) — the isoform starts downstream, losing the canonical N-terminus. The lost region is ordinary canonical-coding sequence, so canonical-frame tools (e.g. AlphaMissense) apply.
  • Extension / uORF / alt-ORF (diff_space = isoform) — the isoform starts upstream or out-of-frame, so the gained region is not in the canonical reading frame. Canonical-frame predictors are meaningless here (AlphaMissense reports these positions as "intronic" / N/A); only frame-aware tools (ESM-C ΔLLR, alt-ORF consequence) apply.

Exploring the gene index

The landing page lists every gene as a card carrying its CDLMPS lineup. Three controls help you find the isoforms that matter:

  • Search — filter by gene symbol or UniProt accession.
  • Filter by function — narrow to genes carrying a given UniProt keyword (a curated functional term such as Kinase, Cell cycle, or DNA-binding). Type to pick from the terms actually present in the catalogue; each gene card also lists its own keywords as clickable chips. Select several to require all of them (or switch to any).
  • Category — keep only genes whose headline (Best:) isoform fires the chosen CDLMPS categories — a letter fires when its LLM verdict is interesting or not-interesting. Selecting C D (with all) keeps genes firing both Conservation and Detection.
  • Sort — interesting index — order genes by their best isoform's net CDLMPS score (green +1, red −1; see below), surfacing the most promising isoforms first.

The metrics behind the criteria

Every measurement that feeds a CDLMPS criterion is described below: what it measures, how it is calculated (formula, algorithm, or tool and version), and a citation where one exists. Entries are grouped by the category and module they support. Where a threshold is provisional and pending genome-wide calibration, this is stated explicitly.

Conservation — evolutionary constraint (C1 · C2 · C3)
phyloP per-base constraint

What it measures. Basewise evolutionary constraint at single-nucleotide resolution: how much slower (purifying selection) or faster (acceleration) a position is evolving than the neutral expectation, measured directly in genomic coordinates over the isoform's own ORF.

How it's calculated. phyloP scores are read from the pre-computed cactus241way BigWig track derived from the Zoonomia 241-mammal Cactus whole-genome alignment. The conservation method (§3.4, "Evolutionary conservation") queries the track at the initiation codon (3 nt) and across the 13-nt Kozak window, and — using the per-ORF genomic exon intervals from the assembly layer's transcript-skeleton walker — computes length-weighted mean phyloP over the isoform-unique region, the canonical-shared region, and the unique/shared enrichment ratio. The mean phyloP over the unique region is the value scored by criterion C3 (threshold ≥ 2.0, a principled anchor for strong purifying selection; the enrichment ratio is context only). A status field distinguishes ok, not_run (no track/config), and no_skeleton (ORF intervals unavailable) so a genuine zero is not confused with a missing measurement.

Refs: phyloP [phylop], Zoonomia [zoonomia]

phastCons per-base constraint

What it measures. The probability that each nucleotide belongs to a conserved element, a complementary (element-based, phylo-HMM) view of constraint alongside phyloP's per-site rate test.

How it's calculated. phastCons scores are read from the pre-computed UCSC phastCons100way BigWig track (100-vertebrate alignment). As with phyloP (§3.4), the method queries the track at the initiation codon (3 nt) and the 13-nt Kozak window, and computes length-weighted mean scores over the isoform-unique region, the canonical-shared region, and their enrichment ratio using the per-ORF genomic exon intervals. phastCons is carried as context and is not the basis of a scored criterion (C3 scores phyloP). The same ok / not_run / no_skeleton status applies.

Refs: phastCons [phastcons]

Reading-frame conservation — mean amino-acid percent identity (mean_pident)

What it measures. The degree to which the isoform-unique open reading frame is preserved as a translatable coding frame across the placental-mammal radiation, summarized as the mean amino-acid percent identity of the unique region across aligned species. This is the scored reading-frame metric and feeds criteria C1 (primates) and C2 (mammals).

How it's calculated. Per §3.4 ("Reading-frame integrity"), the unique-region exon intervals are extracted from the Zoonomia Cactus alignment HAL with hal2maf (see hal2maf). For each aligned species the method tests start-codon conservation, scans for frameshifting indels and premature stop codons, and records the species' amino-acid percent identity to the human ORF. Per-species identities are aggregated separately over the primate and mammalian species sets into a mean percent identity. C1 requires the primate mean ≥ 0.80 and C2 the mammalian mean ≥ 0.50 (both provisional, pending genome-wide calibration). Status values: not_run (no HAL), no_skeleton / no_unique_region (no region to query), no_alignment (region did not align), ok.

Refs: Zoonomia [zoonomia]

Reading-frame conservation — fraction of species with intact frame (frac_intact)

What it measures. The fraction of aligned primate species (and, separately, mammalian species) in which the isoform-unique ORF retains an intact, frameshift- and premature-stop-free reading frame. Context for C1/C2, not the score basis.

How it's calculated. Per §3.4, after each species is classified as frame-intact or disrupted (via the start-codon test plus the frameshift-indel and premature-stop scans on its hal2maf-extracted alignment), the intact calls are aggregated into the fraction of primate and of mammalian species with an intact frame. The methods note explicitly that this intact-frame fraction is "retained as context for the reason string but is no longer the score basis" — C1/C2 are scored on mean percent identity (m-frame-pident), and frac_intact accompanies them in the reason text.

Refs: Zoonomia [zoonomia]

Reading-frame conservation — number of species aligned (n_species_aligned)

What it measures. How many species the unique region could actually be assessed in — the support behind the per-species frame aggregates. A low count signals that the percent-identity and intact-frame summaries rest on few observations.

How it's calculated. Per §3.4, it is the count of species rows returned by the hal2maf extraction over the unique-region intervals that aligned and were analyzed (start-codon, frameshift, premature-stop, percent-identity tests run per species). Species sets are drawn from the curated primate and mammalian lists; species with no alignment over the region are not counted.

Refs: Zoonomia [zoonomia]

Reading-frame conservation — start-codon conservation

What it measures. Whether the isoform's initiation codon is itself conserved in each aligned species — evidence that the alternative TIS, not just the downstream frame, is evolutionarily maintained.

How it's calculated. Per §3.4, for each aligned species the method "tests start-codon conservation" on the hal2maf-extracted MAF of the unique region: it inspects the orthologous trinucleotide at the initiation position and records whether it is preserved. This per-species start-codon call contributes to the frame-intactness classification that drives frac_intact and feeds the percent-identity aggregation.

Refs: Zoonomia [zoonomia]

Reading-frame conservation — deepest-intact species / phylogenetic depth

What it measures. How far back the isoform-unique ORF's coding frame can be traced: the most deeply diverging species in which the frame remains intact, with its phylogenetic depth — a single readout of the evolutionary age of the alternative ORF.

How it's calculated. Per §3.4, among species whose frame the method classifies as intact, it selects the deepest-diverging one and reports its phylogenetic depth, "read from the HAL's own species tree" (parsed from the Cactus HAL rather than a hard-coded tree, so it tracks the alignment release). Reported per radiation (primate and mammalian) as a deepest-intact species plus depth.

Refs: Zoonomia [zoonomia]

hal2maf alignment extraction

What it measures. Not a score itself — the extraction step that supplies the per-species multiple alignment of the isoform-unique region on which all reading-frame metrics (m-frame-pident, m-frame-frac-intact, m-frame-startcodon, m-frame-deepest) are computed.

How it's calculated. Per §3.4, the unique-region exon intervals (from the assembly layer's transcript-skeleton walker) are passed to hal2maf, which slices the Zoonomia Cactus alignment HAL into a per-species multiple alignment format (MAF) block over exactly those genomic intervals. The resulting MAF is then analyzed species-by-species for start-codon conservation, frameshifts, premature stops, and amino-acid percent identity. When no HAL is available the method reports not_run; an empty alignment over the region yields no_alignment.

Refs: Zoonomia [zoonomia]

Detection — initiation & expression (D1 · D2)
Initiation efficiency (TIS counts / gene RNA-seq counts)

What it measures. How productively ribosomes initiate at an alternative start codon relative to the transcriptional output of its gene, per cell line. This is the signal behind criterion D2. It is a normalized ratio of read counts — not a bounded probability and not the Ribo-TISH enrichment p-value; values are not capped at 1 and should be read as a relative initiation rate, not a percentage.

How it's calculated. For each TIS in a sample, the Ribo-TISH raw initiation read count (TISCounts) is divided by that gene's matched RNA-seq read count, the latter taken from the gene-level HTSeq-count table for the sample (§Cross-cell-line aggregation defines initiation efficiency as "the ratio of the TIS read count to the gene's RNA-seq count"). The ratio is carried per cell line through the aggregation step; D2 tests the maximum per-cell-line initiation efficiency against a (provisional, calibration-pending) threshold (§Evidence scoring). Distinct from CPM/RPM below, which normalize against the sample's total mapped reads rather than the gene's own RNA-seq depth.

Refs: Ribo-TISH [ribotish], HTSeq-count [htseq]

CPM / RPM (depth-normalized TIS count)

What it measures. The Ribo-TISH initiation read count for a TIS normalized to sequencing depth, so that initiation events are comparable across the six cell lines despite differing library sizes. Used as the expression-level filter input and reported per cell line in the output table.

How it's calculated. Each TIS's raw read count (TISCounts) is divided by the total number of uniquely-mapped RNA-seq reads in the matched sample — a single grand-total constant summed over all genes in that sample's HTSeq-count table — and scaled to reads per million (§Ribosome-profiling). This RPM value drives the low-expression filter: events with RPM below 0.1 are marked LowReadcounts and dropped (§significance filtering). Unlike initiation efficiency, the denominator is the whole-sample mapped-read total, not the per-gene RNA-seq count.

Refs: Ribo-TISH [ribotish], HTSeq-count [htseq]

TIS-enrichment p-value, frame-preference p-value & Fisher q-value

What it measures. The statistical confidence that an initiation event is a real, frame-coherent ribosome pile-up rather than background. These are the three significance criteria that gate every alternative TIS into the analysis.

How it's calculated. Ribo-TISH emits, per candidate TIS: a two-sided TIS-enrichment p-value comparing the ribosome-footprint pile-up at the candidate codon against local background; a frame-preference p-value testing whether reads concentrate in the codon's reading frame; and a Fisher-combined q-value integrating the two (§Ribosome-profiling). The significance filter marks an event NotSignificant unless all three pass: TIS-enrichment p ≤ 0.01, frame-preference p ≤ 0.01, and Fisher q ≤ 0.05 (§significance filtering). All thresholds are exposed as configurable parameters. The output table additionally carries a one-sided initiation-significance p-value per cell line (§Differential annotation table).

Refs: Ribo-TISH [ribotish]

Cell-line reproducibility count (D1)

What it measures. In how many of the six independently-processed cell lines the same alternative TIS is detected — a reproducibility signal for criterion D1. Recurrence across cell lines argues against a single-sample artifact.

How it's calculated. Each cell line (HeLa, K562, U2OS, and RPE1 asynchronous / quiescent / senescent) is filtered independently; the cross-cell-line combine then takes the union of distinct (gene symbol, transcript, ORF genomic span, start codon) tuples and attaches a per-cell-line expression vector, with zero-filled entries for non-observing samples (§Cross-cell-line aggregation). D1 counts the cell lines in which the TIS was observed and fires when that count is at least 3 (§Evidence scoring). The ≥ 3 cutoff is a fixed, principled anchor — not a provisional/calibration-pending threshold — and is not relaxed by run-script overrides.

Refs: Ribo-TISH [ribotish], GENCODE v49 [gencode]

Kozak context window

What it measures. The 13-nucleotide sequence context around an alternative initiation codon — the determinant of how favourably a ribosome recognizes that start. Supplies the raw window for the Kozak strength scores below and the GC content reported per TIS.

How it's calculated. For each TIS a 13-nt mRNA window spanning positions −9 to +4 is extracted, with the initiation codon occupying positions 9–11 (0-indexed). Extraction is strand-aware: on the plus strand the window is read directly from the primary-assembly GRCh38 FASTA; on the minus strand the reverse-complement of the corresponding plus-strand interval is taken so the string reads 5′→3′ in mRNA orientation (§Kozak-context extraction). The window is populated whether the start is ATG or a near-cognate (CTG/GTG/TTG/ACG/AAG), and the initiation trinucleotide at positions 9–11 is preserved independently in the TIS record. The initiation-context method also records the GC content of this window and the start trinucleotide (§ORF/initiation-context).

Refs: Kozak [kozak]

Kozak strength (Hamming distances to consensus)

What it measures. How closely the 13-nt context matches the optimal Kozak consensus — a proxy for initiation strength. A smaller Hamming distance means a stronger, more consensus-like start.

How it's calculated. The extracted window is scored against the consensus gccgccRccATGG under three weighting schemes derived from Kozak (§ORF/initiation-context): (1) a full-length Hamming distance over all 13 positions; (2) a major-position Hamming distance restricted to the two strongly-conserved positions, −3 and +4; and (3) a partial-weighting score applying weight 1.0 to the major positions and 0.1 to the minor positions. The R at position −3 denotes the purine consensus (A or G). All three are reported per TIS as a site-level annotation.

Refs: Kozak [kozak]

Detection — proteomic evidence (D3)
In-silico tryptic peptide enumeration (isoform-unique)

What it measures. Which tryptic peptides the alternative isoform could yield in a mass-spectrometry experiment, and which of those peptides are unique to the isoform — i.e. do not occur anywhere in the canonical protein's own tryptic digest. Unique peptides are the candidate observations that could, in principle, distinguish the alternative isoform from the canonical in a proteomics experiment, so the uniqueness flag is the primary handle for downstream experimental validation.

How it's calculated. Per §"In-silico proteomic detectability": the protein sequence is digested in silico by cleaving after every lysine (K) or arginine (R) that is not immediately followed by proline (the KP / RP exceptions), allowing up to one missed cleavage. Peptides shorter than 7 or longer than 30 residues — outside the typical tandem-MS observation range — are discarded. Each retained peptide is recorded with its sequence, start position, end position, length, and a boolean uniqueness flag set by testing whether that peptide sequence occurs anywhere in the canonical protein's tryptic digest (unique = absent from canonical). For the canonical protein annotation, uniqueness is undefined (there is no alternative sequence to compare against) and is recorded as None rather than False, to distinguish the semantically-undefined case from an empirically-negative one.

Refs: — (in-silico digest; no external tool)

PepQuery2 spectral validation

What it measures. Whether an enumerated peptide has actually been observed in real mass spectra, elevating it from "in-silico detectable" to "experimentally validated." This is the empirical evidence that the alternative isoform's region is translated and present at the protein level.

How it's calculated. Per §"In-silico proteomic detectability": in-silico detectability alone is not treated as evidence. When pre-computed PepQuery2 validation results are provided, each enumerated peptide already observed in a reanalysis of public proteomics datasets is additionally flagged as experimentally validated. The D3 evidence criterion fires when at least one isoform-unique tryptic peptide is matched in public MS spectra by the PepQuery2 search; D3 is reported as None (not False) until the PepQuery2 precompute exists. PepQuery2 performs unrestricted (open) peptide-spectrum matching against public MS datasets and scores significance against a reference proteome, so an isoform-unique peptide passing PepQuery2 is direct proof of translation.

Refs: PepQuery2 [pepquery]

Truncations: no credit for canonical lost-region peptides

What it measures. A correctness guard ensuring that, for truncated isoforms, mass-spec detectability evidence reflects only sequence the isoform actually contains — never the canonical region that the truncation deletes.

How it's calculated. Per §"In-silico proteomic detectability": the digest and the isoform-uniqueness test are computed on the alternative isoform's own sequence. A truncation's differential region is the canonical lost region (the residues absent from the isoform), so peptides from that lost region are by construction part of the canonical digest and the isoform cannot generate them. Because uniqueness is defined as "absent from the canonical digest," lost-region peptides can never be flagged isoform-unique, and a truncation therefore receives no proteomic-detectability credit for canonical lost-region peptides — only peptides present in the (shorter) isoform sequence are enumerated and tested.

Refs: — (consequence of the digest/uniqueness definition)

Localization & targeting (L1 · L2)
DeepLoc 2.1 — subcellular compartment

What it measures. The predicted primary subcellular localization of a protein (one of ten compartments — nucleus, cytoplasm, mitochondrion, etc.), computed independently for the canonical protein and for each alternative-TIS isoform so that a re-routed N-terminus can be detected.

How it's calculated. Sequences are scored with DeepLoc 2.1 in its "Fast" mode (§"Subcellular localization"). Because DeepLoc's dependencies are pinned to a Python 3.8 stack that conflicts with the rest of the pipeline, it is run inside an isolated conda environment whose stdout the main pipeline consumes. To amortize the fixed cost of ESM embedding, all unique canonical and isoform sequences in a batch are written to a single FASTA, DeepLoc is invoked once, and the per-sequence predictions (primary localization plus the signal-peptide / NLS / mitochondrial / peroxisomal sub-predictions and a membrane-type call) are joined back onto records by sequence identity. A protein-hash-keyed lookup memoizes, so identical sequences across genes or cell lines are predicted exactly once per run. Per-residue ESM embeddings are discarded; only the final categorical and probabilistic outputs are stored.

Refs: DeepLoc 2.1 [deeploc]

DeepLoc localization-feature change (L1)

What it measures. Criterion L1 — whether the alternative isoform is predicted to localize or associate differently from the canonical protein. Labelled "localization features changed."

How it's calculated. The comparator diffs the DeepLoc outputs of the canonical protein against the isoform and fires L1 if any localization feature differs: the primary-compartment prediction, the detected sorting signals (signal-peptide / NLS / mitochondrial / peroxisomal sub-predictions), or the membrane-type call (§"Evidence criteria"). Keying on the full feature set rather than the single headline compartment label catches a re-routed protein even when the top compartment is unchanged.

Refs: DeepLoc 2.1 [deeploc]

SignalP 6.0 — signal peptide

What it measures. Presence and type of an N-terminal secretory signal peptide, predicted per protein (canonical and isoform) — the kind of N-terminal feature an alternative TIS most directly adds or removes.

How it's calculated. Sequences are scored with SignalP 6.0, which predicts the five signal-peptide types using a protein language model (§"Signal and targeting peptides"). It is run once over the batch of unique sequences in its own pinned environment and joined back by sequence identity; for each protein the predicted class, the class probabilities, and the predicted cleavage site are recorded.

Refs: SignalP 6.0 [signalp6]

TargetP 2.0 — transit/targeting peptide

What it measures. The broader N-terminal targeting class of a protein (canonical and isoform): signal peptide (SP), mitochondrial transit peptide (mTP), or chloroplast transit peptide (cTP), with the predicted cleavage site where present.

How it's calculated. Sequences are scored with TargetP 2.0 (§"Signal and targeting peptides"), run once over the batch of unique sequences in its own pinned environment and joined back by sequence identity. For each protein the predicted class, the class probabilities, and the cleavage site are recorded. Running TargetP alongside SignalP lets the targeting criterion catch a gained or lost signal that either tool alone would miss.

Refs: TargetP 2.0 [targetp2]

Targeting-signal change (L2)

What it measures. Criterion L2 — whether the alternative N-terminus gains or loses an N-terminal targeting peptide relative to the canonical protein.

How it's calculated. L2 fires when SignalP 6.0 or TargetP 2.0 disagrees on the canonical versus the isoform call (§"Evidence criteria"; §3.6). Combining both predictors gives complementary coverage — secretory signal peptides (SignalP) and the mitochondrial / chloroplast transit classes (TargetP) — so a targeting change either tool alone would miss is still detected.

Refs: SignalP 6.0 [signalp6], TargetP 2.0 [targetp2]

Mutation landscape — variant effect & burden (M1 · M2)
M1 — gnomAD germline depletion ratio

What it measures. Whether common germline variation avoids the isoform-unique region — a population-genetics signal of purifying constraint. It is one of the two independent inputs that can satisfy criterion M1 (germline tolerance / constraint).

How it's calculated. A genomic-intersection step tags every gnomAD v4.1 exome variant by membership in the isoform-unique versus canonical-shared exon intervals (from the assembly layer's transcript-skeleton walker) and exposes the nucleotide length of each region. The depletion ratio is the per-nucleotide variant density of the unique region divided by that of the shared region: (ngnomAD,unique / ntunique) / (ngnomAD,shared / ntshared). A value below 1 means common variation is depleted from (avoids) the unique region. The ratio is undefined (None) when a denominator is zero. M1 is satisfied when this ratio falls below 0.80 — a provisional threshold, a config-driven placeholder pending calibration on a genome-wide run, not a principled anchor (§"Per-variant effect scoring", §"Evidence scoring framework"). gnomAD v4.1 exomes are preprocessed from per-chromosome VCFs (PASS records, canonical-transcript VEP, non-null gene symbol) into a gene-indexed parquet read by PyArrow filter pushdown (§"Clinical variant burden").

Refs: gnomAD v4.1 [gnomad]

M1 — ESM-C sequence-constraint delta

What it measures. Whether the residues of the isoform-unique region are more constrained by a protein language model than the canonical-shared region — i.e. the language model finds the unique-region sequence less substitutable. This is the second independent signal that can satisfy M1.

How it's calculated. A dedicated language-model module (plm_vep) records the ESM-C (EvolutionaryScale) sequence-constraint profile of the protein itself: the mean wild-type-residue log-probability (the model's log-likelihood of the residue actually present, with no alternate allele) taken over the isoform-unique and canonical-shared regions, plus the count of strongly-constrained positions in each. A higher (nearer zero) log-probability marks a residue the model predicts confidently from its context, and hence a more conserved one — a direction checked against the Zoonomia 241-mammal phyloP track over the same codons (Pearson +0.40, Spearman +0.48 across 6,842 residues, positive in all 27 proteins examined). The constraint delta is the unique-region mean minus the shared-region mean, so it is positive when the isoform-unique region is better predicted than the retained core; it is a difference rather than a ratio because both terms are negative log-probabilities, which makes a ratio invert and diverge as the denominator approaches zero. M1 is satisfied when the delta is at or above 0.0 — a provisional threshold pending genome-wide calibration, not a principled anchor — and is not evaluable on extensions and separate ORFs, whose unique region was never canonical coding sequence and so has no baseline to be contrasted against. This constraint signal is distinct from the per-variant ΔLLR scores (which carry an allele change); it scores the wild-type sequence alone (§"Per-variant effect scoring", §"Evidence scoring framework").

Refs: ESM-C (EvolutionaryScale); masked-marginal LLR method [esmc]

Per-variant ESM-C ΔLLR (masked-marginal effect)

What it measures. The predicted tolerability of a specific amino-acid substitution under a protein language model — how much less likely the alternate residue is than the wild-type residue at that position. More negative ΔLLR indicates a substitution the model tolerates less well. It contributes to the per-variant "damaging" flag (reported as germline/disease context), separate from the M1 constraint profile above.

How it's calculated. ESM-C (EvolutionaryScale) supplies a masked-marginal change in log-likelihood, ΔLLR = log P(alt) − log P(wt), evaluated at the variant's residue from the per-position amino-acid distribution of a single full-protein forward pass (cached on disk by protein hash). A missense variant is flagged damaging when its ΔLLR is at or below −7.5 — the cutoff Brandes et al. established for the analogous ESM-2-class LLR, carried over as a provisional threshold under ESM-C and pending genome-wide recalibration — or when AlphaMissense calls it likely_pathogenic. Loss-of-function consequences (frameshift, stop-gained, splice, start-lost) are damaging on their own, since neither missense predictor sees them. A predicted-damaging gnomAD variant at allele frequency ≥ 10−3 is gated out of the damaging flag (ACMG allele-frequency benign evidence); ClinVar and COSMIC variants are never gated (§"Per-variant effect scoring").

Refs: ESM-C (EvolutionaryScale); masked-marginal LLR method [esmc]; Brandes et al. [brandes]

AlphaMissense pathogenicity canonical frame only

What it measures. DeepMind's calibrated missense pathogenicity for a substitution — a 0–1 score with a likely_pathogenic / ambiguous / likely_benign class. Used as one branch of the per-variant damaging flag.

How it's calculated. The score and class are looked up by genomic coordinate from the AlphaMissense hg38 table. Critical frame caveat: AlphaMissense is defined in the frame of the canonical transcript, so it applies to shared-region variants and to truncation-unique variants (which still lie within the canonical CDS) but is N/A for extension-unique variants, which fall outside the canonical CDS entirely and are re-framed by the alternative TIS. A missense variant is flagged damaging when AlphaMissense calls it likely_pathogenic or its ESM-C ΔLLR ≤ −7.5 (§"Per-variant effect scoring").

Refs: AlphaMissense [alphamissense]

M2 — disease density enrichment (ClinVar + COSMIC)

What it measures. Whether known disease variants concentrate in the isoform-unique region rather than merely being present — i.e. whether the unique N-terminus is a disease hotspot relative to the shared core. This is criterion M2.

How it's calculated. The same genomic-intersection step tags each disease variant (ClinVar + COSMIC) by membership in the isoform-unique versus canonical-shared exon intervals and exposes each region's nucleotide length. The enrichment ratio is the per-nucleotide disease-variant density of the unique region over that of the shared region: (ndisease,unique / ntunique) / (ndisease,shared / ntshared). M2 is satisfied when this ratio is at or above 1.0 — a principled anchor (a fixed zero-point: density-equal is the natural neutral line), not a tuned threshold. The ratio is None when undefined (zero/missing denominator). ClinVar comes from the GRCh38-filtered variant_summary.txt.gz (VCF-format allele columns, genomic plus-strand); COSMIC v102 (GRCh38) from the Sanger Genome Screens Mutant / NonCoding / Targeted Screens distributions, both as filter-pushdown parquet (§"Clinical variant burden", §"Per-variant effect scoring", §"Evidence scoring framework").

Refs: ClinVar [clinvar], COSMIC v102 [cosmic]

Codon-level consequence validation (isoform frame)

What it measures. The true molecular consequence of each queried variant in the alternative isoform's reading frame — because source-database HGVSp strings are written in the canonical frame and are invalid for alternative-TIS events, which can re-frame the protein (extensions) or delete the N-terminus (truncations, uORFs). This is what makes every downstream variant count (M1/M2, LoF gating) frame-correct.

How it's calculated. A dedicated validator re-maps each variant from its genomic position to an isoform-specific coding position using the GTF-derived CDS exon-boundary table for the TIS's transcript. The mapping is strand-aware: on the plus strand, genomic positions within CDS exons get sequential 0-indexed coding positions 5′→3′; on the minus strand, exons are walked in descending genomic order, positions traversed high-to-low, and ref/alt bases complemented before codon assembly. For SNVs, the codon containing the variant is extracted from the isoform's reference coding sequence, translated, and compared with the mutant-codon translation to assign one of synonymous_variant, missense_variant, stop_gained, stop_lost, reference_mismatch. In-frame indels become inframe_insertion / inframe_deletion by length differential; frameshifts are labelled frameshift_variant; equal-length multi-base substitutions are labelled mnv without single-codon translation (deferred). The downstream LoF gate keys on the exact stop_gained, stop_lost, and frameshift_variant terms. Variants outside the isoform's coding interval (upstream of truncated starts, downstream of extended stops, intronic) get protein position None but are retained for re-validation against other transcripts and still counted in the per-protein summary (§"Clinical variant burden and codon-level consequence validation").

Refs: gnomAD [gnomad], ClinVar [clinvar], COSMIC [cosmic]

Predicted structure & biophysics (P1 · P2 · S2)
Differential-region pLDDT (predicted fold confidence)

What it measures. Whether the isoform's differential region — the stretch of sequence that distinguishes the alternative isoform from its canonical counterpart — is predicted to adopt a folded, ordered structure rather than being disordered. This is criterion P1 (structured extension); the biophysical shift is scored separately as S2 (below).

How it's calculated. Each protein is folded with ESMFold2 (EvolutionaryScale; the pipeline's default structure backend, with Boltz-2 and Chai-1 selectable), and the mean per-residue pLDDT over the residues of the differential region is taken. ESMFold2 emits pLDDT natively on a 0–1 scale; P1 is satisfied when this mean is at or above 0.70. The 0.70 cutoff is a principled anchor fixed on biological grounds, not tuned to any dataset (§"Evidence scoring", threshold classes). When the structure prediction is unavailable the criterion returns None.

Refs: ESMFold2 (EvolutionaryScale); Boltz-2, Chai-1 (selectable backends)

Biophysical distinctness (S2)

What it measures. Whether the isoform is physicochemically distinct from the canonical protein, i.e. the canonical→isoform comparison crosses a meaningful threshold on at least one global descriptor. This is criterion S2 — a whole-protein biophysical shift, scored independently of whether the region folds (P1).

How it's calculated. From the canonical-versus-isoform comparison of the eighteen biophysical descriptors, S2 fires when any of three deltas crosses its cutoff (§"Biophysical descriptors"): a GRAVY delta ≥ 0.3, a charged-fraction delta ≥ 0.05, or a disorder delta ≥ 0.05. These three deltas are provisional distinctness cutoffs — config-driven placeholders pending calibration on a genome-wide run, not principled anchors (§"Evidence scoring", threshold classes). When the biophysical comparison is missing the criterion returns None.

Refs: BioPython [biopython], Kyte–Doolittle hydropathy [kd], TOP-IDP disorder scale [topidp]

Isoelectric point (pI)

What it measures. The pH at which the polypeptide carries no net charge — a global summary of its acidic/basic character. Compared canonical vs isoform as a scalar delta.

How it's calculated. pI is computed from an EMBOSS-style pK table by binary search over pH until the net charge is zero (§"Biophysical descriptors"), one of the eighteen per-sequence descriptors.

Refs: —

GRAVY (Kyte–Doolittle hydropathy)

What it measures. Overall hydrophobicity/hydrophilicity of the protein — positive values indicate hydrophobic, negative hydrophilic. The GRAVY delta (≥ 0.3) is one of the three triggers of the S2 biophysical-distinctness flag.

How it's calculated. The mean of the Kyte–Doolittle hydropathy index over all residues (Grand Average of Hydropathy) (§"Biophysical descriptors"). A per-residue hydropathy vector is also retained alongside the scalar for downstream positional comparisons.

Refs: Kyte–Doolittle hydropathy index [kd]

Disorder propensity (TOP-IDP) & fraction of disorder-promoting residues

What it measures. The predicted intrinsic disorder of the polypeptide — how likely it is to lack a fixed fold. The disorder delta (≥ 0.05) is one of the three triggers of the S2 biophysical-distinctness flag, complementing the structural pLDDT signal.

How it's calculated. A TOP-IDP disorder-propensity score is computed from the intrinsic-disorder amino-acid propensity scale of Campen et al., applied residue-by-residue (§"Biophysical descriptors"). A companion descriptor reports the fraction of disorder-promoting residues. The propensity table was implemented directly against the published residue-level values; a per-residue disorder-propensity vector is retained for positional comparisons.

Refs: TOP-IDP scale, Campen et al. [topidp]

Fraction of charged residues

What it measures. The proportion of residues bearing a formal charge (acidic + basic), a determinant of solubility, interaction surfaces, and disorder. The charged-fraction delta (≥ 0.05) is one of the three triggers of the S2 biophysical-distinctness flag.

How it's calculated. Count of charged residues divided by sequence length, one of the eighteen per-sequence descriptors (§"Biophysical descriptors").

Refs: —

Instability index

What it measures. A sequence-derived prediction of in-vitro protein stability based on dipeptide composition.

How it's calculated. Computed with BioPython's ProtParam module (§"Biophysical descriptors").

Refs: BioPython ProtParam [biopython]

Aromatic fraction

What it measures. The proportion of aromatic residues (F/Y/W), relevant to π–π stacking, binding pockets, and phase-separation propensity.

How it's calculated. Aromatic-residue count divided by sequence length, one of the eighteen per-sequence descriptors (§"Biophysical descriptors").

Refs: —

LLPS propensity (composite)

What it measures. Liquid–liquid phase-separation propensity — the tendency of the protein to drive or partition into biomolecular condensates.

How it's calculated. Reported as a composite score combining three sequence-derived components — prion-like-domain content (Q/N/G/S/Y fraction), π–π interaction propensity (F/Y/W/R/H/Q/N), and RG/FG-motif density — together with disorder and low-complexity fraction, under fixed weights (§"Biophysical descriptors"). The three components are emitted alongside the composite. The component propensity tables were implemented directly against published residue-level values.

Refs: —

Sequence-complexity descriptors

What it measures. The compositional complexity of the sequence — distinguishing complex globular sequence from low-complexity / repetitive regions that often co-occur with disorder and condensate biology. These complexity descriptors sit alongside the other seventeen biophysical descriptors (pI, GRAVY, charge, disorder, aromaticity, instability, LLPS) that together define the S2 biophysical-distinctness comparison.

How it's calculated. A family of descriptors (§"Biophysical descriptors"): whole-sequence Shannon entropy and its mean over a sliding 20-residue window; a normalized-complexity score; the fraction of residues in low-complexity regions (Shannon entropy below 2.2 bits over a sliding 12-residue window); the amino-acid diversity of the N-terminal window; and the longest homopolymer run. The Shannon-entropy complexity family was implemented directly from residue-level definitions.

Refs: —

Shared-region backbone RMSD & TM-score (P2)

What it measures. Whether the region the isoform shares with the canonical protein — identical in sequence in both forms — nonetheless folds differently once the alternative N-terminus is added or removed. Criterion P2: a rare, high-interest signal that the differential region reorganizes the retained body rather than simply appending to it.

How it's calculated. Both proteins are folded with ESMFold2, superposed on the shared residues only, and the Cα backbone RMSD over that shared region is taken (with TM-score as a length-normalized companion). Because the shared sequence is identical, a well-behaved isoform folds it almost identically (RMSD ≈ 0); a high RMSD means the extension/truncation perturbs the core fold. P2 fires above a provisional RMSD anchor, only when both structures are confidently folded; uORF/altORF isoforms (no shared region) are not evaluable.

Refs: ESMFold2 (EvolutionaryScale)

SAE interpretability features (S3)

What it measures. Whether the isoform and the canonical protein differ in the learned, human-interpretable features of a protein language model — a mechanistic-interpretability view of what changed, beyond structure and biophysics. Criterion S3 is a presence check: it fires when any interpretable feature is active in one protein but not the other (a feature "gained" or "lost").

How it's calculated. A sparse autoencoder (SAE) decomposes the residual-stream activations of the ESM-C protein language model (default: the 6-billion-parameter model, layer 60) into a dictionary of 16,384 features, of which only the top k = 64 are active at each residue. For the isoform and the canonical protein the pipeline records which features are present (fire on ≥ 1 residue), then compares the two — reporting features present only in the isoform ("gained"), only in the canonical ("lost"), and shared features whose peak activation shifts most between the two. It separately restricts the comparison to the isoform-unique differential region. The S3 score is the binary gained-or-lost check; the feature list itself is descriptive context the LLM reads.

Refs: ESM-C (EvolutionaryScale) [esmc]

Structural characteristics — domains & motifs (S1)
InterProScan member-DB signatures → InterPro entries

What it measures. The protein-domain architecture of each canonical and isoform protein: which functional, sequence, and structural signatures occur, and at what positional span within the protein.

How it's calculated. Domains are annotated with InterProScan 6, run as a Nextflow pipeline that aggregates member-database signatures (Pfam and others) into integrated InterPro entries. To amortize the large per-invocation startup cost, all unique canonical and isoform sequences in a batch are written to a single FASTA, InterProScan is invoked once, and each resulting hit — carrying its member database, signature accession, optional InterPro cross-reference, and start/end coordinates within the protein — is joined back onto the input records by sequence identity. The per-protein record stores the positional hit list plus a summary with a status field, so a protein with zero domains (ok, empty list) is distinguished from one whose scan did not complete.

Refs: InterProScan [interproscan]

S1 — real functional domain gained or lost in the differential region

What it measures. Whether the alternative TIS event adds or removes a genuine functional protein domain in the region that differs between the isoform and the other form — the S1 criterion.

How it's calculated. From the InterProScan hit lists of both forms, S1 counts only real functional domains: hits carrying a genuine InterPro accession, explicitly excluding disorder- and structure-only signatures (MobiDB-lite, coils/coiled-coil, low-complexity, SignalP, Phobius, TMHMM). A domain qualifies only when it starts inside the differential region and is absent from the other form's domain set — i.e. truly gained or lost, not merely shifted in position because the N-terminus moved. Domains present in both forms but repositioned do not count.

Refs: InterProScan [interproscan]

ELM-style linear-motif (SLiM) scanning

What it measures. Short linear motifs (SLiMs) — compact regulatory/binding sequence patterns — present in each protein and their positions, used to detect motifs gained or lost in the differential region.

How it's calculated. SLiMs are identified by regular-expression search against fourteen canonical patterns curated from the ELM database and earlier SLiM literature: CDK-substrate consensus ([ST]P), ATM/ATR-substrate consensus ([ST]Q), 14-3-3 Mode I and Mode II binding motifs, EB1 SxIP microtubule-plus-end-tracking, RGG methylation/RNA-binding, SH3-domain class-I and class-II binding, heme-regulatory motif (C[^C].{2}C[^C]H), cytochrome-c heme-binding (CXXCH), PCNA-interacting PIP-box and APIM, C2H2 zinc-finger, and RING-finger E3-ligase consensus. Each hit is recorded with its start/end coordinates, the literal matching substring, and the motif name; the summary reports per-motif-type density (hits per residue) and the number of residues involved in overlapping hits.

Refs: ELM database [elm]

Evidence-scoring framework
Evidence scoring (CDLMPS categories)

What it measures. A per-TIS summary that condenses every annotation above into fifteen criteria grouped under six categories — Conservation, Detection, Localization, Mutation landscape, Predicted structure, and Structural characteristics. Broadly, C and D speak to whether the isoform is real; L, M, P and S to whether it changes function — but they are read as one panel, not two decoupled scores.

How it's calculated. The framework runs after the comparator, because most criteria read a canonical-versus-isoform (or, for genomic criteria, an isoform-unique-versus-canonical-shared-region) comparison rather than a raw annotation value. Each criterion is evaluated independently against a fixed threshold and returns True / False / not-evaluable; there is no weighting (§"Evidence scoring").

Refs: SwissIsoform methods §Evidence scoring

Tri-state criteria (True / False / None)

What it measures. Whether each individual criterion is satisfied, genuinely unsatisfied, or simply not assessable — so that a low score caused by missing upstream data is distinguishable from one caused by real absence of evidence.

How it's calculated. Every criterion returns True (evidence present), False (evidence genuinely absent), or None (cannot be evaluated — the upstream annotation module did not run or did not emit the required field). None results are excluded from both the count of satisfied criteria and the evaluable count. The evaluable count is therefore the number of criteria that returned True or False, and the satisfied count is the number that returned True (§"Evidence scoring").

Refs: SwissIsoform methods §Evidence scoring

Conservation & Detection (C1–C3, D1–D3)

What it measures. Six independent lines of evidence that the alternative TIS produces a real isoform: cross-species coding conservation (C), and detection — reproducibility, initiation strength, and proteomic evidence (D).

How it's calculated. Each criterion thresholds one upstream measurement over the isoform-unique region. C1 — mean amino-acid percent identity across the primate radiation (reading-frame integrity method, §3.4) ≥ 0.80 (provisional). C2 — the same over the mammalian radiation ≥ 0.50 (provisional). C3 — mean phyloP over the isoform-unique region ≥ 2.0 (absolute strong-purifying-selection statement; the unique-vs-shared enrichment ratio is reported only as context). D1 — TIS detected in ≥ 3 cell lines. D2 — maximum per-cell-line initiation efficiency (TIS read count / gene RNA-seq count) exceeds a threshold. D3 — at least one isoform-unique tryptic peptide matched in public MS spectra by a pre-computed PepQuery2 search; in-silico detectability alone does not count, and D3 is None until the PepQuery2 precompute exists (§"Evidence scoring").

Refs: Zoonomia [zoonomia], phyloP [phylop], PepQuery2 [pepquery]; SwissIsoform methods §Evidence scoring

Localization, Mutation, Structure (L1–L2, M1–M2, P1–P2, S1–S3)

What it measures. Nine independent lines of evidence that the isoform-unique region alters protein function: localization and targeting change (L), mutational constraint and disease burden (M), predicted-structure signals (P), and structural characteristics — domains, biophysics, interpretability features (S).

How it's calculated. L1 — comparator flags a change in any localization feature (DeepLoc primary compartment, sorting signals, or membrane-type call). L2 — SignalP or TargetP disagrees between canonical and isoform (§3.6). M1 — isoform-unique region shows germline constraint by either signal: ESM-C constraint delta (unique − shared mean log P(wt)) ≥ 0.0 (provisional) or gnomAD germline depletion ratio < 0.80 (provisional, §4.5); None when both inputs are missing, and on extensions and separate ORFs, where neither contrast has a canonical-coding baseline. M2 — disease (ClinVar/COSMIC) variant density of the unique region relative to the canonical-shared region ≥ 1.0 (zero-point); None when the ratio is undefined. P1 — mean ESMFold2 pLDDT over the differential region ≥ 0.70 (structured extension). P2 — shared-region Cα RMSD above a provisional anchor, i.e. the retained body re-folds; not evaluable without a shared region. S1 — at least one real InterPro domain (genuine InterPro accession, excluding disorder/structure-only signatures such as MobiDB-lite, coils, low-complexity, SignalP, Phobius, TMHMM) starts inside the differential region and is absent from the other form (§3.5), rather than merely repositioned. S2 — a whole-protein biophysical shift (GRAVY ≥ 0.3, charged-fraction ≥ 0.05, or disorder ≥ 0.05 delta; provisional). S3 — any SAE (ESM-C) interpretability feature gained or lost between isoform and canonical (§"SAE interpretability features"); None when the SAE stage did not run.

Refs: DeepLoc [deeploc], SignalP-6 [signalp6], TargetP-2 [targetp2], gnomAD [gnomad], ClinVar [clinvar], COSMIC [cosmic], AlphaMissense [alphamissense], ESM-C [esmc], InterProScan [interproscan]; SwissIsoform methods §Evidence scoring

The interesting index & the CDLMPS lineup

What it measures. A single at-a-glance readout of how promising an isoform is, shown as the coloured CDLMPS lineup on each gene card and used to rank genes.

How it's calculated. For each isoform, an LLM reads each of the six categories and returns a verdict — interesting, neutral, or not interesting. On the lineup a letter lights green when its category is interesting, red when not interesting, and stays grey when neutral or unscored. The interesting index is the net score: +1 per green letter, −1 per red, neutral = 0. Each gene card headlines its highest-scoring isoform (the "Best:" lineup), and the index page can sort genes by this interesting index. (This editorial index is separate from the underlying True/False criterion scores above — it summarizes the LLM's reading of them.)

Refs: SwissIsoform methods §Evidence scoring

Threshold provenance (principled anchors vs provisional)

What it measures. Which scoring thresholds are biologically fixed versus config-driven placeholders awaiting calibration — so a reader knows which numbers are load-bearing and which may move.

How it's calculated. Principled anchors are fixed on biological grounds and not tuned to any dataset: the C3 phyloP cutoff of 2.0 (strong purifying selection), the P1 pLDDT cutoff of 0.70, the M2 disease-enrichment zero-point of 1.0, and the ≥ 3 cell-line requirement for D1. Provisional thresholds are config-driven placeholders pending calibration on a genome-wide run: the C1/C2 percent-identity cutoffs (0.80 / 0.50), the M1 constraint-enrichment (2.0) and depletion-ratio (0.80) cutoffs, the S2 biophysical-distinctness deltas, and the P2 shared-region RMSD anchor. The curated 13-gene cohort validates the scoring logic, not these calibration points, since a reviewer-selected gene set is not representative of the genome-wide distribution (§"Evidence scoring").

Refs: SwissIsoform methods §Evidence scoring

How a criterion is decided

Every criterion follows the same four-step shape, so the table reads consistently across all fifteen:

  1. Claim — the criterion states one falsifiable thing about the isoform (e.g. "the differential region is under coding-level purifying selection").
  2. Canonical-vs-isoform comparison — the same predictor is run on both the canonical protein and the alternative isoform, and the evidence shown is the difference between them: the differential region against the shared core, or the canonical start against the alternative start. A metric in isolation is never the basis for a call.
  3. Score — the comparison is reduced to ✓ / ✗ / – against a fixed threshold. A criterion whose inputs could not be computed returns – (not evaluable) and is excluded from the denominator, so missing data never counts against an isoform.
  4. Directionality — for the mutation-landscape criteria the sign of the signal is what matters. Germline tolerance (M1) reads in the depletion direction: healthy-population (gnomAD) variation avoiding the unique region, plus high ESM-C sequence constraint, argues the region is functionally load-bearing. Disease overlap (M2) reads in the enrichment direction: ClinVar / COSMIC variants concentrating in the unique region (per nucleotide) versus the shared core ties it to phenotype. These two databases are kept separate — gnomAD is a tolerance catalogue, never disease burden.

Thresholds & validation. A few cutoffs are principled anchors fixed on first principles — strong purifying selection (absolute phyloP ≥ 2.0) for C3, the disease-density zero-point (enrichment ≥ 1.0) for M2, reproducibility across ≥ 3 cell lines for D1, and the structure-confidence floor (pLDDT ≥ 0.70) for P1. The remaining cutoffs are provisional and pending calibration on a genome-wide run; the curated gene set shown here validates the logic of each criterion, not the exact numeric thresholds, which are deliberately not tuned to make a small reviewer-picked set look good.

References

  1. Cheng J, Novati G, Pan J, et al. Accurate proteome-wide missense variant effect prediction with AlphaMissense. Science 381:eadg7492 (2023). doi:10.1126/science.adg7492.
  2. Cock PJA, Antao T, Chang JT, et al. Biopython: freely available Python tools for computational molecular biology and bioinformatics. Bioinformatics 25(11):1422–1423 (2009). doi:10.1093/bioinformatics/btp163.
  3. Brandes N, Goldman G, Wang CH, Ye CJ, Ntranos V. Genome-wide prediction of disease variant effects with a deep protein language model. Nature Genetics 55(9):1512–1522 (2023). doi:10.1038/s41588-023-01465-0.
  4. Landrum MJ, Chitipiralla S, Brown GR, et al. ClinVar: improvements to accessing data. Nucleic Acids Research 48(D1):D835–D844 (2020). doi:10.1093/nar/gkz972.
  5. Sondka Z, Dhir NB, Carvalho-Silva D, et al. COSMIC: a curated database of somatic variants and clinical data for cancer. Nucleic Acids Research 52(D1):D1210–D1217 (2024). doi:10.1093/nar/gkad986.
  6. Thumuluri V, Almagro Armenteros JJ, Johansen AR, Nielsen H, Winther O. DeepLoc 2.0: multi-label subcellular localization prediction using protein language models. Nucleic Acids Research 50(W1):W228–W234 (2022). doi:10.1093/nar/gkac278.
  7. Kumar M, Gouw M, Michael S, et al. ELM — the Eukaryotic Linear Motif resource in 2020. Nucleic Acids Research 48(D1):D296–D306 (2020). doi:10.1093/nar/gkz1030.
  8. Candido S, Hayes T, Derry A, et al. Language Modeling Materializes a World Model of Protein Biology. bioRxiv (2026). doi:10.64898/2026.06.03.729735.
  9. Frankish A, Diekhans M, Jungreis I, et al. GENCODE 2021. Nucleic Acids Research 49(D1):D916–D923 (2021). doi:10.1093/nar/gkaa1087.
  10. Chen S, Francioli LC, Goodrich JK, et al. A genomic mutational constraint map using variation in 76,156 human genomes. Nature 625(7993):92–100 (2024). doi:10.1038/s41586-023-06045-0.
  11. Anders S, Pyl PT, Huber W. HTSeq — a Python framework to work with high-throughput sequencing data. Bioinformatics 31(2):166–169 (2015). doi:10.1093/bioinformatics/btu638.
  12. Jones P, Binns D, Chang H-Y, et al. InterProScan 5: genome-scale protein function classification. Bioinformatics 30(9):1236–1240 (2014). doi:10.1093/bioinformatics/btu031.
  13. Kyte J, Doolittle RF. A simple method for displaying the hydropathic character of a protein. Journal of Molecular Biology 157(1):105–132 (1982). doi:10.1016/0022-2836(82)90515-0.
  14. Kozak M. An analysis of 5′-noncoding sequences from 699 vertebrate messenger RNAs. Nucleic Acids Research 15(20):8125–8148 (1987). doi:10.1093/nar/15.20.8125.
  15. Wen B, Zhang B. PepQuery2 democratizes public MS proteomics data for rapid peptide searching. Nature Communications 14(1):2213 (2023). doi:10.1038/s41467-023-37462-4.
  16. Siepel A, Bejerano G, Pedersen JS, et al. Evolutionarily conserved elements in vertebrate, insect, worm, and yeast genomes. Genome Research 15(8):1034–1050 (2005). doi:10.1101/gr.3715005.
  17. Pollard KS, Hubisz MJ, Rosenbloom KR, Siepel A. Detection of nonneutral substitution rates on mammalian phylogenies. Genome Research 20(1):110–121 (2010). doi:10.1101/gr.097857.109.
  18. Zhang P, He D, Xu Y, et al. Genome-wide identification of translation initiation sites by ribosome profiling. Nature Communications 8(1):1749 (2017). doi:10.1038/s41467-017-01981-8.
  19. Teufel F, Almagro Armenteros JJ, Johansen AR, et al. SignalP 6.0 predicts all five types of signal peptides using protein language models. Nature Biotechnology 40(7):1023–1025 (2022). doi:10.1038/s41587-021-01156-3.
  20. Almagro Armenteros JJ, Salvatore M, Emanuelsson O, et al. Detecting sequence signals in targeting peptides using deep learning. Life Science Alliance 2(5):e201900429 (2019). doi:10.26508/lsa.201900429.
  21. Campen A, Williams RM, Brown CJ, Meng J, Uversky VN, Dunker AK. TOP-IDP-scale: a new amino acid scale measuring propensity for intrinsic disorder. Protein & Peptide Letters 15(9):956–963 (2008). doi:10.2174/092986608785849164.
  22. Christmas MJ, Kaplow IM, Genereux DP, et al. Evolutionary constraint and innovation across hundreds of placental mammals. Science 380(6643):eabn3943 (2023). doi:10.1126/science.abn3943.