MAD2L1
EXTENDED 226 aa (canonical 205 aa) · UniProt Q13257 · CDLMPS
chr4:120066797:-:TTG:ENST00000296509.11
AI summary SignalP predicts a cryptic secretory signal peptide in the extension, but DeepLoc's miscalibrated cytoplasm baseline undercuts a true relocalization claim.
The dominant differential finding is a SignalP-called signal peptide (canonical OTHER → isoform SP, probability delta 0.83) with a cleavage site right at the extension boundary, plus a confidently-folded 19-residue helix (pLDDT ~0.79) in the added segment — together suggestive of a novel N-terminal targeting element. However, this helix sits with very high PAE (~18-22 Å) relative to the retained HORMA core and makes almost no contacts with it, so it reads as a locally folded but structurally unintegrated appendage rather than a docked domain; the shared HORMA core itself is essentially unperturbed (RMSD 1.2 Å but high pTM/pLDDT, TM-score 0.98).
MAD2L1 is established as a nuclear/nuclear-envelope spindle-checkpoint protein that must reach kinetochores and the NPC to function; a genuine secretory signal would be a striking functional divergence, redirecting the protein away from its known checkpoint role entirely. But DeepLoc calls the CANONICAL protein 'Cytoplasm' (not nucleus), which already disagrees with the known nuclear/nuclear-envelope localization — the tool is miscalibrated for this protein, so its isoform read (also Cytoplasm, with reduced confidence and only a peripheral/soluble membrane-association shift) cannot be trusted as evidence of altered trafficking. The SignalP signal-peptide call is intriguing but stands uncorroborated by TargetP (noTP unchanged) and by a calibrated localization tool, and no domain gain/loss or whole-protein biophysical shift accompanies it, so the practical impact on MAD2 checkpoint function (MAD1 binding, kinetochore/NPC localization) remains unresolved rather than demonstrated.
DeepLoc's canonical-protein call (Cytoplasm) contradicts the known nuclear/nuclear-envelope localization, so its localization outputs for this gene are not trustworthy; the SignalP signal-peptide gain is real but lacks TargetP/DeepLoc corroboration and the extension is not evolutionarily conserved, weakening confidence that this is a biologically consequential targeting change.
Folding
Evidence — click any tile for the differential-region detail
LLM reasoning
How well is the isoform-unique region conserved at the amino-acid level across primates (mean percent identity)? High AA identity in close relatives argues the alternative protein is real and translated, not a sequencing or annotation artifact. (Reading-frame intactness is shown as context.)
Comparison
Sequence conservation · primate
mean AA identity over the differential ORF is the score basis; frame intactness across the clade is context
| Conservation metric | Canonical ORF | Differential ORF |
|---|---|---|
| Mean AA identity | 100% | 74% |
| Frame intact (fraction of species) | 96% | 72% |
| Species aligned | 25 | 25 |
| Species frame-intact | 24 | 18 |
| Start codon conserved | 100% | 24% |
| Deepest intact species | Microcebus_murinus | Microcebus_murinus |
| Phylo depth (MRCA) | 7 | 7 |
The same amino-acid identity test across mammals. Conservation over deeper evolutionary distance is stronger evidence the ORF is under selection to be translated. (Reading-frame intactness is shown as context.)
Comparison
Sequence conservation · mammalian
mean AA identity over the differential ORF is the score basis; frame intactness across the clade is context
| Conservation metric | Canonical ORF | Differential ORF |
|---|---|---|
| Mean AA identity | 95% | 43% |
| Frame intact (fraction of species) | 35% | 15% |
| Species aligned | 20 | 20 |
| Species frame-intact | 7 | 3 |
| Start codon conserved | 95% | 12% |
| Deepest intact species | Loxodonta_africana | Microcebus_murinus |
| Phylo depth (MRCA) | 12 | 7 |
phyloP / phastCons measure per-base evolutionary constraint. A high absolute phyloP over the differential region means the sequence itself is under strong purifying selection — evidence it is coding and does something. (Enrichment over the shared core is shown as context only.)
Comparison
Per-base conservation · differential region
absolute mean phyloP over the differential region is the score basis (high = strong purifying selection); the shared column and enrichment are context, not the claim
| Track | Differential | Shared | vs shared |
|---|---|---|---|
| phyloP mean | -0.375 | 4.26 | -0.0881 |
| phastCons mean | 0.0355 | 0.887 | — |
Start site · canonical vs isoform
initiation context + per-base conservation at each start codon
| Property | Canonical | Isoform |
|---|---|---|
| Start codon | ATG | TTG |
| Kozak context (−9..+4) | TCCCTGGCCATGG | GAAGTGCTGTTGG |
| phyloP at start codon | 6.88 | -0.862 |
| phastCons at start codon | 1 | 0 |
| phyloP over Kozak window | 2.69 | 0.0819 |
| phastCons over Kozak window | 0.479 | 0.000154 |
| Kozak mismatch — full consensus | 4 | 8 |
| Kozak window GC content | 0.692 | 0.538 |
LLM reasoning
Is the alternative start used in more than one cell line? Reproducible initiation across independent samples argues against a one-off ribosome artifact.
Comparison
Per-cell-line usage · canonical vs isoform
Fisher q = 6.7e-07
| Cell line | Canonical | Isoform | p-value |
|---|---|---|---|
| HeLa | 62.7 | 2.19 | 0.0107 |
| K562 | 20.5 | 0.835 | 6.7e-07 |
| U2OS | 11 | — | — |
How efficiently ribosomes initiate at this start (TIS) relative to background, per cell line. Strong, reproducible initiation supports a genuine translation event.
Comparison
Start-Site Usage · canonical vs isoform
ribosome initiation efficiency at the canonical start vs this alternative start, per cell line
| Cell line | Canonical | Isoform |
|---|---|---|
| HeLa | 1.5 | 0.0522 |
| K562 | 0.545 | 0.0222 |
| U2OS | 0.174 | — |
Were tryptic peptides unique to the differential region detected by mass spec? Direct peptide evidence is the strongest proof the isoform protein exists.
Comparison
Peptide Evidence (canonical vs isoform)
isoform-unique peptides are direct evidence the alternative protein exists
| Feature | Canonical | Isoform |
|---|---|---|
| Tryptic peptides (in-silico) | 34 | 5 |
| Validated by mass-spec | 0 | 2 |
| Isoform-unique peptides | — | 5 |
Details
Peptide Evidence (canonical vs isoform)
- peptide MEPLWLLSAEWK 0–12
- peptide VLLFVSLAMALQLSR 13–28
- peptide MEPLWLLSAEWKR 0–13
- validated RVLLFVSLAMALQLSR 12–28
- validated VLLFVSLAMALQLSREQGITLR 13–35
LLM reasoning
Do the predicted localization features (DeepLoc prediction, sorting signals, or membrane association) differ between the canonical and the isoform? A protein with changed localization features acts in a different cellular context.
Comparison
Subcellular localization (DeepLoc)
localization features changed (prediction/signals/membrane)
| Property | Canonical | Isoform |
|---|---|---|
| Predicted location | Cytoplasm | Cytoplasm |
| Sorting signals | Nuclear export signal | Nuclear export signal |
| Membrane | Soluble | Peripheral|Soluble |
Do N-terminal targeting signals (SignalP secretion, TargetP mitochondrial/chloroplast) differ between canonical and isoform? N-terminal changes most directly add or remove targeting peptides.
LLM reasoning
Does healthy human germline variation (gnomAD) avoid this region (depletion ratio < 1×), and is it intrinsically constrained (ESM-C constraint delta > 0)? Depletion of population variation plus high sequence constraint mean the region resists change — it is functionally important. Scored only where the unique region is canonical coding sequence, i.e. on truncations. gnomAD is a tolerance catalogue, not a disease one; disease/cancer variants (ClinVar / COSMIC) live in M2.
Comparison
Germline Variants · differential vs shared region
gnomAD (population) variant density per nucleotide — depletion ratio < 1× means healthy human variation avoids the region (constrained), the M1 basis alongside ESM-C constraint
| Variant set | Differential | Shared | Depletion ratio |
|---|---|---|---|
| gnomAD variants | 51 | 283 | 1.8× |
Sequence constraint (ESM-C) · differential vs shared region
lower LLR = more constrained; enrichment = differential / conserved
| Property | Differential | Shared | Enrichment |
|---|---|---|---|
| ESM-C mean LLR | -2.51 | -0.0466 | — |
| Constrained positions | 0 | 1 | — |
Predicted-damaging variants · germline (gnomAD)
AlphaMissense / ESM-C / LoF calls per region (length-normalized)
| Variant set | Differential | Shared | Enrichment |
|---|---|---|---|
| Scorable variants | 48 | 283 | 1.7× |
| Damaging variants | 4 | 144 | 0.27× |
| — of which loss-of-function | 4 | 14 | 2.8× |
| AlphaMissense-pathogenic | 0 | 83 | 0× |
Predictor scores · germline (gnomAD)
scored: 407 ESM-C · 254 AlphaMissense
| Property | Differential | Shared |
|---|---|---|
| Mean ΔLLR (ESM-C) | -0.221 | -5.78 |
| Min ΔLLR (ESM-C) | -2.08 | -14.1 |
| Mean AlphaMissense | — | 0.51 |
Are disease (ClinVar / COSMIC) variants enriched per nucleotide in the differential region versus the shared core (disease enrichment ratio ≥ 1×)? Disease variants concentrating in the unique region tie it to phenotype.
Comparison
Clinical-variant burden · differential vs shared region
counts per region; ratio is density-normalized (variants per nucleotide, differential ÷ shared — ≥1× = disease variants concentrate in the differential region, the M2 basis)
| Variant set | Differential | Shared | Enrichment |
|---|---|---|---|
| Disease variants | 3 | 107 | 0.27× |
| Pathogenic | 0 | 0 | — |
Predicted-damaging variants · disease (ClinVar/COSMIC)
AlphaMissense / ESM-C / LoF calls per region (length-normalized)
| Variant set | Differential | Shared | Enrichment |
|---|---|---|---|
| Scorable variants | 3 | 107 | 0.27× |
| Damaging variants | 0 | 68 | 0× |
| — of which loss-of-function | 0 | 16 | 0× |
| AlphaMissense-pathogenic | 0 | 37 | 0× |
Predictor scores · disease (ClinVar/COSMIC)
scored: 407 ESM-C · 254 AlphaMissense
| Property | Differential | Shared |
|---|---|---|
| Mean ΔLLR (ESM-C) | -0.495 | -7.05 |
| Min ΔLLR (ESM-C) | -0.766 | -13.4 |
| Mean AlphaMissense | — | 0.552 |
LLM reasoning
Does the gained/lost region fold into ordered structure (pLDDT) and shift the protein's biophysical character (charge, hydropathy, disorder)? Structured, biophysically distinct regions are more likely to be functional.
Comparison
Structure (ESMFold2) · canonical vs isoform
TM-score 0.947 · RMSD 0.5 Å · 10 interface contacts
| Metric | Canonical | Isoform |
|---|---|---|
| pLDDT (whole protein) | 0.938 | 0.924 |
Fold confidence · differential vs shared region
mean ESMFold2 pLDDT in each region
| Metric | Differential | Shared | Enrichment |
|---|---|---|---|
| pLDDT | 0.775 | 0.774 | 1 |
The shared region is the stretch of protein identical in both the isoform and the canonical (the canonical body for an extension; the post-truncation body for a truncation). Folded in both contexts it normally comes out nearly identical, so its Cα backbone RMSD ≈ 0. A high shared-region RMSD (superposed on the shared residues only, from the ESMFold2 structures) means the extension or truncation reorganizes how the retained region folds — a rare, high-interest functional signal. TM-score is a length-normalized companion; the check only fires when both structures are confidently folded, and uORF/altORF isoforms (no shared region) are not evaluable.
Comparison
Shared-region structural change
Cα RMSD superposed on the shared residues only; TM-score is length-normalized · shared-region Cα RMSD 1.2 Å · shared TM-score 0.983 · shared region 205 aa · min shared pLDDT 0.938 · global TM-score 0.947 · global RMSD 0.5 Å
| Property | Canonical | Isoform |
|---|---|---|
| Shared-region pLDDT | 0.938 | 0.939 |
Does the differential region contain actual secondary structure — a helix or strand — rather than coil? ESMFold2 predicts coordinates but no secondary structure, so elements are assigned from those coordinates (P-SEA, Cα geometry). Because they are derived from a prediction, each element carries its own mean pLDDT: a geometrically clean helix running through a disordered stretch is geometry fitted to a guess, so BOTH length and confidence are required to score. The direction differs by ORF type — an extension GAINS the element (a candidate functional addition), a truncation LOSES one, and there the element is read off the canonical structure because the removed segment exists only in it. This says nothing about whether the element is integrated with the rest of the fold; the contact and PAE evidence in this same category answers that.
Comparison
Secondary structure · differential vs shared region
helices and strands assigned from the predicted coordinates (P-SEA)
| Metric | Differential | Shared |
|---|---|---|
| Alpha helices | 1 | 3 |
| Beta strands | 0 | 4 |
| Longest element (aa) | 19 | 23 |
| Mean pLDDT | 0.79 | 0.96 |
Elements and coordinates
1 in the differential region, 7 in the shared core — residue numbering is 1-based on the protein holding the region
| Isoform-unique | Shared core |
|---|---|
| alpha helix 9–27 19 aa · pLDDT 0.79 | alpha helix 34–56 23 aa · pLDDT 0.96 |
| — | alpha helix 80–99 20 aa · pLDDT 0.95 |
| — | beta strand 121–127 7 aa · pLDDT 0.98 |
| — | alpha helix 142–162 21 aa · pLDDT 0.95 |
| — | beta strand 169–177 9 aa · pLDDT 0.97 |
| — | beta strand 199–208 10 aa · pLDDT 0.97 |
| — | beta strand 212–221 10 aa · pLDDT 0.98 |
Below threshold
0 in the differential region, 4 in the shared core — shorter than 6 aa or below pLDDT 0.70, so not counted above
| Isoform-unique | Shared core |
|---|---|
| — | beta strand 64–68 5 aa · pLDDT 0.95 |
| — | beta strand 73–77 5 aa · pLDDT 0.96 |
| — | beta strand 107–110 4 aa · pLDDT 0.98 |
| — | beta strand 188–191 4 aa · pLDDT 0.90 |
LLM reasoning
Does the isoform gain or lose a real InterPro functional domain in the differential region (disorder/structural-only signatures excluded)? Gaining or losing a domain changes function directly.
Comparison
Domains & motifs (canonical vs isoform)
gained = only in the isoform; lost = only in the canonical
| Feature | Canonical | Isoform |
|---|---|---|
| InterPro domains | 6 | 6 |
| Short linear motifs | 2 | 2 |
Biophysical character of the isoform-differential region versus the shared canonical core — pI, hydropathy, charge, disorder and related properties. Descriptive; the folding (P1) score keys off the GRAVY / charge / disorder deltas.
Comparison
Biophysics · differential vs shared region
highlighted rows are enriched in the differential region
| Property | Differential | Shared | Enrichment |
|---|---|---|---|
| Isoelectric point (pI) | 7 | 4.69 | 1.49 |
| Hydropathy (GRAVY) | 0.909 | -0.12 | -7.55 |
| Fraction charged | 0.191 | 0.259 | 0.737 |
| Disorder fraction | -0.0492 | 0.0591 | -0.832 |
| Disorder-promoting | 0.429 | 0.493 | 0.87 |
| Low-complexity fraction | 0 | 0 | — |
| Prion-like fraction | 0.0952 | 0.239 | 0.398 |
| LLPS score | 0.0496 | 0.143 | 0.346 |
| π–π propensity | 0.191 | 0.234 | 0.814 |
| Aromaticity | 0.143 | 0.0927 | 1.54 |
| Instability index | 44.2 | 51 | 0.867 |
| Shannon entropy | 3.18 | 4.06 | 0.782 |
| Normalized complexity | 0.735 | 0.94 | 0.782 |
Sparse-autoencoder (SAE) interpretability features on the ESM-C residual stream, comparing the isoform against the canonical protein. This is the S3 criterion — a presence check that fires when any interpretable feature is gained or lost. Only the top 30 features per category (ranked by activation, prevalence ≥ 2) are saved; each table shows the top 10 interpretable features, and generic / "unknown" features are hidden unless you expand it. What are SAE features? →
Part 1 · Global feature comparison
Which SAE features fire in the isoform vs the canonical protein, across the whole sequence.
Isoform-only features — 105 total (prevalence ≥ 2)
| Feature | Label | Description | Activation | Prevalence |
|---|---|---|---|---|
| #3609 | Position-2 N-terminus residue detector | Detector of the identity of the second residue at the extreme N-terminus of proteins, with strongest preference for small residues (especially Ser, Ala, and Gly; also Thr/Pro), i.e., an N-terminal position-2 signal commonly associated with generic, disordered starts and co-translational N-terminus processing | 10.72 | 2 |
| #6484 | Prion-like low-complexity IDRs | Polar low-complexity intrinsically disordered regions (IDRs)—especially Q/N-rich (prion-like) segments and S/T- and glycine-rich tracts with simple SG/TG/PG repeats—typically located in inter-domain linkers and terminal tails of eukaryotic proteins and often harboring Ser/Thr phosphorylation sites. | 10.38 | 22 |
| #448 | N-terminal leader/targeting segments | N-terminal leader/targeting segments: the feature activates on the extreme N-terminus (first 2–5 residues) of proteins, especially the N-region of signal peptides and organelle transit peptides, and more generally on short, disordered/low-complexity N-terminal tails; strongest at residue ~3 and largely independent of amino-acid identity, occurring across all taxa and functions. | 10.02 | 3 |
| #4992 | Small-residue disordered N-terminus | Context-dependent free N-terminus signature: short, disordered amino-terminal stretches beginning with the initiator Met followed by residues compatible with MAP cleavage or N-acetylation (e.g., M–G/A/S/T/V, and other contexts including M–D/E/L); strongest emphasis on small residues at positions 2–3; activation is absent when the native N-terminus is structured or starts with a signal peptide/transmembrane segment. | 8.30 | 2 |
| #11504 | Initiator methionine at M1 | Initiator methionine at the very start of the polypeptide chain (M1), i.e., the translation start residue, independent of protein family, taxonomy, membrane association, or presence of signal/leader/propeptide regions; often annotated as post-translationally removed and typically situated in a flexible, coil-like N-terminus. | 8.24 | 2 |
| #9998 | Weak N+3 positional cue | Absolute N-terminal positional cue centered near the fourth residue (N+3), independent of amino-acid identity; most often within signal/transit peptides of precursors but also present in non-secreted proteins; the effect is weak and often does not produce a discrete detected site. | 8.04 | 3 |
| #7996 | Generic extreme N-terminus detector | Generic extreme N-terminus detector: marks residues around position 5 (occasionally 6–7) at the very start of the polypeptide, most often within the N-region of signal/transit peptides of secreted or organelle-targeted precursors, but also in generally disordered N-termini of cytosolic/nuclear proteins; not motif- or residue-specific | 5.63 | 3 |
| #14822 | Folded-domain interface residue detector | A broadly tuned, weak detector of single residues within folded domains, with a mild bias toward sites involved in binding or assembly (protein–protein interfaces, receptor- or calmodulin-binding patches, and small‑molecule/lipid contacts), often falling in alpha‑helices or at loop/strand edges. | 5.55 | 2 |
| #12227 | Tryptophan residue detector | Residue-level detector for tryptophan (W) side chains, often firing on multiple W positions within a sequence and with a mild bias toward N‑terminal Ws; broadly applicable across taxa and protein classes, with frequent occurrences in short membrane/secreted proteins but also in diverse enzymes (e.g., ATP synthase ε, proteases) and mitochondrial proteins. | 5.48 | 2 |
| #6377 | Interface-adjacent conserved motifs | A "functional-site adjacency" signal that fires at residues within well-ordered secondary structure participating in, or immediately flanking, ligand/cofactor-binding, catalytic, or protein-protein interaction interfaces, with a particularly strong pattern at conserved motif residues of small interaction modules (CCT, SOCS-box, BC-box). | 5.41 | 3 |
| #14646 | Short N-terminal leader segment | Short N-terminal segments within the first ~10–40 amino acids, often Lys/Arg-enriched but sometimes purely hydrophobic. These commonly correspond to Sec-type signal peptides, signal anchors, or other short N-terminal segments preceding/overlapping the first hydrophobic helix, and also include disordered N-terminal tails in soluble proteins. | 5.40 | 14 |
| #7985 | Signal peptides and disordered tails | Compositionally biased, hydrophobic/aromatic-rich segments—often in low-structure regions including N-terminal pre-sequences, flexible linkers/tails, and short exposed stretches within mature chains. These regions are enriched for Leu/Val/Ile/Phe/Tyr with Pro/Gly and often include low-complexity repeats (e.g., Asn-rich tracts in Dictyostelium). The concept spans signal/transit peptides of secreted/membrane/organellar proteins and intrinsically disordered tails in soluble enzymes and viral accessory proteins. | 5.20 | 4 |
| #14493 | Short hydrophobic helical runs | Short hydrophobic helical segments rich in I/V/F (sometimes M), recognized in both membrane-targeting/insertion contexts (signal peptides, signal-anchors, transmembrane helices) and in hydrophobic helical patches within otherwise soluble small proteins. The feature flags contiguous aliphatic/hydrophobic runs across diverse proteins, including viral membrane/accessory proteins, secreted peptide/effector precursors, small ORFs/microproteins, and short basic/uncharacterized human proteins. | 5.05 | 3 |
| #5241 | Early N-terminal positional signal | Generic early N-terminus positional signal peaking at residue ~5–7, found broadly within the disordered N-tail of proteins, including both cleavable leader/targeting segments (signal peptides, propeptides) of precursor proteins and disordered N-terminal regions of mature proteins; residue identity is not specific | 5.00 | 3 |
| #678 | Sparse terminal residue peaks | Sparse residue-level activations distributed across short proteins and protein segments, with frequent firing in disordered N-terminal/C-terminal tails as well as in some short transmembrane or compact domains | 4.93 | 3 |
| #12470 | Aminergic GPCR TM5-TM6 motif | A feature that fires at conserved sequence motifs in class A aminergic GPCRs, with the strongest hits at a specific cytoplasmic loop motif between TM5 and TM6 (around "SSLER(A/V)AEHAQ") and additional weaker peaks scattered through transmembrane helices 4 and 6 and adjacent cytoplasmic regions. | 4.88 | 2 |
| #7384 | Leucine-biased disordered hydrophobic segments | Leucine-biased recognition of intrinsically disordered, low-complexity hydrophobic segments, including disordered insertions within otherwise structured proteins. | 4.82 | 6 |
| #6398 | Polar low-complexity disordered stretches | Polar/small-residue-enriched (often Ser/Thr- and Pro-rich) stretches, frequently within intrinsically disordered or low-complexity regions and N-terminal tails/propeptides; can extend into short, flexible helices in small proteins, and is also seen in localized patches within folded domains | 4.60 | 2 |
| #9395 | Short terminal targeting segment | Terminal or short internal targeting/interaction segment: a short, compositionally biased or hydrophobic patch near a protein terminus, or occasionally an internal low‑complexity loop, that functions as a signal peptide or signal‑anchor (Sec/ER/thylakoid/Tat entry) when hydrophobic, or as a basic/Ser/Lys/Arg/Pro‑rich low‑complexity segment that mediates localization or partner binding in soluble proteins; typically the single prominent terminal segment of short secreted, membrane, or regulatory proteins | 4.58 | 3 |
| #677 | Basic low-complexity disordered regions | Compositionally biased, low-complexity/intrinsically disordered protein segments, typically enriched in basic (Lys/Arg), proline, glycine and serine residues. These often correspond to cytosolic-facing or solvent-exposed disordered tails/linkers, occurring at either terminus or in internal disordered stretches. | 4.49 | 2 |
| #5337 | Pro/RGG-rich low-complexity IDRs | Intrinsically disordered, low‑complexity basic segments at termini and long loops, enriched in Pro/Gly and/or Arg/Ser (including RG/RGG or S/T clusters), common across eukaryotic and viral proteins. | 4.24 | 11 |
| #12751 | Ser/Pro-biased low complexity regions | Compositionally biased regions (IDRs or short low-complexity segments) that are explicitly Ser/Pro-biased or enriched for simple dipeptide repeats (SS/PP, RS/SR, RG/RGG) and basic residues; these tracts occur in viral accessory proteins and micropeptides and as Ser/Pro-rich linkers/termini in diverse proteins. Generic hydrophobic signal peptides without such polarity/repeat bias are not targets. | 4.13 | 2 |
| #3604 | Acidic Pro/Ser-rich disordered regions | Proline/serine-rich low-complexity and disordered regions, frequently with acidic/PEST-like character; the feature also fires on small proteins and on internal stretches outside obvious low-complexity tracts when local Pro/Arg/Ser content is high. | 4.00 | 8 |
| #14891 | Disordered termini and cleavage detector | Residue-level detector of intrinsically disordered, flexible termini and proteolytic processing junctions—especially N-termini (including the first residue of the mature chain after propeptide/leader cleavage)—with no strict amino‑acid specificity and a bias toward small/charged residues; common in small, secreted/precursor and viral proteins but also present in disordered regions of diverse proteins. | 3.99 | 3 |
| #3054 | Basic GP/PTS low-complexity tracts | Intrinsically disordered, low‑complexity segments enriched in glycine/proline and serine/threonine, often containing clusters of basic residues (Lys/Arg) and simple repeats; includes PTS- and GP-rich repeats in secreted/processed peptides and extracellular proteins, as well as basic, Ser/Arg‑rich viral and micropeptide regions. Activation requires extended low‑complexity tracts with these compositions; compact or acidic sequences lacking such tracts are typically not activated. | 3.90 | 4 |
| #8910 | Aliphatic-rich disordered LCRs | Ala/Val/Gly–enriched low‑complexity segments, often within intrinsically disordered or low‑confidence regions of small/unstructured proteins (signal peptides, pro‑peptide extensions, regulatory tails, microproteins), characterized by small aliphatic residues interspersed with sparse acidic residues; composition- rather than function-specific. | 3.86 | 3 |
| #13063 | Serine-rich region detector | Serine residues across a wide range of structural contexts, with a bias toward serine-rich segments in intrinsically disordered or low-complexity regions (often N-terminal tails), but also firing on serines within transmembrane helices and other structured contexts; the feature behaves primarily as a serine detector with some preference for S-rich patches. | 3.84 | 2 |
| #246 | Regulatory low-complexity IDRs | Intrinsically disordered, low-complexity regulatory segments enriched in acidic/serine/proline/glutamine/glycine residues in eukaryotic proteins—most prominently in transcription factors and in regulatory tails/linkers of signaling/scaffold proteins—corresponding to transactivation/repression or partner-interaction regions adjacent to but outside structured DNA-binding or catalytic domains. | 3.83 | 3 |
| #9946 | Disordered N-termini and coils | Detector for intrinsically disordered, low-structure N‑terminal pre-sequences (signal peptides’ N/C regions, organellar transit peptides, and propeptides) and, more generally, flexible coil/low‑pLDDT segments; strongest bias for the first 10–70 residues but can also mark internal loops in large enzymes and short low‑complexity micropeptides. | 3.55 | 2 |
| #14000 | IDR and signal peptide detector | Generic detector of low-complexity/intrinsically disordered segments and short hydrophobic N‑terminal stretches (signal peptides/first TM anchors), with preference for S/T/P/G/A/N- and N/Q‑rich tracts; largely avoids well‑ordered helical cores | 3.53 | 6 |
Canonical-only features — 15 total (prevalence ≥ 2)
| Feature | Label | Description | Activation | Prevalence |
|---|---|---|---|---|
| #10995 | β-strand–rich extracellular interfaces | β-strand–rich interaction surfaces with strong enrichment in secreted/lumenal protein regions, but not exclusive to them; the feature targets β-strands and adjacent loops that are often post-translationally stabilized (N-glycans) in extracellular proteins, and analogous β-elements in certain intracellular trafficking and phosphoinositide-processing proteins. | 3.14 | 2 |
| #15202 | HEAT/ARM alpha-solenoid scaffold | Alpha-solenoid helical-repeat scaffold (importin-β-like HEAT/ARM repeats) used by karyopherins and other large eukaryotic helical-repeat adaptors; the feature recognizes the continuous arrays of amphipathic α-helices that form flexible HEAT/ARM superhelices underlying nucleocytoplasmic transport and assembly of large complexes. | 2.51 | 2 |
| #2180 | TDG/DG strand-end beta-turn | DG-centered beta-turn motif at the end of β-strands—most often the short Thr–Asp–Gly (TDG) sequence—occurring on exposed β→loop or β→α junctions; prominently realized in von Willebrand factor A (VWA) domains but also present in other folds (e.g., ThDP enzymes) | 2.30 | 2 |
| #9216 | Alpha-helical interface scaffolds | Generic recognition of short, well-ordered α-helical interface segments—helices at domain junctions or within all‑α subdomains that scaffold binding/interaction surfaces rather than catalytic motifs | 2.14 | 2 |
| #6268 | Acidic phosphorylation-rich degron IDRs | Extended, phosphorylation-rich, Ser/Thr/Pro- and acidic-biased intrinsically disordered regulatory regions of eukaryotic nuclear cell-cycle and transcription proteins—especially cyclins—often containing PEST-like segments and degron-bearing zones (e.g., D-box/KEN contexts) used for APC/C/SCF-mediated turnover and phospho-dependent regulation. | 2.13 | 2 |
| #2603 | Conserved extramembrane functional hotspots | Conserved non‑transmembrane functional hotspots: the feature highlights family‑signature loci outside membrane spans—either single hallmark residues within soluble enzyme domains or short linear motifs/patches in cytosolic or extracellular/disordered segments of membrane/regulatory proteins—that serve as key determinants of function (e.g., metal/catalytic positioning, structural anchoring, or interaction/gating surfaces). | 2.10 | 2 |
| #554 | Sliding clamp fold detector | Detector of DNA polymerase processivity clamps—PCNA/β-clamp/9-1-1 and related viral processivity factors—keying on the conserved clamp-domain fold, especially repeated loop→β-strand elements within each clamp domain; it cross-activates weakly on short secondary-structure transition segments that mimic these motifs in unrelated proteins | 2.04 | 3 |
| #3108 | Secondary-structure boundary loops | Short secondary-structure boundary segments—loops/turns at the start of helices or edge β-strands and their immediately adjacent residues—often used as flexible recognition/regulatory sites; these regions are enriched in charged/polar residues (frequent Ser/Thr and Lys/Arg/Asp/Glu) with common Gly/Pro and can include modification or disulfide-forming positions. | 2.00 | 3 |
| #1265 | Aromatic hydrophobic hotspots | Generic detector of bulky aromatic hydrophobic side chains—especially tryptophan, tyrosine, and phenylalanine—in helical/hydrophobic microenvironments. In membrane proteins it emphasizes the aromatic belt at transmembrane helix boundaries and ligand/solute-binding cores; in soluble enzymes it marks buried aromatic cores or pockets (including nucleotide-binding sites). Overall this reflects an “aromatic/hydrophobic microenvironment” signal rather than a family-specific motif. | 1.98 | 2 |
| #8073 | Exposed charged loop/helix patch | A solvent-exposed, charged loop/short amphipathic helix patch at secondary-structure junctions (loop–helix–loop or strand–helix), often flexible or partially disordered, enriched in acidic/basic residues and frequently containing aromatic residues (e.g., Trp/Tyr). These segments commonly serve as interaction/recognition or modification-prone patches positioned adjacent to—but not comprising—catalytic motifs, and occur broadly in extracellular/periplasmic enzymes, membrane lumenal domains, and cytosolic adaptor/enzymatic proteins. | 1.96 | 2 |
| #14986 | Extended assembly interface regions | Extended, compositionally biased segments that mediate assembly or anchoring—often helical or coiled-coil but sometimes including β-strand-rich oligomerization interfaces—located in terminal or accessory regions (basic or acidic low-complexity runs, or Leu/Ile/Val/Phe-rich stretches), rather than in compact catalytic cores. | 1.81 | 2 |
| #3190 | HEAT/ARM α-solenoid linkers | HEAT/armadillo-like α-solenoid scaffolds in large eukaryotic assembly and transport factors, with preference for the inter-repeat loops, helix caps, and flexible linkers adjoining HEAT-repeat arrays (karyopherin/importin-β–like folds, nucleoporin scaffolds, PA200/Blm10-type regulators, MMS19- and Ltn1-like HEAT solenoids). | 1.71 | 2 |
| #8874 | Cationic processing-site microdomains | Short cationic/low-complexity microdomains (≈15–30 aa) that are enriched at proteolytic processing junctions of secreted precursors and within assembly-prone extracellular fibrils (e.g., curli, hydrophobins), and that also occur as similar disordered/basic patches in periplasmic and intracellular proteins. These regions are often Gly-rich and/or cationic, frequently near di/tri-basic motifs and may include the first conserved Cys of cystine-stabilized peptides; they are typically disordered or minimally structured until engaged. | 1.71 | 2 |
| #15071 | Outer-membrane beta-barrels | Outer-envelope exported proteins of Gram-negative bacteria, with strongest signal on the transmembrane beta-strand walls of outer-membrane beta-barrel proteins (porins, TonB-dependent receptors, Type V autotransporters, and Omp85/organellar homologs), and secondary enrichment for Sec-exported periplasmic/secreted enzymes (e.g., sulfatases, glycosidases, oxidoreductases) that share the outer-envelope/secretory context. | 1.68 | 2 |
| #6830 | Conserved short hydrophobic motif detector | A feature that activates on short conserved hydrophobic/aromatic motifs in two distinct contexts—(i) a C-terminal beta-strand "GQIPE(L/V)IFY" motif in BTB/POZ-domain proteins, and (ii) an internal "AGG(F/Y)(I/T)Y(T/T)" motif in the lumenal catalytic region of beta-1,3-GalNAc transferase 2 (B3GALNT2) orthologs. | 1.63 | 3 |
Shared features by |Δ| activation — 883 total (prevalence ≥ 2)
| Feature | Label | Description | Δ activation | Iso act. | Canon act. | Iso prev. | Canon prev. |
|---|---|---|---|---|---|---|---|
| #6384 | Widespread activation in folded proteins | Proteins showing broad, diffuse activation across large fractions of their sequence, with peaks frequently occurring in folded regions of single-domain or multi-domain enzymes and structural proteins. | +5.29 | 8.05 | 2.76 | 52 | 12 |
| #9852 | N-terminus activation; hydrophobic helices secondary | Extreme N-terminal regions of proteins, starting at the initiator methionine and extending over a short stretch, with occasional secondary activation at internal hydrophobic/amphipathic helices. | +5.12 | 7.34 | 2.22 | 25 | 4 |
| #220 | Phenylalanine-rich hydrophobic motif detector | Phenylalanine-focused residue identity feature: detects Phe (F) residues, with a preference for F-rich, hydrophobic stretches (e.g., signal peptides and transmembrane helices), but independent of secondary structure or specific function | +3.89 | 10.67 | 6.78 | 10 | 9 |
| #10117 | Diffuse low-complexity/disorder signature | A broadly tuned signature of disordered/low-complexity-containing proteins: the feature activates diffusely across long stretches of large multi-domain proteins, with residue-level peaks scattered through both folded and disordered segments. It is common in large, repeat-rich extracellular/surface proteins but also marks regions in diverse intracellular proteins. | +3.81 | 5.91 | 2.10 | 19 | 3 |
| #10782 | Lysine residues and KK/KR motifs | Lysine (K) feature that fires on K residues across diverse proteins, with notable amplification at Lys-rich, low-complexity/disordered segments and clustered basic motifs (KK/KR). | +3.03 | 7.30 | 4.27 | 14 | 13 |
| #3503 | Cationic/hydrophobic low-complexity segments | Compositionally biased and low-complexity segments enriched in hydrophobic (L/V/I/A), basic (K/R), and Ser/Thr/Pro residues (with underrepresented Trp). The feature marks both cationic Ser/Thr/Pro-enriched segments and hydrophobic leader-like stretches, capturing N-terminal targeting peptides, micropeptides, and internal polybasic/low-complexity motifs in diverse proteins. | +2.86 | 5.78 | 2.93 | 14 | 6 |
| #6176 | Sparse activation in small proteins | Sparse, low-amplitude activation distributed across small or short proteins, with no strict residue preference and peaks landing in a variety of structural contexts. | +2.68 | 5.89 | 3.21 | 8 | 4 |
| #8050 | Low-confidence intrinsic disorder | Residue-level detector of intrinsically disordered/flexible regions characterized by low predicted structural confidence (low pLDDT), often in low-complexity or unstructured tails, independent of specific secondary-structure assignment | +2.65 | 5.07 | 2.43 | 4 | 2 |
| #8412 | Valine-centered residue detector | Detector that fires on valine residues across diverse sequence contexts, including hydrophobic transmembrane helices and signal peptides as well as Val-containing tandem repeats (elastomeric Gly-rich lamprin-like repeats, Pro/Arg-rich disordered repeats) and isolated Val positions in soluble globular proteins. | +2.45 | 7.66 | 5.21 | 21 | 19 |
| #4900 | Arginine-rich segments and residues | Arginine residue identity/basic-tract feature: primarily marks arginine (R) residues—and dense arginine-rich, low-complexity segments—independent of protein family, fold, or function; shows a strong bias for R over other residues (occasional weak lysine signal), appearing both in disordered/repetitive regions and on individual R within structured domains or near transmembrane boundaries. | +2.40 | 6.98 | 4.58 | 11 | 10 |
| #627 | Generic serine detector | Generic serine detector: the feature marks serine residues, with weaker affinity for threonine (and occasionally proline), and is especially prominent in low-complexity S/T-rich stretches; it is context- and structure-agnostic and does not correspond to a specific functional motif. | +2.33 | 6.86 | 4.53 | 19 | 17 |
| #13480 | Alanine and small-residue enrichment | A sequence-composition feature that detects small, non-aromatic residues—especially alanine—and their enrichment, lighting up individual Ala residues and stretches enriched in A/G/S/T/P/V (alanine-rich, low-complexity segments) irrespective of protein function or secondary structure. | +2.30 | 7.16 | 4.86 | 10 | 8 |
| #11526 | Glycine amidation motif detector | Residue-level detector for small/flexible residues—especially glycine—in short, low-structure linkers and proteolytic processing signals of peptide precursors, with a strong preference for the C‑terminal amidation context in which a glycine (amide donor) immediately precedes mono/di‑basic residues (G‑K/R); outside precursors it gives sparse hits on similar small-residue sites in intrinsically disordered regions and occasionally within structured domains across diverse taxa. | +2.19 | 4.06 | 1.88 | 5 | 3 |
| #8319 | Amphipathic recognition helices | Short amphipathic alpha-helical “recognition” segments used for binding (DNA, membranes, or protein partners), including the recognition helix of helix-turn-helix DNA-binding domains and N‑terminal targeting/signal helices | +1.54 | 3.43 | 1.88 | 6 | 2 |
| #6016 | Helix-biased internal methionine detector | Detector for methionine residues, firing on internal Met across diverse proteins with somewhat enhanced response when Met occurs in helical or low-complexity contexts; occasional hits at the initiator Met when embedded in a locally Met- or hydrophobic-rich N-terminus; weak cross-reactivity to other bulky hydrophobics (notably tryptophan). | +1.37 | 6.73 | 5.36 | 3 | 2 |
| #6803 | Transcription factor activation motifs | Short linear interaction motif–like sites in intrinsically disordered regions of transcription factors, often corresponding to activation/cofactor-binding segments (e.g., the Hox Antp-type hexapeptide/YPWM-containing region), with occasional weaker hits in structured DNA-binding domains | +1.30 | 5.91 | 4.61 | 5 | 3 |
| #12662 | Sparse short-stretch disordered activations | Sparse, short-stretch activations distributed across diverse protein contexts, with a tendency to fire in non-conserved insert regions, disordered/low-complexity segments, and short N-terminal propeptide regions, but also occurring within structured domains and transmembrane helices. | +0.96 | 3.05 | 2.09 | 6 | 5 |
| #14980 | Pre-tryptophan motif detector | A short sequence-context feature that fires at single residues positioned approximately two residues N-terminal to a tryptophan, often within an "xxWφ" pattern (where φ is a hydrophobic residue such as L/M/P/R/E/D). The peak residue itself is variable (S, T, K, N, A, F, L, H, etc.), but the downstream Trp is highly consistent. | +0.94 | 3.91 | 2.97 | 5 | 2 |
| #13897 | Proline-rich IDR detector | Detector of proline residues, with strongest signal in proline-rich, intrinsically disordered, low-complexity segments (often at termini, propeptides, and surface-exposed regions). The feature also activates on more isolated prolines in compact protein contexts, including within transmembrane helices, though typically at lower intensity than in extended Pro-rich disorder. | +0.93 | 3.53 | 2.60 | 8 | 4 |
| #15815 | Basic N-terminal targeting patch | Short, basic/polar N‑terminal leader/transit segment immediately after the initiator methionine (positions ~3–6) — i.e., the positively charged/polar start of signal peptides and organelle transit peptides, and more generally low‑acidic, disordered N‑terminal patches (including occasional cysteine‑rich starts) used for targeting, export, rRNA binding, or membrane topology cues | -0.90 | 6.77 | 7.67 | 8 | 4 |
| #189 | Tryptophan residue detector | A residue-identity detector for tryptophan (W) side chains, with minor cross-activation on other aromatics (Y/F) and small hydrophobics, largely independent of position, domain, structure, or taxonomy; frequently encountered in secreted/viral and disordered contexts but not restricted to them. | +0.74 | 8.59 | 7.84 | 5 | 3 |
| #5330 | Low-complexity disordered N-termini | Low-complexity, intrinsically disordered/propeptide-like segments—typically near N-termini—enriched in Gly/Pro/Ser/Ala and basic (Lys/Arg) residues, often in short repeats or polybasic clusters; the feature highlights flexible coils and secondary-structure junctions (helix/strand–coil boundaries) across very small proteins and larger proteins alike. | +0.62 | 2.84 | 2.21 | 8 | 2 |
| #16243 | Structured-region leucine detector | Generic detector of leucine side chains in structured regions of folded domains, with occasional weaker responses to other hydrophobics; largely indifferent to specific functional motifs or protein class. | +0.61 | 6.66 | 6.05 | 26 | 20 |
| #15970 | Serine/threonine-rich disordered linkers | Intrinsically disordered, low‑complexity serine/threonine–rich segments (often containing SP/TP/RS repeats) that function as flexible linkers and phosphorylation‑prone regulatory tracts across diverse proteins, including viral phosphoproteins/nucleocapsid linkers and proline‑rich secreted cell‑wall proteins | +0.50 | 3.25 | 2.75 | 6 | 4 |
| #6544 | TESPA1/ITPRID cassette hotspot | Mammal-specific, family-restricted short internal segment around residues 200–250 shared by TESPA1 and ITPRID proteins; the model selects a single residue-level hotspot within this segment independent of residue identity. | +0.49 | 7.61 | 7.12 | 16 | 16 |
| #14534 | Unknown generic feature | Unknown generic feature | +12.14 | 30.86 | 18.72 | 222 | 196 |
| #9005 | Unknown generic feature | Unknown generic feature | +11.80 | 23.78 | 11.98 | 171 | 156 |
| #1803 | Unknown generic feature | Unknown generic feature | +7.02 | 23.16 | 16.14 | 217 | 196 |
| #9214 | Unknown generic feature | Unknown generic feature | -2.05 | 14.91 | 16.95 | 222 | 200 |
| #14895 | Unknown generic feature | Unknown generic feature | -1.67 | 19.41 | 21.08 | 225 | 204 |
Part 2 · Differential coordinates
Features firing on just the isoform-unique (extension / alt-frame) residues.
Unique-region features — 205 distinct (prevalence ≥ 2)
| Feature | Label | Description | Activation | Prevalence |
|---|---|---|---|---|
| #6484 | Prion-like low-complexity IDRs | Polar low-complexity intrinsically disordered regions (IDRs)—especially Q/N-rich (prion-like) segments and S/T- and glycine-rich tracts with simple SG/TG/PG repeats—typically located in inter-domain linkers and terminal tails of eukaryotic proteins and often harboring Ser/Thr phosphorylation sites. | 10.38 | 21 |
| #448 | N-terminal leader/targeting segments | N-terminal leader/targeting segments: the feature activates on the extreme N-terminus (first 2–5 residues) of proteins, especially the N-region of signal peptides and organelle transit peptides, and more generally on short, disordered/low-complexity N-terminal tails; strongest at residue ~3 and largely independent of amino-acid identity, occurring across all taxa and functions. | 9.78 | 2 |
| #189 | Tryptophan residue detector | A residue-identity detector for tryptophan (W) side chains, with minor cross-activation on other aromatics (Y/F) and small hydrophobics, largely independent of position, domain, structure, or taxonomy; frequently encountered in secreted/viral and disordered contexts but not restricted to them. | 8.59 | 2 |
| #1136 | Broad HORMA-domain activation N-terminal peak | HORMA-domain meiotic/checkpoint proteins, with activation distributed across the folded HORMA domain itself. | 8.23 | 19 |
| #6384 | Widespread activation in folded proteins | Proteins showing broad, diffuse activation across large fractions of their sequence, with peaks frequently occurring in folded regions of single-domain or multi-domain enzymes and structural proteins. | 8.05 | 21 |
| #9998 | Weak N+3 positional cue | Absolute N-terminal positional cue centered near the fourth residue (N+3), independent of amino-acid identity; most often within signal/transit peptides of precursors but also present in non-secreted proteins; the effect is weak and often does not produce a discrete detected site. | 8.04 | 2 |
| #8412 | Valine-centered residue detector | Detector that fires on valine residues across diverse sequence contexts, including hydrophobic transmembrane helices and signal peptides as well as Val-containing tandem repeats (elastomeric Gly-rich lamprin-like repeats, Pro/Arg-rich disordered repeats) and isolated Val positions in soluble globular proteins. | 7.66 | 2 |
| #9852 | N-terminus activation; hydrophobic helices secondary | Extreme N-terminal regions of proteins, starting at the initiator methionine and extending over a short stretch, with occasional secondary activation at internal hydrophobic/amphipathic helices. | 7.34 | 21 |
| #13480 | Alanine and small-residue enrichment | A sequence-composition feature that detects small, non-aromatic residues—especially alanine—and their enrichment, lighting up individual Ala residues and stretches enriched in A/G/S/T/P/V (alanine-rich, low-complexity segments) irrespective of protein function or secondary structure. | 7.16 | 2 |
| #627 | Generic serine detector | Generic serine detector: the feature marks serine residues, with weaker affinity for threonine (and occasionally proline), and is especially prominent in low-complexity S/T-rich stretches; it is context- and structure-agnostic and does not correspond to a specific functional motif. | 6.86 | 2 |
| #16243 | Structured-region leucine detector | Generic detector of leucine side chains in structured regions of folded domains, with occasional weaker responses to other hydrophobics; largely indifferent to specific functional motifs or protein class. | 6.66 | 6 |
| #7572 | Conserved catalytic helix motif | A conserved internal sequence/structural motif within enzyme catalytic domains, prominently exemplified by the "DRL(V/I)G(x)YEE" segment of the glutamate mutase epsilon subunit, where activation peaks on a short hydrophobic-aromatic stretch embedded in an α-helix close to cofactor/substrate binding residues. | 6.00 | 9 |
| #13334 | N-terminal IDR boundary | Recognition of short, low-complexity, intrinsically disordered N-terminal tails immediately upstream of the first folded domain; these segments are often Ser/Thr/Pro/Gly/Ala-biased and sometimes extend a few residues into the initial helix or the very beginning of the first domain (domain boundary/disorder-to-order transition), especially in transcription regulators and co-regulators. | 5.94 | 21 |
| #6803 | Transcription factor activation motifs | Short linear interaction motif–like sites in intrinsically disordered regions of transcription factors, often corresponding to activation/cofactor-binding segments (e.g., the Hox Antp-type hexapeptide/YPWM-containing region), with occasional weaker hits in structured DNA-binding domains | 5.91 | 2 |
| #10117 | Diffuse low-complexity/disorder signature | A broadly tuned signature of disordered/low-complexity-containing proteins: the feature activates diffusely across long stretches of large multi-domain proteins, with residue-level peaks scattered through both folded and disordered segments. It is common in large, repeat-rich extracellular/surface proteins but also marks regions in diverse intracellular proteins. | 5.91 | 15 |
| #6176 | Sparse activation in small proteins | Sparse, low-amplitude activation distributed across small or short proteins, with no strict residue preference and peaks landing in a variety of structural contexts. | 5.89 | 3 |
| #11633 | Assembly scaffold interaction modules | Protein–protein interaction modules in eukaryotic and bacterial assembly/scaffold proteins, prominently HORMA-domain folds and analogous α/β interaction cores, often adjoining low-complexity Ser/Thr/Pro- and Q/N-rich segments, mediating incorporation into large complexes (kinetochore, autophagy scaffolds, ER–mitochondria tethers, RNA-binding assemblies, dynein complexes, Hsp90 co-chaperone systems, COMMD-containing complexes). | 5.82 | 20 |
| #3503 | Cationic/hydrophobic low-complexity segments | Compositionally biased and low-complexity segments enriched in hydrophobic (L/V/I/A), basic (K/R), and Ser/Thr/Pro residues (with underrepresented Trp). The feature marks both cationic Ser/Thr/Pro-enriched segments and hydrophobic leader-like stretches, capturing N-terminal targeting peptides, micropeptides, and internal polybasic/low-complexity motifs in diverse proteins. | 5.78 | 7 |
| #2302 | Seipin post-TM lumenal/propeptide detector | Detector of post-transmembrane lumenal segments in ER membrane proteins of the seipin family, with additional activation on the propeptide regions of TGF-β superfamily precursors (e.g., GDF-10/BMP-3B). | 5.66 | 17 |
| #7996 | Generic extreme N-terminus detector | Generic extreme N-terminus detector: marks residues around position 5 (occasionally 6–7) at the very start of the polypeptide, most often within the N-region of signal/transit peptides of secreted or organelle-targeted precursors, but also in generally disordered N-termini of cytosolic/nuclear proteins; not motif- or residue-specific | 5.63 | 2 |
| #15815 | Basic N-terminal targeting patch | Short, basic/polar N‑terminal leader/transit segment immediately after the initiator methionine (positions ~3–6) — i.e., the positively charged/polar start of signal peptides and organelle transit peptides, and more generally low‑acidic, disordered N‑terminal patches (including occasional cysteine‑rich starts) used for targeting, export, rRNA binding, or membrane topology cues | 5.52 | 4 |
| #9000 | Glutamate-biased acidic tract detector | Detector of glutamate identity and glutamate-enriched acidic tracts: strong activation on E residues, especially within acidic, low‑complexity/disordered regions; weaker, sporadic responses on D; largely domain-, function-, and taxonomy-agnostic. | 5.49 | 3 |
| #12227 | Tryptophan residue detector | Residue-level detector for tryptophan (W) side chains, often firing on multiple W positions within a sequence and with a mild bias toward N‑terminal Ws; broadly applicable across taxa and protein classes, with frequent occurrences in short membrane/secreted proteins but also in diverse enzymes (e.g., ATP synthase ε, proteases) and mitochondrial proteins. | 5.48 | 2 |
| #14646 | Short N-terminal leader segment | Short N-terminal segments within the first ~10–40 amino acids, often Lys/Arg-enriched but sometimes purely hydrophobic. These commonly correspond to Sec-type signal peptides, signal anchors, or other short N-terminal segments preceding/overlapping the first hydrophobic helix, and also include disordered N-terminal tails in soluble proteins. | 5.40 | 14 |
| #14534 | Unknown generic feature | Unknown generic feature | 30.86 | 21 |
| #9005 | Unknown generic feature | Unknown generic feature | 23.78 | 21 |
| #1803 | Unknown generic feature | Unknown generic feature | 23.16 | 21 |
| #14895 | Unknown generic feature | Unknown generic feature | 19.41 | 21 |
| #9214 | Unknown generic feature | Unknown generic feature | 14.91 | 21 |
| #9194 | Unknown generic feature | Unknown generic feature | 13.80 | 21 |
Top-K sparse-autoencoder activations on the ESM-C residual stream. Activation is a feature's peak value over its region; Prevalence is the number of residues it fires on; Δ activation is isoform − canonical. Only the top 30 features per category are saved; generic / "unknown" features are hidden until you expand a table. Read more about SAE features →
Clinical variants
Differential region — N-terminal extension (isoform-unique)
| Pos (iso) | AA change | Consequence | Source | Clin. sig. | AF (gnomAD) | Impact | AlphaMissense | ESM-C ΔLLR | Link |
|---|---|---|---|---|---|---|---|---|---|
| 0 | L→W | missense_variant | gnomAD | — | 2.78e-06 | — | N/A | — | chr4-120066796-A-C |
| 0 | L→V | missense_variant | gnomAD | — | 9.40e-07 | — | N/A | — | chr4-120066797-A-C |
| 0 | L→L | synonymous_variant | gnomAD | — | 5.64e-06 | — | N/A | — | chr4-120066797-A-G |
| 1 | E→D | missense_variant | gnomAD | — | 8.87e-07 | — | N/A | -0.69 | chr4-120066792-C-A |
| 1 | E→D | missense_variant | gnomAD | — | 3.55e-06 | — | N/A | -0.69 | chr4-120066792-C-G |
| 1 | E→E | synonymous_variant | gnomAD | — | 2.66e-06 | — | N/A | 0.00 | chr4-120066792-C-T |
| 2 | P→L | missense_variant | gnomAD | — | 6.13e-06 | — | N/A | 0.86 | chr4-120066790-G-A |
| 2 | P→Q | missense_variant | gnomAD | — | 4.38e-06 | — | N/A | -0.28 | chr4-120066790-G-T |
| 3 | L→L | synonymous_variant | gnomAD | — | 8.19e-07 | — | N/A | 0.00 | chr4-120066788-G-A |
| 4 | W→* | stop_gained | gnomAD | — | 3.19e-06 | LoF | — | — | chr4-120066783-C-T |
| 4 | W→R | missense_variant | gnomAD | — | 8.13e-07 | — | N/A | 1.16 | chr4-120066785-A-G |
| 5 | L→F | missense_variant | gnomAD | — | 7.84e-07 | — | N/A | -0.91 | chr4-120066780-C-A |
| 5 | L→W | missense_variant | gnomAD | — | 7.86e-07 | — | N/A | -1.79 | chr4-120066781-A-C |
| 6 | — | frameshift_variant | gnomAD | — | 2.19e-04 | LoF | — | — | chr4-120066778-A-AGC |
| 6 | L→P | missense_variant | gnomAD | — | 5.47e-06 | — | N/A | -0.67 | chr4-120066778-A-G |
| 6 | L→L | synonymous_variant | COSMIC | — | — | — | N/A | 0.00 | COSV56636693 |
| 7 | S→S | synonymous_variant | gnomAD | — | 3.74e-06 | — | N/A | 0.00 | chr4-120066774-G-A |
| 7 | S→S | synonymous_variant | gnomAD | — | 2.24e-06 | — | N/A | 0.00 | chr4-120066774-G-C |
| 7 | S→S | synonymous_variant | gnomAD | — | 8.97e-06 | — | N/A | 0.00 | chr4-120066774-G-T |
| 8 | A→A | synonymous_variant | gnomAD | — | 4.41e-06 | — | N/A | 0.00 | chr4-120066771-C-T |
| 8 | A→T | missense_variant | gnomAD | — | 1.50e-06 | — | N/A | -0.98 | chr4-120066773-C-T |
| 9 | E→E | synonymous_variant | gnomAD | — | 7.24e-07 | — | N/A | 0.00 | chr4-120066768-C-T |
| 9 | E→G | missense_variant | gnomAD | — | 7.29e-07 | — | N/A | 0.94 | chr4-120066769-T-C |
| 9 | E→Q | missense_variant | gnomAD | — | 7.35e-07 | — | N/A | -0.45 | chr4-120066770-C-G |
| 10 | W→R | missense_variant | gnomAD | — | 7.25e-07 | — | N/A | 1.77 | chr4-120066767-A-T |
| 11 | K→N | missense_variant | gnomAD | — | 7.10e-07 | — | N/A | -0.84 | chr4-120066762-C-A |
| 11 | K→N | missense_variant | gnomAD | — | 7.10e-07 | — | N/A | -0.84 | chr4-120066762-C-G |
| 11 | K→R | missense_variant | gnomAD | — | 4.97e-06 | — | N/A | 1.09 | chr4-120066763-T-C |
| 11 | K→E | missense_variant | gnomAD | — | 2.14e-06 | — | N/A | -0.06 | chr4-120066764-T-C |
| 12 | R→R | synonymous_variant | gnomAD | — | 1.27e-05 | — | N/A | 0.00 | chr4-120066759-G-A |
| 12 | R→C | missense_variant | gnomAD | — | 1.42e-06 | — | N/A | -0.77 | chr4-120066761-G-A |
| 12 | R→G | missense_variant | gnomAD | — | 2.84e-06 | — | N/A | -0.33 | chr4-120066761-G-C |
| 12 | R→C | missense_variant | COSMIC | — | — | — | N/A | -0.77 | COSV56636912 |
| 13 | V→V | synonymous_variant | gnomAD | — | 8.48e-06 | — | N/A | 0.00 | chr4-120066756-C-T |
| 13 | V→L | missense_variant | gnomAD | — | 2.82e-06 | — | N/A | 0.66 | chr4-120066758-C-A |
| 14 | L→F | missense_variant | gnomAD | — | 9.81e-06 | — | N/A | -0.50 | chr4-120066755-G-A |
| 14 | L→V | missense_variant | gnomAD | — | 1.68e-05 | — | N/A | -0.67 | chr4-120066755-G-C |
| 15 | — | frameshift_variant | gnomAD | — | 6.99e-07 | LoF | — | — | chr4-120066750-CA-C |
| 15 | L→S | missense_variant | gnomAD | — | 6.99e-07 | — | N/A | -0.30 | chr4-120066751-A-G |
| 16 | F→S | missense_variant | gnomAD | — | 1.39e-06 | — | N/A | 1.08 | chr4-120066748-A-G |
| 17 | — | frameshift_variant | gnomAD | — | 6.95e-07 | LoF | — | — | chr4-120066744-CACAA-C |
| 17 | V→E | missense_variant | gnomAD | — | 6.95e-07 | — | N/A | -1.14 | chr4-120066745-A-T |
| 17 | V→L | missense_variant | gnomAD | — | 2.09e-06 | — | N/A | 0.95 | chr4-120066746-C-A |
| 17 | V→L | missense_variant | gnomAD | — | 6.97e-07 | — | N/A | 0.95 | chr4-120066746-C-G |
| 18 | S→F | missense_variant | gnomAD | — | 1.39e-06 | — | N/A | -0.80 | chr4-120066742-G-A |
| 18 | S→Y | missense_variant | gnomAD | — | 2.78e-06 | — | N/A | -2.08 | chr4-120066742-G-T |
| 19 | L→L | synonymous_variant | gnomAD | — | 6.93e-07 | — | N/A | 0.00 | chr4-120066738-C-T |
| 19 | L→R | missense_variant | gnomAD | — | 3.47e-06 | — | N/A | -0.45 | chr4-120066739-A-C |
| 19 | L→Q | missense_variant | gnomAD | — | 1.39e-06 | — | N/A | -1.66 | chr4-120066739-A-T |
| 19 | L→L | synonymous_variant | gnomAD | — | 1.39e-06 | — | N/A | 0.00 | chr4-120066740-G-A |
| 20 | A→A | synonymous_variant | gnomAD | — | 2.77e-06 | — | N/A | 0.00 | chr4-120066735-G-A |
| 20 | A→V | missense_variant | gnomAD | — | 6.93e-07 | — | N/A | -0.58 | chr4-120066736-G-A |
| 20 | A→D | missense_variant | gnomAD | — | 6.93e-07 | — | N/A | -1.67 | chr4-120066736-G-T |
| 20 | A→S | missense_variant | COSMIC | — | — | — | N/A | -0.72 | COSV99633801 |
54 variants in the differential region.
Shared canonical core — sequence common to canonical and isoform; AlphaMissense applies here
| Pos (iso) | AA change | Consequence | Source | Clin. sig. | AF (gnomAD) | Impact | AlphaMissense | ESM-C ΔLLR | Link |
|---|---|---|---|---|---|---|---|---|---|
| 21 | M→T | missense_variant | gnomAD | — | 1.38e-06 | damaging | — | -9.37 | chr4-120066733-A-G |
| 21 | M→K | missense_variant | gnomAD | — | 6.92e-07 | damaging | — | -10.81 | chr4-120066733-A-T |
| 21 | M→V | missense_variant | gnomAD | — | 6.92e-07 | damaging | — | -8.87 | chr4-120066734-T-C |
| 21 | — | frameshift_variant | COSMIC | — | — | LoF | — | — | COSV99634044 |
| 22 | A→A | synonymous_variant | gnomAD | — | 1.38e-06 | — | — | 0.00 | chr4-120066729-C-T |
| 22 | A→V | missense_variant | gnomAD | — | 1.11e-05 | damaging | likely_pathogenic (0.60) | -7.98 | chr4-120066730-G-A |
| 22 | A→E | missense_variant | gnomAD | — | 6.92e-07 | damaging | ambiguous (0.34) | -8.11 | chr4-120066730-G-T |
| 22 | A→S | missense_variant | gnomAD | — | 6.91e-07 | — | likely_benign (0.13) | -5.20 | chr4-120066731-C-A |
| 22 | A→T | missense_variant | gnomAD | — | 1.73e-05 | — | likely_benign (0.33) | -5.05 | chr4-120066731-C-T |
| 23 | L→L | synonymous_variant | gnomAD | — | 6.90e-07 | — | — | 0.00 | chr4-120066726-C-T |
| 23 | L→R | missense_variant | COSMIC | — | — | — | likely_benign (0.05) | -7.25 | COSV99633902 |
| 24 | Q→* | stop_gained | gnomAD | — | 1.38e-06 | LoF | — | — | chr4-120066725-G-A |
| 25 | L→L | synonymous_variant | gnomAD | — | 1.38e-06 | — | — | 0.00 | chr4-120066720-G-A |
| 25 | L→L | synonymous_variant | gnomAD | — | 1.38e-06 | — | — | 0.00 | chr4-120066720-G-C |
| 25 | L→H | missense_variant | gnomAD | — | 1.38e-06 | damaging | likely_benign (0.14) | -8.68 | chr4-120066721-A-T |
| 25 | L→F | missense_variant | gnomAD | — | 3.45e-06 | — | likely_benign (0.11) | -6.81 | chr4-120066722-G-A |
| 25 | L→V | missense_variant | gnomAD | — | 1.38e-06 | — | likely_benign (0.06) | -6.46 | chr4-120066722-G-C |
| 25 | L→I | missense_variant | gnomAD | — | 2.07e-06 | — | likely_benign (0.07) | -6.90 | chr4-120066722-G-T |
| 25 | L→L | synonymous_variant | COSMIC | — | — | — | — | 0.00 | COSV56637843 |
| 26 | S→F | missense_variant | gnomAD | — | 6.90e-07 | damaging | ambiguous (0.34) | -8.62 | chr4-120066718-G-A |
| 27 | R→R | synonymous_variant | gnomAD | — | 4.83e-06 | — | — | 0.00 | chr4-120066714-C-T |
| 27 | R→Q | missense_variant | gnomAD | — | 2.21e-05 | — | likely_benign (0.17) | -6.23 | chr4-120066715-C-T |
| 27 | R→W | missense_variant | gnomAD | — | 6.90e-07 | damaging | ambiguous (0.52) | -8.23 | chr4-120066716-G-A |
| 27 | R→G | missense_variant | gnomAD | — | 1.24e-05 | — | likely_benign (0.12) | -6.55 | chr4-120066716-G-C |
| 27 | R→Q | missense_variant | ClinVar | Uncertain significance | — | — | likely_benign (0.17) | -6.23 | ClinVar:3541916 |
| 28 | E→E | synonymous_variant | gnomAD | — | 6.89e-07 | — | — | 0.00 | chr4-120066711-C-T |
| 28 | E→V | missense_variant | gnomAD | — | 6.89e-07 | damaging | likely_benign (0.32) | -9.29 | chr4-120066712-T-A |
| 28 | — | frameshift_variant | gnomAD | — | 6.90e-07 | LoF | — | — | chr4-120066712-TC-T |
| 28 | E→K | missense_variant | gnomAD | — | 4.83e-06 | damaging | ambiguous (0.44) | -8.23 | chr4-120066713-C-T |
| 29 | Q→Q | synonymous_variant | gnomAD | — | 6.89e-07 | — | — | 0.00 | chr4-120066708-C-T |
| 29 | Q→Q | synonymous_variant | ClinVar | Likely benign | — | — | — | 0.00 | ClinVar:2655059 |
| 29 | Q→* | stop_gained | COSMIC | — | — | LoF | — | — | COSV56636507 |
| 30 | G→G | synonymous_variant | gnomAD | — | 6.89e-07 | — | — | 0.00 | chr4-120066705-T-C |
| 30 | G→R | missense_variant | gnomAD | — | 2.76e-06 | damaging | likely_pathogenic (0.89) | -8.56 | chr4-120066707-C-T |
| 30 | G→R | missense_variant | ClinVar | Uncertain significance | — | damaging | likely_pathogenic (0.89) | -8.56 | ClinVar:2481387 |
| 30 | G→G | synonymous_variant | ClinVar | Likely benign | — | — | — | 0.00 | ClinVar:2655058 |
| 31 | I→M | missense_variant | gnomAD | — | 6.89e-07 | damaging | likely_pathogenic (0.80) | -10.37 | chr4-120066702-G-C |
| 31 | I→T | missense_variant | gnomAD | — | 1.38e-06 | damaging | likely_pathogenic (0.98) | -10.87 | chr4-120066703-A-G |
| 31 | I→V | missense_variant | gnomAD | — | 6.89e-06 | — | likely_benign (0.28) | -7.12 | chr4-120066704-T-C |
| 31 | I→V | missense_variant | ClinVar | Uncertain significance | — | — | likely_benign (0.28) | -7.12 | ClinVar:2490789 |
| 31 | I→I | synonymous_variant | COSMIC | — | — | — | — | 0.00 | COSV106426012 |
| 31 | I→N | missense_variant | COSMIC | — | — | damaging | likely_pathogenic (0.98) | -12.25 | COSV56636743 |
| 32 | T→T | synonymous_variant | gnomAD | — | 6.89e-07 | — | — | 0.00 | chr4-120066699-G-A |
| 33 | L→P | missense_variant | gnomAD | — | 8.95e-06 | damaging | likely_pathogenic (0.99) | -11.44 | chr4-120066697-A-G |
| 33 | L→M | missense_variant | gnomAD | — | 6.89e-07 | damaging | likely_pathogenic (0.81) | -12.94 | chr4-120066698-G-T |
| 33 | L→L | synonymous_variant | ClinVar | Likely benign | — | — | — | 0.00 | ClinVar:2655057 |
| 33 | L→M | missense_variant | COSMIC | — | — | damaging | likely_pathogenic (0.81) | -12.94 | COSV99633896 |
| 34 | R→R | synonymous_variant | gnomAD | — | 6.89e-07 | — | — | 0.00 | chr4-120066693-G-T |
| 35 | G→G | synonymous_variant | gnomAD | — | 2.75e-06 | — | — | 0.00 | chr4-120066690-C-G |
| 35 | G→G | synonymous_variant | gnomAD | — | 1.24e-05 | — | — | 0.00 | chr4-120066690-C-T |
| 35 | G→R | missense_variant | gnomAD | — | 6.89e-07 | damaging | likely_pathogenic (1.00) | -10.12 | chr4-120066692-C-G |
| 35 | G→R | missense_variant | COSMIC | — | — | damaging | likely_pathogenic (1.00) | -10.12 | COSV56637063 |
| 37 | A→A | synonymous_variant | gnomAD | — | 6.89e-07 | — | — | 0.00 | chr4-120066684-G-C |
| 37 | A→A | synonymous_variant | gnomAD | — | 1.38e-06 | — | — | 0.00 | chr4-120066684-G-T |
| 37 | A→V | missense_variant | gnomAD | — | 6.89e-07 | damaging | likely_pathogenic (0.92) | -8.25 | chr4-120066685-G-A |
| 37 | A→S | missense_variant | gnomAD | — | 6.89e-07 | — | likely_benign (0.29) | -7.50 | chr4-120066686-C-A |
| 37 | A→T | missense_variant | gnomAD | — | 6.89e-07 | damaging | likely_pathogenic (0.78) | -7.62 | chr4-120066686-C-T |
| 38 | E→K | missense_variant | gnomAD | — | 1.38e-06 | damaging | likely_pathogenic (0.68) | -9.80 | chr4-120066683-C-T |
| 38 | E→E | synonymous_variant | ClinVar | Likely benign | — | — | — | 0.00 | ClinVar:2655056 |
| 39 | I→M | missense_variant | gnomAD | — | 1.38e-06 | — | likely_benign (0.16) | -7.18 | chr4-120066678-G-C |
| 39 | I→I | synonymous_variant | gnomAD | — | 6.89e-07 | — | — | 0.00 | chr4-120066678-G-T |
| 40 | V→A | missense_variant | gnomAD | — | 6.89e-07 | damaging | likely_pathogenic (0.96) | -10.94 | chr4-120066676-A-G |
| 40 | V→L | missense_variant | gnomAD | — | 2.07e-06 | damaging | likely_pathogenic (0.97) | -11.31 | chr4-120066677-C-G |
| 41 | A→A | synonymous_variant | gnomAD | — | 6.89e-07 | — | — | 0.00 | chr4-120066672-G-A |
| 41 | A→T | missense_variant | gnomAD | — | 6.89e-07 | damaging | likely_pathogenic (0.67) | -7.68 | chr4-120066674-C-T |
| 41 | A→V | missense_variant | COSMIC | — | — | damaging | likely_pathogenic (0.90) | -8.81 | COSV56637607 |
| 42 | E→E | synonymous_variant | gnomAD | — | 8.27e-06 | — | — | 0.00 | chr4-120066669-C-T |
| 42 | E→D | missense_variant | COSMIC | — | — | — | ambiguous (0.48) | -7.37 | COSV99633714 |
| 43 | F→L | missense_variant | gnomAD | — | 6.89e-07 | damaging | likely_pathogenic (1.00) | -10.12 | chr4-120066668-A-G |
| 43 | F→F | synonymous_variant | COSMIC | — | — | — | — | 0.00 | COSV56636758 |
| 43 | F→L | missense_variant | COSMIC | — | — | damaging | likely_pathogenic (1.00) | -10.12 | COSV56636790 |
| 44 | F→F | synonymous_variant | gnomAD | — | 6.89e-07 | — | — | 0.00 | chr4-120066663-G-A |
| 44 | F→S | missense_variant | gnomAD | — | 2.07e-06 | damaging | likely_pathogenic (0.99) | -14.12 | chr4-120066664-A-G |
| 45 | S→S | synonymous_variant | gnomAD | — | 6.85e-07 | — | — | 0.00 | chr4-120065817-T-A |
| 45 | S→A | missense_variant | gnomAD | — | 1.38e-06 | damaging | likely_benign (0.10) | -7.52 | chr4-120066662-A-C |
| 45 | — | frameshift_variant | gnomAD | — | 2.07e-06 | LoF | — | — | chr4-120066662-AG-A |
| 45 | S→L | missense_variant | COSMIC | — | — | damaging | likely_benign (0.27) | -8.56 | COSV56637555 |
| 46 | F→F | synonymous_variant | gnomAD | — | 6.85e-07 | — | — | 0.00 | chr4-120065814-G-A |
| 46 | F→L | missense_variant | gnomAD | — | 8.90e-06 | damaging | likely_pathogenic (0.99) | -8.68 | chr4-120065814-G-C |
| 47 | G→G | synonymous_variant | gnomAD | — | 3.63e-05 | — | — | 0.00 | chr4-120065811-G-A |
| 47 | G→D | missense_variant | gnomAD | — | 6.85e-07 | damaging | likely_pathogenic (1.00) | -12.62 | chr4-120065812-C-T |
| 47 | G→S | missense_variant | gnomAD | — | 2.05e-06 | damaging | likely_pathogenic (0.79) | -7.87 | chr4-120065813-C-T |
| 47 | G→G | synonymous_variant | COSMIC | — | — | — | — | 0.00 | COSV56636895 |
| 48 | I→I | synonymous_variant | gnomAD | — | 6.85e-07 | — | — | 0.00 | chr4-120065808-G-A |
| 48 | I→F | missense_variant | gnomAD | — | 6.85e-07 | damaging | likely_pathogenic (0.85) | -11.62 | chr4-120065810-T-A |
| 52 | L→L | synonymous_variant | gnomAD | — | 1.57e-05 | — | — | 0.00 | chr4-120065796-T-C |
| 53 | Y→* | stop_gained | gnomAD | — | 6.84e-07 | LoF | — | — | chr4-120065793-A-T |
| 53 | Y→H | missense_variant | gnomAD | — | 6.84e-07 | damaging | likely_pathogenic (0.96) | -8.81 | chr4-120065795-A-G |
| 55 | R→H | missense_variant | gnomAD | — | 1.37e-06 | damaging | likely_pathogenic (0.99) | -11.31 | chr4-120065788-C-T |
| 55 | R→C | missense_variant | COSMIC | — | — | damaging | likely_pathogenic (0.99) | -11.44 | COSV56636453 |
| 56 | G→G | synonymous_variant | gnomAD | — | 2.05e-06 | — | — | 0.00 | chr4-120065784-G-C |
| 57 | I→M | missense_variant | gnomAD | — | 2.05e-06 | damaging | likely_pathogenic (0.67) | -8.43 | chr4-120065781-T-C |
| 57 | I→V | missense_variant | gnomAD | — | 6.84e-07 | — | likely_benign (0.13) | -6.25 | chr4-120065783-T-C |
| 57 | I→L | missense_variant | gnomAD | — | 6.84e-07 | — | likely_benign (0.23) | -7.18 | chr4-120065783-T-G |
| 57 | I→M | missense_variant | COSMIC | — | — | damaging | likely_pathogenic (0.67) | -8.43 | COSV56637152 |
| 58 | Y→S | missense_variant | gnomAD | — | 2.05e-06 | damaging | likely_pathogenic (0.96) | -12.62 | chr4-120065779-T-G |
| 58 | Y→C | missense_variant | ClinVar | Uncertain significance | — | damaging | likely_pathogenic (0.91) | -12.06 | ClinVar:2337132 |
| 59 | P→P | synonymous_variant | gnomAD | — | 6.84e-07 | — | — | 0.00 | chr4-120065775-T-C |
| 60 | S→S | synonymous_variant | gnomAD | — | 6.84e-06 | — | — | 0.00 | chr4-120065772-A-T |
| 60 | S→C | missense_variant | gnomAD | — | 2.05e-06 | damaging | likely_benign (0.15) | -9.02 | chr4-120065773-G-C |
| 62 | T→T | synonymous_variant | gnomAD | — | 2.05e-06 | — | — | 0.00 | chr4-120065766-G-A |
| 62 | T→N | missense_variant | gnomAD | — | 6.84e-07 | damaging | likely_benign (0.16) | -8.18 | chr4-120065767-G-T |
| 64 | T→I | missense_variant | gnomAD | — | 3.42e-06 | — | likely_benign (0.26) | -6.92 | chr4-120065761-G-A |
| 64 | T→I | missense_variant | ClinVar | Uncertain significance | — | — | likely_benign (0.26) | -6.92 | ClinVar:2314971 |
| 65 | R→G | missense_variant | gnomAD | — | 6.84e-07 | damaging | likely_pathogenic (0.82) | -9.50 | chr4-120065759-G-C |
| 65 | R→R | synonymous_variant | gnomAD | — | 6.84e-07 | — | — | 0.00 | chr4-120065759-G-T |
| 65 | R→* | stop_gained | COSMIC | — | — | LoF | — | — | COSV56636783 |
| 66 | V→V | synonymous_variant | gnomAD | — | 1.37e-06 | — | — | 0.00 | chr4-120065754-C-T |
| 66 | V→A | missense_variant | gnomAD | — | 6.84e-07 | — | ambiguous (0.39) | -5.87 | chr4-120065755-A-G |
| 67 | Q→P | missense_variant | gnomAD | — | 2.05e-06 | damaging | likely_benign (0.27) | -8.50 | chr4-120065752-T-G |
| 67 | Q→* | stop_gained | COSMIC | — | — | LoF | — | — | COSV99634048 |
| 67 | Q→Q | synonymous_variant | COSMIC | — | — | — | — | 0.00 | COSV56636984 |
| 67 | Q→R | missense_variant | COSMIC | — | — | damaging | likely_benign (0.32) | -9.25 | COSV105896543 |
| 68 | K→* | stop_gained | COSMIC | — | — | LoF | — | — | COSV56636610 |
| 69 | Y→Y | synonymous_variant | gnomAD | — | 4.31e-05 | — | — | 0.00 | chr4-120065745-G-A |
| 69 | Y→C | missense_variant | COSMIC | — | — | damaging | likely_pathogenic (0.98) | -9.81 | COSV56636603 |
| 70 | G→R | missense_variant | COSMIC | — | — | damaging | likely_pathogenic (0.97) | -10.25 | COSV106426007 |
| 70 | G→* | stop_gained | COSMIC | — | — | LoF | — | — | COSV56636827 |
| 71 | L→F | missense_variant | gnomAD | — | 1.37e-06 | damaging | likely_pathogenic (0.91) | -9.62 | chr4-120065741-G-A |
| 72 | T→T | synonymous_variant | gnomAD | — | 1.37e-06 | — | — | 0.00 | chr4-120065736-G-A |
| 72 | T→T | synonymous_variant | gnomAD | — | 6.84e-07 | — | — | 0.00 | chr4-120065736-G-C |
| 72 | T→I | missense_variant | gnomAD | — | 6.84e-07 | damaging | likely_pathogenic (0.77) | -8.68 | chr4-120065737-G-A |
| 73 | L→F | missense_variant | gnomAD | — | 6.84e-07 | damaging | likely_pathogenic (0.77) | -9.87 | chr4-120065733-C-G |
| 73 | — | frameshift_variant | gnomAD | — | 8.21e-06 | LoF | — | — | chr4-120065733-CA-C |
| 73 | L→W | missense_variant | gnomAD | — | 1.37e-06 | damaging | likely_pathogenic (0.98) | -13.31 | chr4-120065734-A-C |
| 74 | L→F | missense_variant | gnomAD | — | 6.84e-07 | damaging | likely_benign (0.26) | -8.69 | chr4-120065732-G-A |
| 75 | V→L | missense_variant | COSMIC | — | — | — | likely_benign (0.32) | -7.34 | COSV56636970 |
| 77 | T→A | missense_variant | gnomAD | — | 2.12e-05 | — | likely_benign (0.10) | -5.80 | chr4-120065723-T-C |
| 80 | E→Q | missense_variant | gnomAD | — | 6.84e-07 | damaging | likely_benign (0.13) | -7.58 | chr4-120065714-C-G |
| 82 | — | frameshift_variant | gnomAD | — | 6.84e-07 | LoF | — | — | chr4-120065708-TG-T |
| 83 | K→R | missense_variant | gnomAD | — | 6.84e-07 | — | likely_benign (0.08) | -5.92 | chr4-120065704-T-C |
| 83 | K→* | stop_gained | gnomAD | — | 6.84e-07 | LoF | — | — | chr4-120065705-T-A |
| 84 | Y→F | missense_variant | gnomAD | — | 2.74e-06 | — | ambiguous (0.38) | -6.94 | chr4-120065701-T-A |
| 84 | Y→C | missense_variant | gnomAD | — | 6.84e-07 | damaging | likely_pathogenic (0.96) | -10.00 | chr4-120065701-T-C |
| 84 | Y→N | missense_variant | COSMIC | — | — | damaging | likely_pathogenic (0.98) | -10.81 | COSV99633708 |
| 85 | L→L | synonymous_variant | gnomAD | — | 1.37e-06 | — | — | 0.00 | chr4-120065697-T-C |
| 85 | L→L | synonymous_variant | gnomAD | — | 1.37e-06 | — | — | 0.00 | chr4-120065697-T-G |
| 85 | — | frameshift_variant | gnomAD | — | 6.84e-07 | LoF | — | — | chr4-120065697-TA-T |
| 85 | L→I | missense_variant | gnomAD | — | 1.21e-04 | damaging | likely_benign (0.17) | -8.31 | chr4-120065699-G-T |
| 86 | N→S | missense_variant | gnomAD | — | 2.74e-06 | — | likely_benign (0.07) | -3.68 | chr4-120065695-T-C |
| 86 | N→N | synonymous_variant | COSMIC | — | — | — | — | 0.00 | COSV108109490 |
| 87 | N→N | synonymous_variant | gnomAD | — | 2.05e-06 | — | — | 0.00 | chr4-120065691-A-G |
| 87 | N→Y | missense_variant | gnomAD | — | 1.37e-06 | damaging | likely_benign (0.16) | -8.23 | chr4-120065693-T-A |
| 87 | N→Y | missense_variant | ClinVar | Uncertain significance | — | damaging | likely_benign (0.16) | -8.23 | ClinVar:3869470 |
| 88 | V→G | missense_variant | gnomAD | — | 6.84e-07 | damaging | likely_pathogenic (0.89) | -11.56 | chr4-120065689-A-C |
| 88 | V→L | missense_variant | gnomAD | — | 1.37e-06 | damaging | likely_pathogenic (0.88) | -8.69 | chr4-120065690-C-A |
| 90 | E→K | missense_variant | COSMIC | — | — | damaging | likely_benign (0.14) | -7.81 | COSV56636425 |
| 91 | Q→Q | synonymous_variant | gnomAD | — | 1.57e-05 | — | — | 0.00 | chr4-120065679-T-C |
| 92 | L→L | synonymous_variant | gnomAD | — | 4.11e-06 | — | — | 0.00 | chr4-120065676-C-T |
| 93 | K→R | missense_variant | gnomAD | — | 6.84e-07 | — | likely_benign (0.10) | -4.08 | chr4-120065674-T-C |
| 94 | D→G | missense_variant | gnomAD | — | 6.88e-07 | — | likely_benign (0.12) | -5.71 | chr4-120062095-T-C |
| 94 | D→H | missense_variant | gnomAD | — | 6.85e-07 | damaging | likely_benign (0.28) | -9.40 | chr4-120065672-C-G |
| 95 | W→L | missense_variant | COSMIC | — | — | damaging | likely_pathogenic (0.93) | -10.50 | COSV99633853 |
| 96 | L→L | synonymous_variant | gnomAD | — | 6.86e-07 | — | — | 0.00 | chr4-120062088-T-C |
| 96 | — | frameshift_variant | gnomAD | — | 6.18e-06 | LoF | — | — | chr4-120062090-A-AC |
| 97 | Y→C | missense_variant | gnomAD | — | 2.74e-06 | — | likely_benign (0.07) | -5.45 | chr4-120062086-T-C |
| 97 | Y→F | missense_variant | COSMIC | — | — | — | likely_benign (0.08) | -4.57 | COSV56636540 |
| 97 | Y→H | missense_variant | COSMIC | — | — | — | likely_benign (0.13) | -6.64 | COSV56637585 |
| 98 | K→N | missense_variant | gnomAD | — | 1.51e-05 | — | likely_benign (0.19) | -6.30 | chr4-120062082-C-G |
| 98 | K→K | synonymous_variant | gnomAD | — | 3.43e-06 | — | — | 0.00 | chr4-120062082-C-T |
| 98 | K→R | missense_variant | gnomAD | — | 1.23e-05 | — | likely_benign (0.07) | -5.96 | chr4-120062083-T-C |
| 98 | K→* | stop_gained | COSMIC | — | — | LoF | — | — | COSV56637871 |
| 99 | C→Y | missense_variant | gnomAD | — | 1.37e-06 | damaging | ambiguous (0.51) | -8.62 | chr4-120062080-C-T |
| 99 | C→S | missense_variant | gnomAD | — | 2.06e-06 | — | ambiguous (0.41) | -7.19 | chr4-120062081-A-T |
| 100 | S→S | synonymous_variant | gnomAD | — | 1.37e-06 | — | — | 0.00 | chr4-120062076-T-A |
| 100 | S→S | synonymous_variant | gnomAD | — | 6.85e-07 | — | — | 0.00 | chr4-120062076-T-C |
| 100 | S→L | missense_variant | gnomAD | — | 6.85e-07 | — | likely_benign (0.10) | -6.53 | chr4-120062077-G-A |
| 100 | S→* | stop_gained | gnomAD | — | 6.85e-07 | LoF | — | — | chr4-120062077-G-C |
| 100 | S→L | missense_variant | COSMIC | — | — | — | likely_benign (0.10) | -6.53 | COSV56636668 |
| 102 | Q→R | missense_variant | gnomAD | — | 2.06e-06 | damaging | likely_benign (0.32) | -8.62 | chr4-120062071-T-C |
| 102 | Q→* | stop_gained | COSMIC | — | — | LoF | — | — | COSV56636702 |
| 104 | L→L | synonymous_variant | gnomAD | — | 1.37e-06 | — | — | 0.00 | chr4-120062064-C-A |
| 104 | L→L | synonymous_variant | COSMIC | — | — | — | — | 0.00 | COSV56637689 |
| 105 | V→V | synonymous_variant | gnomAD | — | 6.85e-07 | — | — | 0.00 | chr4-120062061-A-T |
| 105 | V→F | missense_variant | gnomAD | — | 6.85e-07 | damaging | likely_pathogenic (0.99) | -12.31 | chr4-120062063-C-A |
| 106 | V→L | missense_variant | gnomAD | — | 6.85e-07 | — | likely_benign (0.27) | -7.25 | chr4-120062060-C-G |
| 106 | V→G | missense_variant | COSMIC | — | — | damaging | likely_pathogenic (0.92) | -11.81 | COSV106426017 |
| 107 | V→V | synonymous_variant | gnomAD | — | 4.79e-06 | — | — | 0.00 | chr4-120062055-A-G |
| 107 | V→I | missense_variant | gnomAD | — | 6.16e-06 | damaging | likely_benign (0.22) | -8.50 | chr4-120062057-C-T |
| 108 | I→I | synonymous_variant | gnomAD | — | 4.11e-06 | — | — | 0.00 | chr4-120062052-G-A |
| 108 | I→T | missense_variant | gnomAD | — | 1.37e-06 | damaging | likely_pathogenic (0.98) | -8.56 | chr4-120062053-A-G |
| 110 | N→D | missense_variant | gnomAD | — | 6.16e-06 | — | likely_benign (0.07) | -6.80 | chr4-120062048-T-C |
| 110 | N→D | missense_variant | ClinVar | Uncertain significance | — | — | likely_benign (0.07) | -6.80 | ClinVar:3121965 |
| 111 | I→V | missense_variant | gnomAD | — | 6.09e-05 | — | likely_benign (0.06) | -4.88 | chr4-120062045-T-C |
| 111 | I→N | missense_variant | COSMIC | — | — | damaging | likely_benign (0.31) | -8.57 | COSV56636977 |
| 111 | I→V | missense_variant | COSMIC | — | — | — | likely_benign (0.06) | -4.88 | COSV99047961 |
| 113 | S→R | missense_variant | gnomAD | — | 6.85e-07 | damaging | ambiguous (0.48) | -9.10 | chr4-120062037-A-C |
| 113 | S→S | synonymous_variant | gnomAD | — | 1.37e-06 | — | — | 0.00 | chr4-120062037-A-G |
| 113 | S→T | missense_variant | gnomAD | — | 6.85e-07 | — | likely_benign (0.06) | -3.75 | chr4-120062038-C-G |
| 113 | S→N | missense_variant | gnomAD | — | 1.16e-05 | — | likely_benign (0.07) | -5.35 | chr4-120062038-C-T |
| 113 | S→R | missense_variant | COSMIC | — | — | damaging | ambiguous (0.48) | -9.10 | COSV56636802 |
| 114 | G→V | missense_variant | gnomAD | — | 6.85e-07 | damaging | ambiguous (0.40) | -10.24 | chr4-120062035-C-A |
| 114 | G→D | missense_variant | gnomAD | — | 1.37e-06 | damaging | likely_benign (0.21) | -8.74 | chr4-120062035-C-T |
| 115 | E→K | missense_variant | gnomAD | — | 1.44e-05 | damaging | likely_benign (0.33) | -9.75 | chr4-120062033-C-T |
| 115 | E→K | missense_variant | COSMIC | — | — | damaging | likely_benign (0.33) | -9.75 | COSV105174397 |
| 116 | V→V | synonymous_variant | gnomAD | — | 6.85e-07 | — | — | 0.00 | chr4-120062028-G-A |
| 117 | L→L | synonymous_variant | gnomAD | — | 1.37e-06 | — | — | 0.00 | chr4-120062025-C-T |
| 117 | L→V | missense_variant | gnomAD | — | 6.84e-07 | damaging | likely_benign (0.15) | -8.75 | chr4-120062027-G-C |
| 119 | R→I | missense_variant | gnomAD | — | 6.85e-07 | damaging | likely_pathogenic (0.99) | -13.75 | chr4-120062020-C-A |
| 121 | Q→R | missense_variant | gnomAD | — | 2.74e-06 | damaging | likely_pathogenic (0.85) | -10.81 | chr4-120062014-T-C |
| 121 | Q→P | missense_variant | gnomAD | — | 6.85e-07 | damaging | likely_pathogenic (0.99) | -12.12 | chr4-120062014-T-G |
| 122 | F→L | missense_variant | gnomAD | — | 6.85e-07 | damaging | likely_pathogenic (1.00) | -10.44 | chr4-120062012-A-G |
| 123 | D→V | missense_variant | gnomAD | — | 6.85e-07 | damaging | likely_pathogenic (0.81) | -10.75 | chr4-120062008-T-A |
| 123 | D→N | missense_variant | COSMIC | — | — | damaging | ambiguous (0.38) | -9.00 | COSV109419265 |
| 124 | I→I | synonymous_variant | gnomAD | — | 4.11e-06 | — | — | 0.00 | chr4-120062004-A-T |
| 124 | I→T | missense_variant | gnomAD | — | 3.42e-06 | damaging | likely_pathogenic (0.79) | -9.06 | chr4-120062005-A-G |
| 125 | E→Q | missense_variant | gnomAD | — | 6.85e-07 | damaging | likely_benign (0.18) | -7.62 | chr4-120062003-C-G |
| 126 | C→Y | missense_variant | gnomAD | — | 2.74e-06 | damaging | likely_pathogenic (0.82) | -9.50 | chr4-120061999-C-T |
| 127 | D→E | missense_variant | gnomAD | — | 2.06e-06 | damaging | likely_pathogenic (0.61) | -8.56 | chr4-120061995-G-C |
| 128 | K→M | missense_variant | COSMIC | — | — | damaging | ambiguous (0.41) | -9.25 | COSV56637840 |
| 129 | T→T | synonymous_variant | gnomAD | — | 1.24e-05 | — | — | 0.00 | chr4-120061989-A-G |
| 129 | T→T | synonymous_variant | COSMIC | — | — | — | — | 0.00 | COSV56637816 |
| 130 | A→A | synonymous_variant | gnomAD | — | 6.88e-07 | — | — | 0.00 | chr4-120061986-T-G |
| 130 | A→V | missense_variant | gnomAD | — | 6.89e-07 | — | likely_benign (0.10) | -5.11 | chr4-120061987-G-A |
| 130 | A→S | missense_variant | gnomAD | — | 2.07e-06 | — | likely_benign (0.12) | -7.11 | chr4-120061988-C-A |
| 130 | A→V | missense_variant | COSMIC | — | — | — | likely_benign (0.10) | -5.11 | COSV56636463 |
| 131 | K→E | missense_variant | gnomAD | — | 1.38e-06 | damaging | likely_benign (0.13) | -7.83 | chr4-120061985-T-C |
| 132 | D→D | synonymous_variant | gnomAD | — | 6.90e-07 | — | — | 0.00 | chr4-120061980-A-G |
| 132 | D→Y | missense_variant | gnomAD | — | 3.45e-06 | damaging | likely_benign (0.26) | -8.36 | chr4-120061982-C-A |
| 133 | D→Y | missense_variant | gnomAD | — | 6.90e-07 | damaging | likely_benign (0.19) | -7.57 | chr4-120061979-C-A |
| 134 | S→S | synonymous_variant | gnomAD | — | 5.60e-06 | — | — | 0.00 | chr4-120060977-A-G |
| 134 | S→I | missense_variant | ClinVar | — | — | damaging | likely_benign (0.15) | -7.89 | ClinVar:4301842 |
| 134 | S→N | missense_variant | COSMIC | — | — | — | likely_benign (0.12) | -5.11 | COSV105174411 |
| 135 | A→A | synonymous_variant | gnomAD | — | 1.39e-06 | — | — | 0.00 | chr4-120060974-T-C |
| 135 | A→V | missense_variant | gnomAD | — | 2.79e-06 | — | likely_benign (0.07) | -3.28 | chr4-120060975-G-A |
| 135 | A→A | synonymous_variant | COSMIC | — | — | — | — | 0.00 | COSV56636648 |
| 136 | P→L | missense_variant | gnomAD | — | 2.08e-06 | — | likely_benign (0.25) | -6.75 | chr4-120060972-G-A |
| 136 | P→A | missense_variant | gnomAD | — | 1.39e-06 | damaging | likely_benign (0.13) | -7.56 | chr4-120060973-G-C |
| 136 | P→T | missense_variant | COSMIC | — | — | damaging | likely_benign (0.21) | -7.84 | COSV99633915 |
| 137 | R→S | missense_variant | gnomAD | — | 6.89e-07 | damaging | likely_pathogenic (0.92) | -8.87 | chr4-120060968-T-A |
| 137 | R→I | missense_variant | gnomAD | — | 6.92e-07 | damaging | likely_pathogenic (0.76) | -10.80 | chr4-120060969-C-A |
| 137 | R→G | missense_variant | gnomAD | — | 2.07e-06 | damaging | likely_pathogenic (0.70) | -7.90 | chr4-120060970-T-C |
| 139 | K→K | synonymous_variant | gnomAD | — | 6.87e-07 | — | — | 0.00 | chr4-120060962-C-T |
| 139 | — | frameshift_variant | gnomAD | — | 6.87e-07 | LoF | — | — | chr4-120060962-CT-C |
| 139 | K→M | missense_variant | gnomAD | — | 1.99e-05 | damaging | likely_pathogenic (0.82) | -10.31 | chr4-120060963-T-A |
| 139 | K→E | missense_variant | gnomAD | — | 6.88e-07 | damaging | likely_pathogenic (0.96) | -11.06 | chr4-120060964-T-C |
| 140 | S→S | synonymous_variant | gnomAD | — | 1.37e-06 | — | — | 0.00 | chr4-120060959-A-G |
| 140 | S→F | missense_variant | gnomAD | — | 5.70e-05 | damaging | likely_pathogenic (0.80) | -9.05 | chr4-120060960-G-A |
| 140 | S→C | missense_variant | gnomAD | — | 1.37e-06 | — | likely_benign (0.21) | -7.18 | chr4-120060960-G-C |
| 140 | S→C | missense_variant | ClinVar | Uncertain significance | — | — | likely_benign (0.21) | -7.18 | ClinVar:2624283 |
| 141 | Q→H | missense_variant | gnomAD | — | 2.06e-06 | — | likely_benign (0.21) | -6.53 | chr4-120060956-C-G |
| 141 | Q→P | missense_variant | gnomAD | — | 3.29e-05 | — | likely_benign (0.10) | -7.21 | chr4-120060957-T-G |
| 142 | K→K | synonymous_variant | gnomAD | — | 6.86e-07 | — | — | 0.00 | chr4-120060953-T-C |
| 142 | K→T | missense_variant | gnomAD | — | 6.86e-07 | damaging | ambiguous (0.37) | -8.37 | chr4-120060954-T-G |
| 142 | K→E | missense_variant | gnomAD | — | 6.86e-07 | damaging | ambiguous (0.42) | -7.78 | chr4-120060955-T-C |
| 145 | Q→R | missense_variant | gnomAD | — | 3.43e-06 | damaging | likely_benign (0.30) | -9.00 | chr4-120060945-T-C |
| 145 | Q→* | stop_gained | COSMIC | — | — | LoF | — | — | COSV99633770 |
| 146 | D→N | missense_variant | gnomAD | — | 6.85e-07 | damaging | likely_benign (0.15) | -8.18 | chr4-120060943-C-T |
| 146 | D→N | missense_variant | ClinVar | Uncertain significance | — | damaging | likely_benign (0.15) | -8.18 | ClinVar:3121966 |
| 148 | I→V | missense_variant | gnomAD | — | 2.06e-06 | damaging | likely_pathogenic (0.65) | -9.37 | chr4-120060937-T-C |
| 148 | I→I | synonymous_variant | COSMIC | — | — | — | — | 0.00 | COSV56637007 |
| 148 | I→V | missense_variant | COSMIC | — | — | damaging | likely_pathogenic (0.65) | -9.37 | COSV56636872 |
| 149 | R→L | missense_variant | gnomAD | — | 2.06e-06 | damaging | likely_pathogenic (0.96) | -10.50 | chr4-120060933-C-A |
| 149 | R→H | missense_variant | gnomAD | — | 6.85e-06 | damaging | likely_pathogenic (0.74) | -8.75 | chr4-120060933-C-T |
| 149 | R→C | missense_variant | gnomAD | — | 3.43e-05 | damaging | likely_pathogenic (0.61) | -8.56 | chr4-120060934-G-A |
| 149 | R→G | missense_variant | gnomAD | — | 6.85e-07 | damaging | likely_pathogenic (0.86) | -10.06 | chr4-120060934-G-C |
| 149 | R→S | missense_variant | gnomAD | — | 6.85e-07 | damaging | likely_pathogenic (0.93) | -9.37 | chr4-120060934-G-T |
| 149 | R→C | missense_variant | COSMIC | — | — | damaging | likely_pathogenic (0.61) | -8.56 | COSV99633932 |
| 150 | S→A | missense_variant | gnomAD | — | 6.85e-07 | — | likely_benign (0.12) | -6.47 | chr4-120060931-A-C |
| 150 | S→* | stop_gained | COSMIC | — | — | LoF | — | — | COSV56636444 |
| 151 | V→V | synonymous_variant | gnomAD | — | 6.85e-07 | — | — | 0.00 | chr4-120060926-C-G |
| 152 | I→M | missense_variant | gnomAD | — | 8.22e-06 | damaging | likely_benign (0.27) | -8.94 | chr4-120060923-G-C |
| 153 | R→K | missense_variant | COSMIC | — | — | damaging | likely_pathogenic (0.79) | -10.37 | COSV56637106 |
| 153 | R→T | missense_variant | COSMIC | — | — | damaging | likely_pathogenic (1.00) | -13.44 | COSV56636656 |
| 154 | Q→Q | synonymous_variant | COSMIC | — | — | — | — | 0.00 | COSV56637581 |
| 156 | T→T | synonymous_variant | gnomAD | — | 6.86e-07 | — | — | 0.00 | chr4-120060911-T-C |
| 156 | T→I | missense_variant | COSMIC | — | — | damaging | likely_pathogenic (0.94) | -9.81 | COSV56637834 |
| 157 | A→A | synonymous_variant | gnomAD | — | 6.86e-07 | — | — | 0.00 | chr4-120060908-A-C |
| 157 | A→D | missense_variant | COSMIC | — | — | damaging | likely_pathogenic (1.00) | -11.12 | COSV56637574 |
| 158 | T→T | synonymous_variant | gnomAD | — | 1.10e-05 | — | — | 0.00 | chr4-120060905-C-T |
| 158 | T→M | missense_variant | gnomAD | — | 2.06e-06 | damaging | likely_pathogenic (0.69) | -9.56 | chr4-120060906-G-A |
| 158 | T→T | synonymous_variant | COSMIC | — | — | — | — | 0.00 | COSV56637199 |
| 158 | T→M | missense_variant | COSMIC | — | — | damaging | likely_pathogenic (0.69) | -9.56 | COSV99633752 |
| 159 | V→V | synonymous_variant | gnomAD | — | 6.87e-07 | — | — | 0.00 | chr4-120060902-C-T |
| 159 | V→V | synonymous_variant | COSMIC | — | — | — | — | 0.00 | COSV56637231 |
| 159 | V→A | missense_variant | COSMIC | — | — | damaging | likely_pathogenic (0.97) | -10.50 | COSV56637826 |
| 163 | P→P | synonymous_variant | gnomAD | — | 6.18e-02 | — | — | 0.00 | chr4-120060890-T-C |
| 163 | P→L | missense_variant | gnomAD | — | 6.89e-07 | damaging | likely_pathogenic (0.99) | -11.69 | chr4-120060891-G-A |
| 163 | P→P | synonymous_variant | COSMIC | — | — | — | — | 0.00 | COSV56636944 |
| 164 | L→L | synonymous_variant | gnomAD | — | 6.89e-07 | — | — | 0.00 | chr4-120060887-C-G |
| 166 | E→D | missense_variant | COSMIC | — | — | — | likely_benign (0.09) | -6.00 | COSV56637746 |
| 166 | E→Q | missense_variant | COSMIC | — | — | damaging | likely_benign (0.34) | -8.68 | COSV56636515 |
| 167 | V→A | missense_variant | gnomAD | — | 6.91e-07 | — | likely_benign (0.09) | -3.35 | chr4-120060879-A-G |
| 167 | V→F | missense_variant | gnomAD | — | 6.91e-07 | damaging | likely_benign (0.17) | -8.22 | chr4-120060880-C-A |
| 168 | S→F | missense_variant | gnomAD | — | 6.92e-07 | damaging | likely_benign (0.33) | -8.87 | chr4-120060876-G-A |
| 168 | S→Y | missense_variant | gnomAD | — | 1.38e-06 | damaging | likely_benign (0.29) | -9.62 | chr4-120060876-G-T |
| 168 | — | frameshift_variant | gnomAD | — | 6.92e-07 | LoF | — | — | chr4-120060876-GA-G |
| 169 | C→Y | missense_variant | gnomAD | — | 3.45e-06 | damaging | likely_pathogenic (0.97) | -10.19 | chr4-120060290-C-T |
| 169 | C→Y | missense_variant | ClinVar | Uncertain significance | — | damaging | likely_pathogenic (0.97) | -10.19 | ClinVar:4050418 |
| 169 | C→R | missense_variant | COSMIC | — | — | damaging | likely_pathogenic (0.98) | -11.37 | COSV56637100 |
| 170 | S→L | missense_variant | gnomAD | — | 6.88e-07 | damaging | ambiguous (0.39) | -9.55 | chr4-120060287-G-A |
| 170 | S→* | stop_gained | COSMIC | — | — | LoF | — | — | COSV56637188 |
| 172 | D→E | missense_variant | gnomAD | — | 1.37e-06 | damaging | likely_pathogenic (0.91) | -8.31 | chr4-120060280-A-T |
| 172 | D→N | missense_variant | COSMIC | — | — | damaging | likely_pathogenic (0.70) | -9.06 | COSV108814783 |
| 173 | L→L | synonymous_variant | gnomAD | — | 2.06e-06 | — | — | 0.00 | chr4-120060277-C-T |
| 173 | L→P | missense_variant | gnomAD | — | 2.06e-06 | damaging | likely_pathogenic (1.00) | -11.37 | chr4-120060278-A-G |
| 174 | L→L | synonymous_variant | gnomAD | — | 2.06e-05 | — | — | 0.00 | chr4-120060274-C-T |
| 174 | L→V | missense_variant | gnomAD | — | 2.06e-06 | damaging | likely_pathogenic (0.82) | -9.62 | chr4-120060276-G-C |
| 175 | — | frameshift_variant | COSMIC | — | — | LoF | — | — | COSV99633904 |
| 176 | Y→D | missense_variant | gnomAD | — | 6.86e-07 | damaging | likely_pathogenic (0.99) | -11.12 | chr4-120060270-A-C |
| 177 | T→I | missense_variant | gnomAD | — | 6.86e-07 | damaging | likely_pathogenic (0.96) | -6.87 | chr4-120060266-G-A |
| 177 | T→A | missense_variant | gnomAD | — | 6.85e-07 | — | ambiguous (0.36) | -6.75 | chr4-120060267-T-C |
| 177 | T→A | missense_variant | COSMIC | — | — | — | ambiguous (0.36) | -6.75 | COSV56637751 |
| 179 | — | frameshift_variant | gnomAD | — | 4.11e-06 | LoF | — | — | chr4-120060260-TTGTC-T |
| 180 | D→H | missense_variant | gnomAD | — | 4.11e-06 | damaging | likely_pathogenic (0.68) | -8.75 | chr4-120060258-C-G |
| 181 | L→F | missense_variant | gnomAD | — | 1.37e-06 | damaging | likely_benign (0.33) | -8.37 | chr4-120060253-C-G |
| 181 | L→S | missense_variant | gnomAD | — | 1.37e-06 | damaging | likely_benign (0.29) | -8.87 | chr4-120060254-A-G |
| 181 | L→M | missense_variant | gnomAD | — | 2.05e-06 | — | likely_benign (0.10) | -7.28 | chr4-120060255-A-T |
| 181 | L→F | missense_variant | COSMIC | — | — | damaging | likely_benign (0.33) | -8.37 | COSV99633667 |
| 181 | L→W | missense_variant | COSMIC | — | — | damaging | likely_pathogenic (0.59) | -11.50 | COSV56636532 |
| 182 | V→V | synonymous_variant | gnomAD | — | 6.84e-07 | — | — | 0.00 | chr4-120060250-A-T |
| 182 | V→L | missense_variant | gnomAD | — | 7.53e-06 | — | likely_benign (0.11) | -5.09 | chr4-120060252-C-G |
| 183 | V→V | synonymous_variant | gnomAD | — | 4.79e-06 | — | — | 0.00 | chr4-120060247-T-C |
| 183 | V→L | missense_variant | gnomAD | — | 1.37e-06 | damaging | likely_pathogenic (0.61) | -8.00 | chr4-120060249-C-A |
| 185 | — | frameshift_variant | COSMIC | — | — | LoF | — | — | COSV56637918 |
| 185 | — | frameshift_variant | COSMIC | — | — | LoF | — | — | COSV56637019 |
| 185 | E→* | stop_gained | COSMIC | — | — | LoF | — | — | COSV99633990 |
| 187 | W→S | missense_variant | gnomAD | — | 6.84e-07 | damaging | likely_pathogenic (1.00) | -12.37 | chr4-120060236-C-G |
| 187 | W→G | missense_variant | COSMIC | — | — | damaging | likely_pathogenic (0.98) | -11.81 | COSV56637715 |
| 188 | E→K | missense_variant | COSMIC | — | — | damaging | likely_pathogenic (0.75) | -9.25 | COSV99633670 |
| 189 | E→K | missense_variant | gnomAD | — | 2.74e-06 | damaging | likely_pathogenic (0.92) | -9.69 | chr4-120060231-C-T |
| 190 | S→S | synonymous_variant | gnomAD | — | 6.85e-07 | — | — | 0.00 | chr4-120060226-C-A |
| 190 | S→S | synonymous_variant | gnomAD | — | 1.57e-05 | — | — | 0.00 | chr4-120060226-C-T |
| 190 | S→L | missense_variant | COSMIC | — | — | damaging | likely_pathogenic (0.69) | -10.25 | COSV104613725 |
| 191 | G→E | missense_variant | gnomAD | — | 6.85e-07 | damaging | likely_pathogenic (0.89) | -9.12 | chr4-120060224-C-T |
| 192 | — | frameshift_variant | gnomAD | — | 6.85e-07 | LoF | — | — | chr4-120060220-TG-T |
| 192 | P→T | missense_variant | gnomAD | — | 9.58e-06 | damaging | likely_pathogenic (0.89) | -10.31 | chr4-120060222-G-T |
| 192 | P→T | missense_variant | COSMIC | — | — | damaging | likely_pathogenic (0.89) | -10.31 | COSV56637625 |
| 194 | F→L | missense_variant | gnomAD | — | 6.85e-07 | damaging | likely_pathogenic (0.77) | -6.34 | chr4-120060216-A-G |
| 195 | I→I | synonymous_variant | gnomAD | — | 6.85e-07 | — | — | 0.00 | chr4-120060211-A-T |
| 195 | I→V | missense_variant | gnomAD | — | 4.11e-06 | — | likely_benign (0.13) | -6.56 | chr4-120060213-T-C |
| 196 | T→T | synonymous_variant | gnomAD | — | 6.85e-07 | — | — | 0.00 | chr4-120060208-G-A |
| 196 | T→T | synonymous_variant | gnomAD | — | 6.85e-07 | — | — | 0.00 | chr4-120060208-G-C |
| 197 | N→N | synonymous_variant | gnomAD | — | 6.85e-07 | — | — | 0.00 | chr4-120060205-A-G |
| 197 | N→S | missense_variant | gnomAD | — | 6.85e-07 | — | likely_benign (0.07) | -4.55 | chr4-120060206-T-C |
| 198 | S→F | missense_variant | gnomAD | — | 3.42e-06 | damaging | likely_pathogenic (0.85) | -10.37 | chr4-120060203-G-A |
| 198 | S→A | missense_variant | gnomAD | — | 1.37e-06 | damaging | likely_benign (0.10) | -9.31 | chr4-120060204-A-C |
| 198 | S→T | missense_variant | gnomAD | — | 6.85e-07 | damaging | likely_benign (0.33) | -10.69 | chr4-120060204-A-T |
| 199 | E→G | missense_variant | gnomAD | — | 1.51e-05 | damaging | likely_pathogenic (0.86) | -10.19 | chr4-120060200-T-C |
| 200 | E→G | missense_variant | gnomAD | — | 6.85e-07 | damaging | likely_pathogenic (0.73) | -9.69 | chr4-120060197-T-C |
| 200 | E→K | missense_variant | gnomAD | — | 6.85e-07 | damaging | likely_pathogenic (0.73) | -9.00 | chr4-120060198-C-T |
| 201 | V→I | missense_variant | gnomAD | — | 6.85e-07 | damaging | ambiguous (0.49) | -9.00 | chr4-120060195-C-T |
| 202 | R→L | missense_variant | gnomAD | — | 6.85e-07 | damaging | likely_pathogenic (0.76) | -8.18 | chr4-120060191-C-A |
| 202 | R→H | missense_variant | gnomAD | — | 7.53e-06 | — | likely_benign (0.32) | -6.15 | chr4-120060191-C-T |
| 202 | R→C | missense_variant | gnomAD | — | 1.64e-05 | — | ambiguous (0.44) | -6.15 | chr4-120060192-G-A |
| 202 | R→L | missense_variant | COSMIC | — | — | damaging | likely_pathogenic (0.76) | -8.18 | COSV99633797 |
| 202 | R→C | missense_variant | COSMIC | — | — | — | ambiguous (0.44) | -6.15 | COSV56637784 |
| 203 | L→F | missense_variant | gnomAD | — | 1.10e-05 | damaging | likely_pathogenic (0.90) | -9.69 | chr4-120060189-G-A |
| 203 | L→R | missense_variant | COSMIC | — | — | damaging | likely_pathogenic (0.99) | -13.06 | COSV99633921 |
| 204 | R→L | missense_variant | gnomAD | — | 1.37e-06 | damaging | likely_pathogenic (0.97) | -10.50 | chr4-120060185-C-A |
| 204 | R→H | missense_variant | gnomAD | — | 7.53e-06 | damaging | likely_pathogenic (0.71) | -8.25 | chr4-120060185-C-T |
| 204 | R→C | missense_variant | gnomAD | — | 2.74e-06 | damaging | likely_pathogenic (0.86) | -8.19 | chr4-120060186-G-A |
| 204 | R→H | missense_variant | ClinVar | Uncertain significance | — | damaging | likely_pathogenic (0.71) | -8.25 | ClinVar:3541917 |
| 204 | R→C | missense_variant | COSMIC | — | — | damaging | likely_pathogenic (0.86) | -8.19 | COSV56637566 |
| 205 | S→S | synonymous_variant | gnomAD | — | 9.59e-06 | — | — | 0.00 | chr4-120060181-T-C |
| 205 | S→L | missense_variant | gnomAD | — | 6.85e-07 | damaging | likely_pathogenic (0.86) | -12.12 | chr4-120060182-G-A |
| 205 | S→* | stop_gained | COSMIC | — | — | LoF | — | — | COSV99633830 |
| 206 | F→F | synonymous_variant | gnomAD | — | 6.85e-07 | — | — | 0.00 | chr4-120060178-A-G |
| 206 | F→L | missense_variant | gnomAD | — | 6.85e-07 | damaging | likely_pathogenic (1.00) | -9.62 | chr4-120060180-A-G |
| 209 | T→I | missense_variant | gnomAD | — | 6.85e-07 | damaging | likely_pathogenic (0.89) | -7.96 | chr4-120060170-G-A |
| 209 | T→I | missense_variant | COSMIC | — | — | damaging | likely_pathogenic (0.89) | -7.96 | COSV106096082 |
| 210 | I→V | missense_variant | gnomAD | — | 5.62e-03 | — | likely_benign (0.10) | -5.18 | chr4-120060168-T-C |
| 210 | I→V | missense_variant | ClinVar | Benign | — | — | likely_benign (0.10) | -5.18 | ClinVar:782943 |
| 211 | H→H | synonymous_variant | gnomAD | — | 6.86e-07 | — | — | 0.00 | chr4-120060163-G-A |
| 211 | H→R | missense_variant | COSMIC | — | — | damaging | likely_pathogenic (0.99) | -11.31 | COSV56637130 |
| 213 | V→V | synonymous_variant | gnomAD | — | 6.86e-07 | — | — | 0.00 | chr4-120060157-T-C |
| 214 | N→N | synonymous_variant | gnomAD | — | 1.37e-06 | — | — | 0.00 | chr4-120060154-A-G |
| 215 | S→N | missense_variant | gnomAD | — | 1.37e-06 | damaging | likely_pathogenic (0.93) | -10.75 | chr4-120060152-C-T |
| 215 | S→G | missense_variant | COSMIC | — | — | damaging | likely_benign (0.28) | -9.44 | COSV56636733 |
| 216 | M→K | missense_variant | gnomAD | — | 7.56e-06 | damaging | ambiguous (0.44) | -9.44 | chr4-120060149-A-T |
| 217 | V→M | missense_variant | COSMIC | — | — | damaging | likely_pathogenic (1.00) | -11.62 | COSV56636857 |
| 218 | A→A | synonymous_variant | gnomAD | — | 2.07e-06 | — | — | 0.00 | chr4-120060142-G-A |
| 218 | A→A | synonymous_variant | gnomAD | — | 6.91e-07 | — | — | 0.00 | chr4-120060142-G-C |
| 218 | A→V | missense_variant | gnomAD | — | 6.91e-07 | damaging | likely_pathogenic (0.78) | -8.25 | chr4-120060143-G-A |
| 218 | A→T | missense_variant | gnomAD | — | 2.07e-06 | — | ambiguous (0.42) | -6.56 | chr4-120060144-C-T |
| 219 | Y→Y | synonymous_variant | gnomAD | — | 6.90e-07 | — | — | 0.00 | chr4-120060139-G-A |
| 220 | K→R | missense_variant | gnomAD | — | 2.07e-06 | — | likely_benign (0.06) | -5.31 | chr4-120060137-T-C |
| 220 | K→E | missense_variant | gnomAD | — | 6.90e-07 | damaging | likely_pathogenic (0.69) | -11.37 | chr4-120060138-T-C |
| 221 | I→M | missense_variant | gnomAD | — | 2.07e-06 | — | likely_benign (0.09) | -6.84 | chr4-120060133-A-C |
| 221 | I→I | synonymous_variant | gnomAD | — | 6.90e-07 | — | — | 0.00 | chr4-120060133-A-G |
| 222 | P→S | missense_variant | COSMIC | — | — | — | likely_benign (0.13) | -5.24 | COSV56636813 |
| 223 | V→V | synonymous_variant | gnomAD | — | 6.91e-07 | — | — | 0.00 | chr4-120060127-G-A |
| 223 | V→L | missense_variant | gnomAD | — | 4.15e-06 | — | likely_benign (0.09) | -5.62 | chr4-120060129-C-G |
| 223 | V→I | missense_variant | gnomAD | — | 4.84e-06 | — | likely_benign (0.07) | -4.18 | chr4-120060129-C-T |
| 223 | V→I | missense_variant | COSMIC | — | — | — | likely_benign (0.07) | -4.18 | COSV56636962 |
| 224 | N→S | missense_variant | gnomAD | — | 1.38e-06 | — | likely_benign (0.06) | -3.94 | chr4-120060125-T-C |
| 225 | D→D | synonymous_variant | gnomAD | — | 1.38e-05 | — | — | 0.00 | chr4-120060121-G-A |
| 225 | D→Y | missense_variant | gnomAD | — | 1.38e-06 | damaging | likely_benign (0.28) | -7.90 | chr4-120060123-C-A |
| 225 | D→N | missense_variant | gnomAD | — | 2.77e-06 | — | likely_benign (0.08) | -6.37 | chr4-120060123-C-T |
| 225 | D→N | missense_variant | ClinVar | Uncertain significance | — | — | likely_benign (0.08) | -6.37 | ClinVar:2539636 |
390 variants in the shared canonical core.
LoF = frameshift / stop-gain / splice-disrupting — inherently loss-of-function, flagged by consequence (AlphaMissense and ESM-C score only missense/substitutions, so they are blank here by design, not by absence of impact). AlphaMissense is computed in the canonical reading frame, so it scores the shared core but reads N/A across an isoform-unique extension — use ESM-C ΔLLR there.