UBE2M
EXTENDED 248 aa (canonical 183 aa) · UniProt P61081 · CDLMPS
chr19:58558576:-:ACG:ENST00000253023.8
AI summary N-terminal extension adds a disordered, Gly/Ala-rich tail with one loosely-docked helix, but no domain, localization, or biophysical change to the UBE2M core.
The added 65-residue N-terminus folds moderately well in isolation (mean pLDDT 0.79) and contains one qualifying helix, but that helix shows weak, diffuse contacts and high PAE (~12-19 Å) against the shared catalytic core, meaning its orientation relative to the retained E2 domain is essentially unresolved rather than a stably integrated new structural module. DeepLoc, SignalP/TargetP, InterPro domains, and whole-protein biophysics all read unchanged versus canonical, so no mechanism finding here contradicts or extends the known nuclear/cytosolic, cullin-neddylating function.
UBE2M's activity depends on its UBC catalytic core (charged by the NEDD8 E1, engaged by DCN1 via a short N-terminal peptide, catalyzing via C111) and on nuclear/cytosolic access to cullin-RING ligases. Because the shared core is structurally unperturbed (low RMSD, high shared pLDDT/pTM) and no domain, compartment, or biophysical shift accompanies the extension, this isoform's added tail does not appear to alter catalytic core integrity, subcellular distribution, or the DCN1-binding surface as currently modeled; its loosely-docked helix is a structural curiosity without demonstrated functional consequence for cullin neddylation or the described non-cullin substrate interactions.
Extension is real (multi-cell-line detection, MS-validated peptides, strong conservation/phyloP) but germline-constraint and disease-variant metrics are uninterpretable/depleted for this never-coding-turned-coding region, and the one candidate structural feature (helix) has high inter-region PAE, so no tier-2 mechanism passes the confidence bar for tagging.
Folding
Evidence — click any tile for the differential-region detail
LLM reasoning
How well is the isoform-unique region conserved at the amino-acid level across primates (mean percent identity)? High AA identity in close relatives argues the alternative protein is real and translated, not a sequencing or annotation artifact. (Reading-frame intactness is shown as context.)
Comparison
Sequence conservation · primate
mean AA identity over the differential ORF is the score basis; frame intactness across the clade is context
| Conservation metric | Canonical ORF | Differential ORF |
|---|---|---|
| Mean AA identity | 100% | 96% |
| Frame intact (fraction of species) | 84% | 25% |
| Species aligned | 25 | 24 |
| Species frame-intact | 21 | 6 |
| Start codon conserved | 96% | 91% |
| Deepest intact species | Microcebus_murinus | Callithrix_jacchus |
| Phylo depth (MRCA) | 7 | 6 |
The same amino-acid identity test across mammals. Conservation over deeper evolutionary distance is stronger evidence the ORF is under selection to be translated. (Reading-frame intactness is shown as context.)
Comparison
Sequence conservation · mammalian
mean AA identity over the differential ORF is the score basis; frame intactness across the clade is context
| Conservation metric | Canonical ORF | Differential ORF |
|---|---|---|
| Mean AA identity | 99% | 85% |
| Frame intact (fraction of species) | 65% | 17% |
| Species aligned | 20 | 18 |
| Species frame-intact | 13 | 3 |
| Start codon conserved | 100% | 92% |
| Deepest intact species | Echinops_telfairi | Bos_taurus |
| Phylo depth (MRCA) | 12 | 11 |
phyloP / phastCons measure per-base evolutionary constraint. A high absolute phyloP over the differential region means the sequence itself is under strong purifying selection — evidence it is coding and does something. (Enrichment over the shared core is shown as context only.)
Comparison
Per-base conservation · differential region
absolute mean phyloP over the differential region is the score basis (high = strong purifying selection); the shared column and enrichment are context, not the claim
| Track | Differential | Shared | vs shared |
|---|---|---|---|
| phyloP mean | 4.01 | 5.63 | 0.713 |
| phastCons mean | 0.322 | 0.897 | — |
Start site · canonical vs isoform
initiation context + per-base conservation at each start codon
| Property | Canonical | Isoform |
|---|---|---|
| Start codon | ATG | ACG |
| Kozak context (−9..+4) | GGCGGCAGGATGA | AGCCGGACTACGG |
| phyloP at start codon | 6.16 | 6.44 |
| phastCons at start codon | 1 | 0.208 |
| phyloP over Kozak window | 6.63 | 4.56 |
| phastCons over Kozak window | 1 | 0.0758 |
| Kozak mismatch — full consensus | 5 | 7 |
| Kozak window GC content | 0.692 | 0.692 |
LLM reasoning
Is the alternative start used in more than one cell line? Reproducible initiation across independent samples argues against a one-off ribosome artifact.
Comparison
Per-cell-line usage · canonical vs isoform
Fisher q = 2.61e-21
| Cell line | Canonical | Isoform | p-value |
|---|---|---|---|
| HeLa | — | 10.6 | 0.000142 |
| K562 | 7.03 | 4.94 | 2.61e-21 |
| U2OS | 0.81 | 1.18 | 1.37e-13 |
| RPE1 Async | 1.2 | 1.09 | 6.1e-05 |
| RPE1 Que | 0.811 | 0.316 | 1.2e-13 |
| RPE1 Sen | 1.13 | 1.09 | 5.15e-06 |
How efficiently ribosomes initiate at this start (TIS) relative to background, per cell line. Strong, reproducible initiation supports a genuine translation event.
Comparison
Start-Site Usage · canonical vs isoform
ribosome initiation efficiency at the canonical start vs this alternative start, per cell line
| Cell line | Canonical | Isoform |
|---|---|---|
| HeLa | — | 0.119 |
| K562 | 0.152 | 0.107 |
| U2OS | 0.0303 | 0.0441 |
| RPE1 Async | 0.0233 | 0.0211 |
| RPE1 Que | 0.0281 | 0.0109 |
| RPE1 Sen | 0.0204 | 0.0196 |
Were tryptic peptides unique to the differential region detected by mass spec? Direct peptide evidence is the strongest proof the isoform protein exists.
Comparison
Peptide Evidence (canonical vs isoform)
isoform-unique peptides are direct evidence the alternative protein exists
| Feature | Canonical | Isoform |
|---|---|---|
| Tryptic peptides (in-silico) | 29 | 14 |
| Validated by mass-spec | 0 | 5 |
| Isoform-unique peptides | — | 14 |
Details
Peptide Evidence (canonical vs isoform)
- peptide MGAEAGR 0–7
- peptide GAEAGR 1–7
- validated VAAAAEEAAAAGPR 22–36
- peptide SGGDAGGGGGPGGR 36–50
- peptide GGGSGGGGGR 55–65
- peptide MGAEAGRAVGAER 0–13
- peptide GAEAGRAVGAER 1–13
- validated AVGAERSGAAR 7–18
- validated SGAARQAGR 13–22
- peptide QAGRVAAAAEEAAAAGPR 18–36
- peptide VAAAAEEAAAAGPRSGGDAGGGGGPGGR 22–50
- validated SGGDAGGGGGPGGRGPGPR 36–55
- validated GPGPRGGGSGGGGGR 50–65
- peptide GGGSGGGGGRMIK 55–68
LLM reasoning
Do the predicted localization features (DeepLoc prediction, sorting signals, or membrane association) differ between the canonical and the isoform? A protein with changed localization features acts in a different cellular context.
Comparison
Subcellular localization (DeepLoc)
no localization-feature change
| Property | Canonical | Isoform |
|---|---|---|
| Predicted location | Cytoplasm | Cytoplasm |
| Sorting signals | Nuclear localization signal | Nuclear localization signal |
| Membrane | Peripheral|Soluble | Peripheral|Soluble |
Do N-terminal targeting signals (SignalP secretion, TargetP mitochondrial/chloroplast) differ between canonical and isoform? N-terminal changes most directly add or remove targeting peptides.
LLM reasoning
Does healthy human germline variation (gnomAD) avoid this region (depletion ratio < 1×), and is it intrinsically constrained (ESM-C constraint delta > 0)? Depletion of population variation plus high sequence constraint mean the region resists change — it is functionally important. Scored only where the unique region is canonical coding sequence, i.e. on truncations. gnomAD is a tolerance catalogue, not a disease one; disease/cancer variants (ClinVar / COSMIC) live in M2.
Comparison
Germline Variants · differential vs shared region
gnomAD (population) variant density per nucleotide — depletion ratio < 1× means healthy human variation avoids the region (constrained), the M1 basis alongside ESM-C constraint
| Variant set | Differential | Shared | Depletion ratio |
|---|---|---|---|
| gnomAD variants | 91 | 215 | 1.2× |
Sequence constraint (ESM-C) · differential vs shared region
lower LLR = more constrained; enrichment = differential / conserved
| Property | Differential | Shared | Enrichment |
|---|---|---|---|
| ESM-C mean LLR | -1.69 | -0.0159 | — |
| Constrained positions | 0 | 0 | — |
Predicted-damaging variants · germline (gnomAD)
AlphaMissense / ESM-C / LoF calls per region (length-normalized)
| Variant set | Differential | Shared | Enrichment |
|---|---|---|---|
| Scorable variants | 71 | 209 | 0.96× |
| Damaging variants | 4 | 54 | 0.21× |
| — of which loss-of-function | 4 | 2 | 5.6× |
| AlphaMissense-pathogenic | 0 | 30 | 0× |
Predictor scores · germline (gnomAD)
scored: 377 ESM-C · 186 AlphaMissense
| Property | Differential | Shared |
|---|---|---|
| Mean ΔLLR (ESM-C) | -1.45 | -3.64 |
| Min ΔLLR (ESM-C) | -5.89 | -13.9 |
| Mean AlphaMissense | — | 0.419 |
Are disease (ClinVar / COSMIC) variants enriched per nucleotide in the differential region versus the shared core (disease enrichment ratio ≥ 1×)? Disease variants concentrating in the unique region tie it to phenotype.
Comparison
Clinical-variant burden · differential vs shared region
counts per region; ratio is density-normalized (variants per nucleotide, differential ÷ shared — ≥1× = disease variants concentrate in the differential region, the M2 basis)
| Variant set | Differential | Shared | Enrichment |
|---|---|---|---|
| Disease variants | 5 | 113 | 0.12× |
| Pathogenic | 0 | 0 | — |
Predicted-damaging variants · disease (ClinVar/COSMIC)
AlphaMissense / ESM-C / LoF calls per region (length-normalized)
| Variant set | Differential | Shared | Enrichment |
|---|---|---|---|
| Scorable variants | 5 | 111 | 0.13× |
| Damaging variants | 0 | 69 | 0× |
| — of which loss-of-function | 0 | 13 | 0× |
| AlphaMissense-pathogenic | 0 | 47 | 0× |
Predictor scores · disease (ClinVar/COSMIC)
scored: 377 ESM-C · 186 AlphaMissense
| Property | Differential | Shared |
|---|---|---|
| Mean ΔLLR (ESM-C) | -0.717 | -7.34 |
| Min ΔLLR (ESM-C) | -2.95 | -14.4 |
| Mean AlphaMissense | — | 0.599 |
LLM reasoning
Does the gained/lost region fold into ordered structure (pLDDT) and shift the protein's biophysical character (charge, hydropathy, disorder)? Structured, biophysically distinct regions are more likely to be functional.
Comparison
Structure (ESMFold2) · canonical vs isoform
TM-score 0.827 · RMSD 1.44 Å · 9 interface contacts
| Metric | Canonical | Isoform |
|---|---|---|
| pLDDT (whole protein) | 0.869 | 0.873 |
Fold confidence · differential vs shared region
mean ESMFold2 pLDDT in each region
| Metric | Differential | Shared | Enrichment |
|---|---|---|---|
| pLDDT | 0.787 | 0.751 | 1 |
The shared region is the stretch of protein identical in both the isoform and the canonical (the canonical body for an extension; the post-truncation body for a truncation). Folded in both contexts it normally comes out nearly identical, so its Cα backbone RMSD ≈ 0. A high shared-region RMSD (superposed on the shared residues only, from the ESMFold2 structures) means the extension or truncation reorganizes how the retained region folds — a rare, high-interest functional signal. TM-score is a length-normalized companion; the check only fires when both structures are confidently folded, and uORF/altORF isoforms (no shared region) are not evaluable.
Comparison
Shared-region structural change
Cα RMSD superposed on the shared residues only; TM-score is length-normalized · shared-region Cα RMSD 1.48 Å · shared TM-score 0.947 · shared region 183 aa · min shared pLDDT 0.869 · global TM-score 0.827 · global RMSD 1.44 Å
| Property | Canonical | Isoform |
|---|---|---|
| Shared-region pLDDT | 0.869 | 0.904 |
Does the differential region contain actual secondary structure — a helix or strand — rather than coil? ESMFold2 predicts coordinates but no secondary structure, so elements are assigned from those coordinates (P-SEA, Cα geometry). Because they are derived from a prediction, each element carries its own mean pLDDT: a geometrically clean helix running through a disordered stretch is geometry fitted to a guess, so BOTH length and confidence are required to score. The direction differs by ORF type — an extension GAINS the element (a candidate functional addition), a truncation LOSES one, and there the element is read off the canonical structure because the removed segment exists only in it. This says nothing about whether the element is integrated with the rest of the fold; the contact and PAE evidence in this same category answers that.
Comparison
Secondary structure · differential vs shared region
helices and strands assigned from the predicted coordinates (P-SEA)
| Metric | Differential | Shared |
|---|---|---|
| Alpha helices | 1 | 4 |
| Beta strands | 0 | 2 |
| Longest element (aa) | 17 | 13 |
| Mean pLDDT | 0.75 | 0.96 |
Elements and coordinates
1 in the differential region, 6 in the shared core — residue numbering is 1-based on the protein holding the region
| Isoform-unique | Shared core |
|---|---|
| alpha helix 16–32 17 aa · pLDDT 0.75 | alpha helix 94–105 12 aa · pLDDT 0.92 |
| — | beta strand 125–131 7 aa · pLDDT 0.98 |
| — | beta strand 139–147 9 aa · pLDDT 0.98 |
| — | alpha helix 190–202 13 aa · pLDDT 0.91 |
| — | alpha helix 212–219 8 aa · pLDDT 0.96 |
| — | alpha helix 222–234 13 aa · pLDDT 0.97 |
Below threshold
3 in the differential region, 2 in the shared core — shorter than 6 aa or below pLDDT 0.70, so not counted above
| Isoform-unique | Shared core |
|---|---|
| beta strand 2–9 8 aa · pLDDT 0.70 | alpha helix 72–78 7 aa · pLDDT 0.56 |
| beta strand 38–42 5 aa · pLDDT 0.88 | beta strand 112–116 5 aa · pLDDT 0.93 |
| beta strand 57–61 5 aa · pLDDT 0.85 | — |
LLM reasoning
Does the isoform gain or lose a real InterPro functional domain in the differential region (disorder/structural-only signatures excluded)? Gaining or losing a domain changes function directly.
Comparison
Domains & motifs (canonical vs isoform)
gained = only in the isoform; lost = only in the canonical
| Feature | Canonical | Isoform |
|---|---|---|
| InterPro domains | 10 | 10 |
| Short linear motifs | 2 | 2 |
Biophysical character of the isoform-differential region versus the shared canonical core — pI, hydropathy, charge, disorder and related properties. Descriptive; the folding (P1) score keys off the GRAVY / charge / disorder deltas.
Comparison
Biophysics · differential vs shared region
highlighted rows are enriched in the differential region
| Property | Differential | Shared | Enrichment |
|---|---|---|---|
| Isoelectric point (pI) | 12 | 7.59 | 1.58 |
| Hydropathy (GRAVY) | -0.565 | -0.517 | 1.09 |
| Fraction charged | 0.2 | 0.279 | 0.718 |
| Disorder fraction | 0.221 | 0.108 | 2.04 |
| Disorder-promoting | 0.954 | 0.541 | 1.76 |
| Low-complexity fraction | 0.969 | 0 | — |
| Prion-like fraction | 0.446 | 0.284 | 1.57 |
| LLPS score | 0.386 | 0.159 | 2.43 |
| π–π propensity | 0.139 | 0.262 | 0.528 |
| Aromaticity | 0 | 0.0984 | 0 |
| Instability index | 44 | 40.5 | 1.09 |
| Shannon entropy | 2.53 | 4.11 | 0.616 |
| Normalized complexity | 0.586 | 0.952 | 0.616 |
Sparse-autoencoder (SAE) interpretability features on the ESM-C residual stream, comparing the isoform against the canonical protein. This is the S3 criterion — a presence check that fires when any interpretable feature is gained or lost. Only the top 30 features per category (ranked by activation, prevalence ≥ 2) are saved; each table shows the top 10 interpretable features, and generic / "unknown" features are hidden unless you expand it. What are SAE features? →
Part 1 · Global feature comparison
Which SAE features fire in the isoform vs the canonical protein, across the whole sequence.
Isoform-only features — 158 total (prevalence ≥ 2)
| Feature | Label | Description | Activation | Prevalence |
|---|---|---|---|---|
| #448 | N-terminal leader/targeting segments | N-terminal leader/targeting segments: the feature activates on the extreme N-terminus (first 2–5 residues) of proteins, especially the N-region of signal peptides and organelle transit peptides, and more generally on short, disordered/low-complexity N-terminal tails; strongest at residue ~3 and largely independent of amino-acid identity, occurring across all taxa and functions. | 13.58 | 2 |
| #15545 | Low-complexity disordered regions | Intrinsic disorder/low-complexity segments enriched in small, polar and charged residues (S/T/N/G/A/E/D/R/K), especially N-terminal tails, C-terminal tails, and inter-domain regions that contain short repeating motifs; common in nucleic-acid-associated regulators and many viral proteins, and largely absent from structured domains | 11.84 | 16 |
| #3609 | Position-2 N-terminus residue detector | Detector of the identity of the second residue at the extreme N-terminus of proteins, with strongest preference for small residues (especially Ser, Ala, and Gly; also Thr/Pro), i.e., an N-terminal position-2 signal commonly associated with generic, disordered starts and co-translational N-terminus processing | 10.99 | 2 |
| #7996 | Generic extreme N-terminus detector | Generic extreme N-terminus detector: marks residues around position 5 (occasionally 6–7) at the very start of the polypeptide, most often within the N-region of signal/transit peptides of secreted or organelle-targeted precursors, but also in generally disordered N-termini of cytosolic/nuclear proteins; not motif- or residue-specific | 10.54 | 2 |
| #7903 | Arginine-rich polybasic patches | Basic polycationic patches enriched in arginine (often with lysine) within low-complexity/disordered regions, frequently N-terminal. These include classical monopartite NLSs, nucleic acid–binding basic tails, signal peptide n-regions and cytosolic juxtamembrane segments (positive-inside rule), and basic clusters in secreted precursors; typically depleted of acidic residues. | 8.23 | 8 |
| #678 | Sparse terminal residue peaks | Sparse residue-level activations distributed across short proteins and protein segments, with frequent firing in disordered N-terminal/C-terminal tails as well as in some short transmembrane or compact domains | 7.66 | 16 |
| #14493 | Short hydrophobic helical runs | Short hydrophobic helical segments rich in I/V/F (sometimes M), recognized in both membrane-targeting/insertion contexts (signal peptides, signal-anchors, transmembrane helices) and in hydrophobic helical patches within otherwise soluble small proteins. The feature flags contiguous aliphatic/hydrophobic runs across diverse proteins, including viral membrane/accessory proteins, secreted peptide/effector precursors, small ORFs/microproteins, and short basic/uncharacterized human proteins. | 7.42 | 3 |
| #11429 | G/S-rich disordered low-complexity | Intrinsically disordered, low‑complexity regions enriched in glycine and serine (with frequent threonine and Q/N tracts)—i.e., SG‑repeat and polar low‑complexity segments—common in eukaryotic regulators and some viral proteins, and largely absent from folded catalytic cores. | 7.32 | 29 |
| #11159 | Glycine residue identity detector | Residue-identity detector for glycine (G), largely context-agnostic but with a tendency to fire on glycines in flexible/disordered, often N-terminal or membrane-proximal regions; also active on glycines within transmembrane helices; broadly taxon- and function-independent. | 7.19 | 25 |
| #11504 | Initiator methionine at M1 | Initiator methionine at the very start of the polypeptide chain (M1), i.e., the translation start residue, independent of protein family, taxonomy, membrane association, or presence of signal/leader/propeptide regions; often annotated as post-translationally removed and typically situated in a flexible, coil-like N-terminus. | 6.97 | 2 |
| #3642 | Gly/Ala/Pro-rich disordered segments | Glycine/alanine/proline-rich low-complexity stretches, frequently within annotated disordered regions, that share a recurring small-residue context (Gly-Val/Ala-Xaa-Ala/Pro patterns). | 6.91 | 4 |
| #3454 | Disordered C-terminal G/P/S repeats | Short G/P/S-rich low-complexity segments within intrinsically disordered C-terminal tails of bacterial proteins, frequently containing repeated GxG or SGIG/GPG-like sub-motifs with sparsely interspersed R residues | 6.75 | 21 |
| #8488 | Ala/Thr-rich N-terminal disorder | Ala/Thr-enriched composition feature with a bias for low-complexity intrinsically disordered regions and N-terminal prepro/signal-peptide segments; the feature reflects small-residue (A/T, secondarily S/P) composition often in flexible regions but also appears at A/T residues in some structured contexts. | 6.66 | 16 |
| #11915 | Disordered small-residue tandem repeats | Low-complexity, intrinsically disordered tandem-repeat tracts enriched in small/polar residues (Ser/Thr/Ala/Gly/Pro, often acidic), highlighting simple S/T/A/G/P-rich repeats and flexible non-catalytic tails rather than folded domains; basic and aromatic residues are disfavored. | 6.64 | 25 |
| #9946 | Disordered N-termini and coils | Detector for intrinsically disordered, low-structure N‑terminal pre-sequences (signal peptides’ N/C regions, organellar transit peptides, and propeptides) and, more generally, flexible coil/low‑pLDDT segments; strongest bias for the first 10–70 residues but can also mark internal loops in large enzymes and short low‑complexity micropeptides. | 6.62 | 8 |
| #9852 | N-terminus activation; hydrophobic helices secondary | Extreme N-terminal regions of proteins, starting at the initiator methionine and extending over a short stretch, with occasional secondary activation at internal hydrophobic/amphipathic helices. | 6.49 | 61 |
| #9234 | Charged low-complexity disordered regions | Intrinsically disordered, low-complexity segments enriched for Lys and acidic residues (E/D), plus N/Q, that include both polyampholyte and basic-residue tracts; typically flexible terminal tails and linker regions in diverse eukaryotic and viral proteins, including many small/secreted proteins and membrane-protein loops | 6.48 | 7 |
| #13345 | Low-complexity glycine-rich repeats | Small-residue–biased low-complexity repeat regions—typically intrinsically disordered and enriched for Gly and/or Ser/Thr with frequent acidic (Asp/Glu) residues—serving as flexible linkers or mesh-forming domains (e.g., FG repeats in nucleoporins, PTS- or Ser/Thr-rich repeats in mucins/adhesins, glycine-biased linkers of secreted/cell-wall proteins); also includes N/Q-rich LCRs in some taxa (e.g., Apicomplexa) and occasionally captures glycine-periodic repeats such as Gly–X–Y in collagen-like sequences. | 5.92 | 27 |
| #5337 | Pro/RGG-rich low-complexity IDRs | Intrinsically disordered, low‑complexity basic segments at termini and long loops, enriched in Pro/Gly and/or Arg/Ser (including RG/RGG or S/T clusters), common across eukaryotic and viral proteins. | 5.42 | 50 |
| #2548 | Sticker-spacer tandem-repeat IDRs | Intrinsically disordered, low-complexity tandem-repeat regions (“sticky–spacer” IDRs) enriched in small aliphatic/proline residues with polar/acidic content (P/A/G/V with S/T and D/E), common in secreted/extracellular scaffolds and terminal tails of cytosolic proteins; the feature especially highlights hydrophobic “sticker” residues (often Val/Ala/Phe) embedded within repeat-rich, compositionally biased segments. | 5.34 | 20 |
| #13063 | Serine-rich region detector | Serine residues across a wide range of structural contexts, with a bias toward serine-rich segments in intrinsically disordered or low-complexity regions (often N-terminal tails), but also firing on serines within transmembrane helices and other structured contexts; the feature behaves primarily as a serine detector with some preference for S-rich patches. | 5.26 | 3 |
| #11226 | N/Q-rich low-complexity IDRs | Intrinsically disordered low-complexity regions specifically enriched for long asparagine and/or glutamine/serine/threonine runs; Pro/Gly content alone is insufficient. Activation is typically in extended linker or terminal tracts outside folded domains and is most common in organisms/proteins that harbor pronounced N/Q-rich repeats. | 5.14 | 40 |
| #5482 | N-terminal targeting presequences | Universal eukaryotic N-terminal targeting presequences: the feature detects short, cleavable leader regions at the extreme N-terminus that direct proteins to organelles or the secretory pathway—especially chloroplast/apicoplast transit peptides and thylakoid lumen signals, but also mitochondrial targeting peptides and classical signal peptides. These segments are Ser/Thr- and small/hydrophobic–rich, enriched in Lys/Arg and depleted of acidic residues, typically low-structure/low-confidence and ending at the maturation cleavage site. | 4.71 | 8 |
| #9962 | S/T/P-rich disordered regulatory tails | Intrinsically disordered, low‑complexity regions enriched in serine, threonine, proline and polar/charged residues—flexible regulatory linkers/tails and propeptide segments that often host short linear motifs (e.g., phosphorylation- and proline‑rich motifs) and proteolytic processing sites; signal is absent from well‑folded catalytic domains and is common across taxa | 4.71 | 11 |
| #10117 | Diffuse low-complexity/disorder signature | A broadly tuned signature of disordered/low-complexity-containing proteins: the feature activates diffusely across long stretches of large multi-domain proteins, with residue-level peaks scattered through both folded and disordered segments. It is common in large, repeat-rich extracellular/surface proteins but also marks regions in diverse intracellular proteins. | 4.69 | 36 |
| #10810 | Polar/charged low-complexity regions | Compositionally biased, intrinsically disordered low-complexity segments enriched for polar/charged residues (Q/N/S/T and K/E/D), including homopolymeric tracts (polyQ/polyN), QA/N/S/T repeats, and Lys/acidic clusters; observed across both eukaryotic and bacterial proteins. | 4.66 | 41 |
| #5323 | Acidic/polar IDR hotspots | Residue-level detector for acidic/polar hotspots within intrinsically disordered regions (often N‑terminal), with a strong preference for Asp, Asn, His, and Tyr; highlights acidic low‑complexity stretches and flexible loops rather than structured domains | 4.57 | 2 |
| #3278 | PRQH-rich disordered tails | Compositionally biased, intrinsically disordered low‑complexity segments enriched in Pro/Arg/Gln/His (frequent PR/PQ tracts, Arg- and His-clusters), with occasional sensitivity to Leu‑rich helical stretches (signal peptides or leucine zippers); typically terminal, widespread across taxa, and common in nucleic‑acid–binding and secreted proteins. | 4.54 | 7 |
| #15970 | Serine/threonine-rich disordered linkers | Intrinsically disordered, low‑complexity serine/threonine–rich segments (often containing SP/TP/RS repeats) that function as flexible linkers and phosphorylation‑prone regulatory tracts across diverse proteins, including viral phosphoproteins/nucleocapsid linkers and proline‑rich secreted cell‑wall proteins | 4.52 | 4 |
| #7553 | Eukaryotic N-terminal disordered regulatory tails | Eukaryotic N-terminal intrinsically disordered, low-complexity segments of diverse composition (often including S/P/T/G runs, acidic/basic stretches, or proline clusters) that act as regulatory tails/linkers for phosphorylation-dependent signaling and protein–protein interactions. | 4.29 | 44 |
Canonical-only features — 36 total (prevalence ≥ 2)
| Feature | Label | Description | Activation | Prevalence |
|---|---|---|---|---|
| #11775 | Disordered C-terminal tail hotspots | Short hotspots in intrinsically disordered terminal tails—most often the cytosolic C-termini of membrane proteins (especially GPCRs and single-pass adhesion proteins), but also analogous termini of soluble proteins. | 5.52 | 2 |
| #13847 | Generic polar coil boundary marker | Sparse residue-level feature firing at isolated polar/structured-context positions, often Ser/Thr/Asn/Asp but also basic residues (Arg/Lys), typically in coil or boundary segments of moderately well-structured regions; appears to mark a generic local structural context rather than a function-specific site. | 4.12 | 2 |
| #5613 | Junction loops flanking ligand pockets | Short loop/turn segments at secondary-structure junctions that flank or line ligand/cofactor/metal-binding pockets across diverse enzymes, with a characteristic peak on a conserved His/Tyr within an E-T-H/Y-G–type motif in SAM-binding methyltransferases. | 3.31 | 4 |
| #8387 | Alpha-helix boundary recognition | Alpha-helix boundary recognition: residues at helix termini (especially C‑caps) and immediately adjacent loop/cap positions, i.e., helix–loop junctions, rather than catalytic or ligand-binding sites | 3.19 | 2 |
| #11248 | Short beta-strand loop clusters | Short β-strand/loop segments in non-Ig β-sheet-containing enzymes, including bacterial glycoside hydrolases (amylopullulanases, neopullulanase, isomaltosyltransferase) and eukaryotic molybdo-flavoprotein nitrate reductases. | 2.51 | 2 |
| #11643 | Aromatic/cysteine interfacial anchors | Short, aromatic- and cysteine-enriched interfacial helices/patches that mediate contacts at lipid membranes or protein/protein/nucleic-acid interfaces; typified by Trp/Phe clusters and/or palmitoylation‑prone Cys pairs/clusters; these segments act as interfacial anchors or contact elements rather than core transmembrane helices | 2.25 | 2 |
| #10752 | Short solvent-exposed connector loops | Generic structural signal for short, solvent-exposed loop/turn connectors between secondary structure elements—especially beta-beta hairpin loops and helix-strand junctions—enriched in small/charged residues (S/T, N, D/E, K/R); often adjacent to functional sites but not family-specific | 2.08 | 2 |
| #8694 | Acidic S/T-rich IDRs | Acidic, serine/threonine-rich low-complexity intrinsically disordered regions (often with S/T-P motifs) in eukaryotic nuclear/chromatin-associated proteins, typically residing in long N- or C-terminal regulatory tails enriched for phosphorylation sites. | 2.04 | 2 |
| #9952 | Ordered helical assembly scaffold | A broad “ordered helical/assembly scaffold” signature: long, structured segments—often alpha-helical (TM helices, coiled-coils, helical hairpins) and stable cores of ligand/effector-binding domains—that mark architectural elements used for signal transfer and assembly across diverse molecular machines (signaling, secretion, transport, nucleotide/quinone metabolism, and viral replication), rather than specific catalytic micro-motifs. | 2.02 | 4 |
| #6235 | Aromatic pocket-gating loop SLiM | A short, aromatic/hydrophobic, helix- or strand-capping loop or linear motif that either lines the entrance/gate of small-molecule/cofactor pockets or forms an exposed docking patch for protein–protein interactions. It recurs across enzymes and helix-grip lipid/ligand carriers (GNAT/NAT acetyltransferases, LuxI AHL synthases, HPPK, Bet v 1/START-like), and also on β-grasp ubiquitin-like modifiers and certain IDR-embedded degrons (e.g., Aux/IAA). Segments are enriched in hydrophobic/aromatic residues with interspersed charges and often feature an aromatic–acidic dyad (W next to D/E) or short W/Y–P–P motifs; when annotated, they map to binding or docking residues. | 1.97 | 2 |
| #630 | Cofactor binding loop | A short loop/coil segment lining enzyme active/cofactor-binding pockets, commonly near nucleotide- or cofactor-binding residues across glycosyltransferases, nucleotidyltransferases, and other nucleotide/cofactor-processing enzymes. | 1.93 | 2 |
| #1226 | Short basic/aromatic interaction patches | Short basic/aromatic–enriched segments, typically at or just beyond catalytic domains and especially the retroelement integrase C‑terminal module (integrase CTD/GPY‑F region), that often occur as flexible coils or short helices and serve as DNA/chromatin- or partner‑interaction patches; the feature also recognizes analogous Trp‑rich motifs (e.g., PWWP‑like) in diverse proteins. | 1.92 | 2 |
| #6360 | N-terminal low-complexity IDRs | Eukaryotic intrinsically disordered, low‑complexity N‑terminal tails enriched in Lys/Arg/Ser/Asp/Glu that precede structured domains of RNA biogenesis/translation and chromatin‑remodeling factors (nucleolar/ssu‑processome, spliceosome/DEAD‑box helicases, H2A.Z chaperones); these flexible segments are typical interaction/assembly regions. | 1.89 | 2 |
| #10131 | Conserved aromatic packing hotspots | Aromatic-residue hotspot detector: the feature marks phenylalanine/tyrosine sites that create conserved hydrophobic packing or aromatic-pocket/cage elements within many globular domains and repeat scaffolds, often at analogous positions across duplicated/repeat units. | 1.88 | 2 |
| #2494 | Structured-to-disordered boundary segments | A short, terminal or domain-edge coil/loop segment—most often a C‑terminal tail or a loop-to-helix/strand capping region—that marks the transition from a structured core to a flexible/low‑complexity extension. These segments frequently show reduced structural confidence (low pLDDT), are enriched in charged and glycine residues, and occasionally embed specific motifs (e.g., metal-binding cysteines), but typically lack discrete annotated active sites. | 1.86 | 4 |
| #4945 | Nucleotidyltransferase phosphate-binding loop | Short glycine‑biased active‑site loop of nucleotidyltransferases that positions the nucleotide triphosphate and helps coordinate catalytic metal ions—i.e., the phosphate‑binding/positioning loop immediately flanking the conserved acidic metal‑binding residues across diverse NTase folds (polymerase/palm, PAP/TUT/CCA‑adding, CD‑NTase/cGAS/OAS, adenylyl/uridylyltransferases, and cytidylyltransferases); often exhibits GxG/GG and nearby aromatic/basic residues. | 1.79 | 2 |
| #5761 | Basic helical small protein activation | a broad, whole-protein activation feature for small bacterial and organellar proteins, often helix-rich and Lys/Arg-enriched, with peaks landing on functional helices that engage macromolecular targets — including the HTH DNA-binding modules of transcriptional regulators, internal transmembrane helices of small inner-membrane enzymes, and helical domains of nucleic-acid/protein-binding proteins. | 1.78 | 2 |
| #2308 | RNase H-like DDE nuclease core | RNase H–like two‑metal‑ion nuclease catalytic core shared by retroviral integrases and related mobile‑element nucleases (integrases/transposases) and RNase H domains, with strongest signal in the C‑terminal portion of the integrase catalytic domain and adjacent helices surrounding the acidic DDE/DDD metal‑binding site | 1.76 | 2 |
| #1601 | Charged helix-capping junction motif | A short, sequence-level motif marking secondary-structure junctions—especially coil/β-strand to α-helix starts, helix capping/termination sites, and β‑hairpin turns—characterized by polar/charged residues with helix-breakers (e.g., Lys/Arg/Glu/His with Pro/Gly/Thr/Ser) and frequent di‑basic Lys pairs. The motif often sits at C-termini or domain boundaries and is frequently adjacent to functional sites (ligand-binding pockets, disulfide/metal linkages), making it a broadly used structural transition/capping signal across taxa and localizations. | 1.74 | 3 |
| #5381 | SF2 helicase motif VI loop | Arginine/glycine–rich “motif VI” loop of SF2 helicases in the C‑terminal RecA-like domain—the arginine‑finger–bearing linker (e.g., HRIGRTGR/RTGRFGR/QTRGR) that couples ATP hydrolysis to nucleic‑acid binding/translocation; recognized as a short flexible helix→loop→β-strand junction conserved across RNA and DNA helicase/translocase families (including chromatin remodelers and viral SF2 helicases) | 1.72 | 2 |
| #10820 | Short helical activation segments | Short helical or mixed secondary-structure segments within catalytic enzymes (notably M16 family zinc metallopeptidases such as stromal processing peptidase, presequence protease/pitrilysin, and PqqE) and some non-catalytic mitochondrial complex subunits. Activation tends to localize to a C-terminal or internal region of the mature chain, distinct from the active site. | 1.71 | 2 |
| #13047 | Membrane-docking amphipathic helices | Short helical patches (often amphipathic and enriched in basic/aromatic residues) embedded within intrinsically disordered linkers or tails, typically positioned immediately C-terminal to signal/translocation motifs, adjacent to transmembrane segments, or at the edge of folded domains/coiled-coils. These helices mediate membrane association or adaptor docking in secretion, trafficking, and virulence contexts. | 1.69 | 2 |
| #10816 | Terminal low-complexity interaction tails | Terminal low‑complexity interaction tails (predominantly C‑terminal) that are disordered‑to‑helical segments enriched in polar/charged residues with interspersed hydrophobics, carrying short linear motifs or key residues used for partner docking (protein or nucleic acid), oligomerization, or proximity to regulatory redox sites | 1.66 | 4 |
| #2980 | Extracellular and lumenal receptor-binding modules | Extracellular and organelle-lumenal recognition/adhesion modules and their flexible linkers in secreted, surface-exposed, or organelle-targeted proteins—most prominently CBMs and phage tail-spike receptor-binding regions—together with related beta-rich recognition domains (e.g., MAM, MD-2/ML/NPC2, toxin receptor-binding β-domains) and N-terminal segments immediately following signal or transit peptides; the feature avoids catalytic glycosidase cores and transmembranes and favors glycine/serine/threonine- and acidic-rich loops studded with aromatics used for glycan/lipid/protein ligand binding. | 1.65 | 3 |
| #3402 | Charged low-complexity interface patches | Short, low-complexity, charged/polar segments and adjacent helical patches (often containing Lys/Arg clusters), used widely across proteins for assembly, targeting, or intra-/inter-subunit contacts and nucleic-acid binding interfaces | 1.64 | 2 |
| #6746 | Charged regions and conserved motifs | Charged/polar interaction segments and conserved motifs: the feature highlights mixed-charge segments rich in S/T/E/D/K/R/H that mediate assembly and binding, and also peaks on specific histidine/acidic motifs embedded in structured domains (e.g., HxHxDH in metallo-β-lactamases, GH–WD repeat blades). | 1.51 | 2 |
| #3175 | Short disordered C-terminal tail | Short disordered C-terminal (or near-terminal) tracts (~7–10 aa) that drive a single localized activation peak. Compositional bias varies — A/G/S/P-rich, polybasic (K/R-rich), mixed basic/acidic, or cysteine-rich repeat motifs are all observed — so the feature is best characterized by its positional/structural context (short disordered tails or terminal segments) rather than by a single residue composition. | 1.47 | 2 |
| #14850 | Active-site cofactor-binding loops | Short, active-site–adjacent segments that bind or coordinate small-molecule cofactors and metal centers—especially FMN/PLP, F420, 2-oxoglutarate, Zn2+, and [4Fe-4S]—and closely related substrate-recognition motifs; typically β-strand/loop or β-turn–β elements (occasionally a short helix), enriched in Gly plus aromatic and charged residues, and located in mid-to-late regions of the protein. | 1.46 | 3 |
| #11022 | Conserved loop motif activations | Short conserved sequence motifs in loop/turn regions of diverse enzymes, frequently containing a mix of basic (R/Q), polar (T/N/S), and acidic (D/E) residues, located at internal positions away from the canonical catalytic residues | 1.46 | 2 |
| #13374 | Post-cleavage extracellular edge hotspots | Signature of extracytoplasmic/envelope-associated structural regions—preferentially the mature (post-cleavage) portions of virion and surface/secreted proteins and single-pass membrane anchors—with peaks at secondary-structure edges (helix/strand caps, loop→strand transitions and within beta-strands); avoided in signal peptides and well-folded accessory C-terminal domains | 1.44 | 2 |
Shared features by |Δ| activation — 1008 total (prevalence ≥ 2)
| Feature | Label | Description | Δ activation | Iso act. | Canon act. | Iso prev. | Canon prev. |
|---|---|---|---|---|---|---|---|
| #1135 | Homopolymeric low-complexity tracts | Compositionally biased, low-complexity sequence segments characterized by homopolymeric residue runs; the feature detects local stretches of repeated single amino acids regardless of chemistry (polar, basic, or hydrophobic), most often within disordered or unstructured contexts but also occasionally within folded domains where such runs occur. | +10.06 | 11.83 | 1.77 | 63 | 4 |
| #6484 | Prion-like low-complexity IDRs | Polar low-complexity intrinsically disordered regions (IDRs)—especially Q/N-rich (prion-like) segments and S/T- and glycine-rich tracts with simple SG/TG/PG repeats—typically located in inter-domain linkers and terminal tails of eukaryotic proteins and often harboring Ser/Thr phosphorylation sites. | +9.76 | 12.23 | 2.47 | 63 | 3 |
| #6125 | Disordered PTM/cleavage SLiMs | Short linear motifs in intrinsically disordered/low-complexity regions, often S/T/P/G- and aromatic-rich, that include proteolytic-processing/PTM hotspots and flexible linkers adjacent to transmembrane segments; the feature favors flexible coils or short amphipathic helices and systematically avoids hydrophobic transmembrane cores | +6.19 | 8.67 | 2.48 | 11 | 8 |
| #4992 | Small-residue disordered N-terminus | Context-dependent free N-terminus signature: short, disordered amino-terminal stretches beginning with the initiator Met followed by residues compatible with MAP cleavage or N-acetylation (e.g., M–G/A/S/T/V, and other contexts including M–D/E/L); strongest emphasis on small residues at positions 2–3; activation is absent when the native N-terminus is structured or starts with a signal peptide/transmembrane segment. | +6.16 | 17.31 | 11.15 | 5 | 2 |
| #6398 | Polar low-complexity disordered stretches | Polar/small-residue-enriched (often Ser/Thr- and Pro-rich) stretches, frequently within intrinsically disordered or low-complexity regions and N-terminal tails/propeptides; can extend into short, flexible helices in small proteins, and is also seen in localized patches within folded domains | +5.35 | 8.00 | 2.65 | 19 | 11 |
| #9000 | Glutamate-biased acidic tract detector | Detector of glutamate identity and glutamate-enriched acidic tracts: strong activation on E residues, especially within acidic, low‑complexity/disordered regions; weaker, sporadic responses on D; largely domain-, function-, and taxonomy-agnostic. | +4.54 | 9.37 | 4.82 | 20 | 15 |
| #13480 | Alanine and small-residue enrichment | A sequence-composition feature that detects small, non-aromatic residues—especially alanine—and their enrichment, lighting up individual Ala residues and stretches enriched in A/G/S/T/P/V (alanine-rich, low-complexity segments) irrespective of protein function or secondary structure. | +4.50 | 9.25 | 4.75 | 22 | 6 |
| #11023 | Low-complexity disordered regions | Intrinsically disordered, low‑complexity segments—often N‑terminal tails or leader/signal regions—enriched in Ser/Pro/Gly/Arg and predicted as coils/low‑confidence structure; the feature marks flexible, poorly structured regions rather than a specific function | +4.39 | 6.06 | 1.67 | 8 | 2 |
| #583 | Asp-biased disordered region detector | Detector of Asp residues embedded in low-complexity/disordered segments, with a strong bias for Asp over Glu; fires broadly in polar or acidic-enriched tracts and only weakly on isolated Asp in structured domains. | +4.39 | 9.83 | 5.44 | 11 | 10 |
| #9395 | Short terminal targeting segment | Terminal or short internal targeting/interaction segment: a short, compositionally biased or hydrophobic patch near a protein terminus, or occasionally an internal low‑complexity loop, that functions as a signal peptide or signal‑anchor (Sec/ER/thylakoid/Tat entry) when hydrophobic, or as a basic/Ser/Lys/Arg/Pro‑rich low‑complexity segment that mediates localization or partner binding in soluble proteins; typically the single prominent terminal segment of short secreted, membrane, or regulatory proteins | +4.22 | 6.31 | 2.09 | 7 | 4 |
| #8412 | Valine-centered residue detector | Detector that fires on valine residues across diverse sequence contexts, including hydrophobic transmembrane helices and signal peptides as well as Val-containing tandem repeats (elastomeric Gly-rich lamprin-like repeats, Pro/Arg-rich disordered repeats) and isolated Val positions in soluble globular proteins. | +4.14 | 10.31 | 6.17 | 11 | 9 |
| #6356 | Ser/Gly/Pro-rich low-complexity IDRs | Often fires on intrinsically disordered, low-complexity regions (IDRs) enriched in Ser/Gly/Pro and basic (Lys/Arg) residues, frequently in secreted/viral proteins, but also activates on short peptides and small folded proteins with no clear sequence motif. Includes Ser-rich PTM-prone segments within unstructured regions. | +4.07 | 7.25 | 3.18 | 32 | 7 |
| #14891 | Disordered termini and cleavage detector | Residue-level detector of intrinsically disordered, flexible termini and proteolytic processing junctions—especially N-termini (including the first residue of the mature chain after propeptide/leader cleavage)—with no strict amino‑acid specificity and a bias toward small/charged residues; common in small, secreted/precursor and viral proteins but also present in disordered regions of diverse proteins. | +4.04 | 6.91 | 2.87 | 21 | 2 |
| #5021 | Polybasic intrinsically disordered regions | Low-complexity, often intrinsically disordered regions in short proteins, precursors, and microproteins across taxa, including viral accessory proteins, neuropeptide/hormone precursors, sperm nuclear proteins, flexible enzyme tails/inserts, and antisense-derived/uncharacterized microproteins. Activation is biased toward, but not restricted to, basic (Lys/Arg) clusters, with frequent peaks also at Pro/Ser/Gly residues in disordered context. | +3.97 | 5.52 | 1.54 | 11 | 3 |
| #3926 | Intrinsic disorder/low-complexity detector | Intrinsic disorder/low-complexity detector: activates on flexible, polar/charged, Q/N/S/T/E/D/K/R- and G/P-enriched segments—often N‑terminal tails—typical of intrinsically disordered regions in small secreted peptides/toxins and viral or eukaryotic regulatory proteins, including Q/N‑rich tracts; signals low-confidence coils or labile helices rather than folded domains. | +3.95 | 6.25 | 2.30 | 9 | 6 |
| #14056 | Q/H-rich low-complexity IDRs | Intrinsically disordered, low‑complexity regions enriched for glutamine and histidine (often with proline/glycine runs)—flexible activation/linker segments that frequently flank structured cores (e.g., DNA‑binding or catalytic domains) in eukaryotic proteins. They are especially common in plant transcription factors, but also occur in secreted peptide precursors and in short terminal tails of enzymes (e.g., Fe(II)/2OG dioxygenases). The feature avoids folded domains and metal‑binding motifs and highlights long polar, compositionally biased coils (including poly‑Q/His patches). | +3.92 | 7.27 | 3.35 | 13 | 5 |
| #11526 | Glycine amidation motif detector | Residue-level detector for small/flexible residues—especially glycine—in short, low-structure linkers and proteolytic processing signals of peptide precursors, with a strong preference for the C‑terminal amidation context in which a glycine (amide donor) immediately precedes mono/di‑basic residues (G‑K/R); outside precursors it gives sparse hits on similar small-residue sites in intrinsically disordered regions and occasionally within structured domains across diverse taxa. | +3.80 | 6.59 | 2.79 | 18 | 12 |
| #2692 | Compositionally biased disordered LCRs | Feature targets compositionally biased, intrinsically disordered low‑complexity regions with long contiguous runs strongly enriched in small/polar (Gly/Ser/Asn/Thr) or acidic (Asp/Glu) residues; occasional activation on highly basic protamine‑like LCRs. Mere disorder, generic tails, or coiled‑coils are insufficient without such compositional bias. | +3.79 | 7.07 | 3.29 | 33 | 8 |
| #1722 | S/T/Pro-rich disordered regions | Intrinsically disordered, low-complexity sequence elements enriched in Ser/Thr/Pro/polar residues, characteristic of flexible linkers, regulatory tails, and polar/low-complexity tracts across diverse taxa. | +3.69 | 6.80 | 3.12 | 22 | 13 |
| #15526 | Polar low-complexity disordered regions | Intrinsically disordered, low‑complexity regions enriched in polar/acidic and amide residues—especially Q/E/S/T/P/G—often including proline‑rich tracts and S/T clusters; these segments are flexible, cysteine‑poor, and typically lie outside structured catalytic cores (e.g., disordered tails, propeptides, and regulatory regions). | +3.62 | 6.65 | 3.03 | 7 | 7 |
| #13897 | Proline-rich IDR detector | Detector of proline residues, with strongest signal in proline-rich, intrinsically disordered, low-complexity segments (often at termini, propeptides, and surface-exposed regions). The feature also activates on more isolated prolines in compact protein contexts, including within transmembrane helices, though typically at lower intensity than in extended Pro-rich disorder. | +3.62 | 8.07 | 4.45 | 16 | 9 |
| #10931 | Polar/proline-rich IDR tails | Compositionally biased, low-complexity segments enriched in polar/proline residues (Q/N/H/P with frequent S/G), typically in intrinsically disordered regions and often in C-terminal tails of transcription factors and other proteins | +3.47 | 6.05 | 2.58 | 6 | 4 |
| #3054 | Basic GP/PTS low-complexity tracts | Intrinsically disordered, low‑complexity segments enriched in glycine/proline and serine/threonine, often containing clusters of basic residues (Lys/Arg) and simple repeats; includes PTS- and GP-rich repeats in secreted/processed peptides and extracellular proteins, as well as basic, Ser/Arg‑rich viral and micropeptide regions. Activation requires extended low‑complexity tracts with these compositions; compact or acidic sequences lacking such tracts are typically not activated. | +3.46 | 6.25 | 2.79 | 22 | 10 |
| #2038 | Disordered low-complexity regulatory regions | Low-complexity, intrinsically disordered regulatory regions enriched for serine/threonine and glutamine/asparagine (often with glycine/proline/histidine runs), typically located in N- or C-terminal tails and flexible linkers of eukaryotic proteins; these segments include transcriptional activation/repression regions and phosphorylation-modulated interaction sites, while structured catalytic or DNA/RNA-binding domains are not targeted. | +3.44 | 5.39 | 1.96 | 60 | 5 |
| #7583 | Short disordered low-complexity segments | Short, intrinsically disordered low-complexity segments—including propeptide regions and other disordered stretches—with a bias toward small (Gly/Pro/Ser/Ala/Asn) and basic (Lys/Arg) residues | +3.23 | 6.88 | 3.65 | 22 | 12 |
| #9998 | Weak N+3 positional cue | Absolute N-terminal positional cue centered near the fourth residue (N+3), independent of amino-acid identity; most often within signal/transit peptides of precursors but also present in non-secreted proteins; the effect is weak and often does not produce a discrete detected site. | +3.10 | 12.38 | 9.28 | 3 | 3 |
| #787 | Basic disordered microprotein regions | The feature activates within small proteins, viral accessory/regulatory proteins, and short ORFs, often on disordered or low-complexity segments. Activation is frequently associated with basic-residue–containing patches (Lys/Arg) but is not restricted to a particular position in the protein. | +3.07 | 4.97 | 1.90 | 19 | 2 |
| #1803 | Unknown generic feature | Unknown generic feature | +11.54 | 23.66 | 12.12 | 245 | 179 |
| #14534 | Unknown generic feature | Unknown generic feature | +8.38 | 24.86 | 16.48 | 246 | 182 |
| #9005 | Unknown generic feature | Unknown generic feature | +8.22 | 21.06 | 12.84 | 209 | 157 |
Part 2 · Differential coordinates
Features firing on just the isoform-unique (extension / alt-frame) residues.
Unique-region features — 278 distinct (prevalence ≥ 2)
| Feature | Label | Description | Activation | Prevalence |
|---|---|---|---|---|
| #4992 | Small-residue disordered N-terminus | Context-dependent free N-terminus signature: short, disordered amino-terminal stretches beginning with the initiator Met followed by residues compatible with MAP cleavage or N-acetylation (e.g., M–G/A/S/T/V, and other contexts including M–D/E/L); strongest emphasis on small residues at positions 2–3; activation is absent when the native N-terminus is structured or starts with a signal peptide/transmembrane segment. | 17.31 | 3 |
| #5725 | UBC/E2-like fold recognition | UBC/E2-like fold recognition across ubiquitin and ubiquitin-like conjugation systems, capturing catalytically active E2s and inactive homologs (UEV, RWD) and the cognate ESCRT components that harbor them; the signal corresponds to the conserved UBC-core surface, with a bias toward the N-terminal α-helix, and can extend to physicochemically similar acidic, α-helical blocks that mimic this surface | 16.41 | 64 |
| #6484 | Prion-like low-complexity IDRs | Polar low-complexity intrinsically disordered regions (IDRs)—especially Q/N-rich (prion-like) segments and S/T- and glycine-rich tracts with simple SG/TG/PG repeats—typically located in inter-domain linkers and terminal tails of eukaryotic proteins and often harboring Ser/Thr phosphorylation sites. | 12.23 | 63 |
| #15545 | Low-complexity disordered regions | Intrinsic disorder/low-complexity segments enriched in small, polar and charged residues (S/T/N/G/A/E/D/R/K), especially N-terminal tails, C-terminal tails, and inter-domain regions that contain short repeating motifs; common in nucleic-acid-associated regulators and many viral proteins, and largely absent from structured domains | 11.84 | 16 |
| #1135 | Homopolymeric low-complexity tracts | Compositionally biased, low-complexity sequence segments characterized by homopolymeric residue runs; the feature detects local stretches of repeated single amino acids regardless of chemistry (polar, basic, or hydrophobic), most often within disordered or unstructured contexts but also occasionally within folded domains where such runs occur. | 11.83 | 61 |
| #8412 | Valine-centered residue detector | Detector that fires on valine residues across diverse sequence contexts, including hydrophobic transmembrane helices and signal peptides as well as Val-containing tandem repeats (elastomeric Gly-rich lamprin-like repeats, Pro/Arg-rich disordered repeats) and isolated Val positions in soluble globular proteins. | 10.31 | 2 |
| #5241 | Early N-terminal positional signal | Generic early N-terminus positional signal peaking at residue ~5–7, found broadly within the disordered N-tail of proteins, including both cleavable leader/targeting segments (signal peptides, propeptides) of precursor proteins and disordered N-terminal regions of mature proteins; residue identity is not specific | 10.04 | 2 |
| #10704 | N-terminal PEST-like IDRs | Low-complexity, often acidic/proline-rich intrinsically disordered N-terminal tails of intracellular eukaryotic proteins, frequently overlapping annotated PEST-like segments. | 9.84 | 39 |
| #9000 | Glutamate-biased acidic tract detector | Detector of glutamate identity and glutamate-enriched acidic tracts: strong activation on E residues, especially within acidic, low‑complexity/disordered regions; weaker, sporadic responses on D; largely domain-, function-, and taxonomy-agnostic. | 9.37 | 4 |
| #13480 | Alanine and small-residue enrichment | A sequence-composition feature that detects small, non-aromatic residues—especially alanine—and their enrichment, lighting up individual Ala residues and stretches enriched in A/G/S/T/P/V (alanine-rich, low-complexity segments) irrespective of protein function or secondary structure. | 9.25 | 16 |
| #15815 | Basic N-terminal targeting patch | Short, basic/polar N‑terminal leader/transit segment immediately after the initiator methionine (positions ~3–6) — i.e., the positively charged/polar start of signal peptides and organelle transit peptides, and more generally low‑acidic, disordered N‑terminal patches (including occasional cysteine‑rich starts) used for targeting, export, rRNA binding, or membrane topology cues | 9.12 | 4 |
| #627 | Generic serine detector | Generic serine detector: the feature marks serine residues, with weaker affinity for threonine (and occasionally proline), and is especially prominent in low-complexity S/T-rich stretches; it is context- and structure-agnostic and does not correspond to a specific functional motif. | 8.69 | 3 |
| #9237 | Glycine detector in disordered regions | Residue-identity detector for glycine (G), with a mild bias toward intrinsically disordered/low-complexity segments (often including very N-terminal residues); pan-taxonomic and function-agnostic, but activation is sparse and frequently absent, becoming detectable mainly in glycine-rich low-complexity contexts; occasional weak responses to other small/polar residues. | 8.52 | 25 |
| #7903 | Arginine-rich polybasic patches | Basic polycationic patches enriched in arginine (often with lysine) within low-complexity/disordered regions, frequently N-terminal. These include classical monopartite NLSs, nucleic acid–binding basic tails, signal peptide n-regions and cytosolic juxtamembrane segments (positive-inside rule), and basic clusters in secreted precursors; typically depleted of acidic residues. | 8.23 | 8 |
| #13897 | Proline-rich IDR detector | Detector of proline residues, with strongest signal in proline-rich, intrinsically disordered, low-complexity segments (often at termini, propeptides, and surface-exposed regions). The feature also activates on more isolated prolines in compact protein contexts, including within transmembrane helices, though typically at lower intensity than in extended Pro-rich disorder. | 8.07 | 5 |
| #6398 | Polar low-complexity disordered stretches | Polar/small-residue-enriched (often Ser/Thr- and Pro-rich) stretches, frequently within intrinsically disordered or low-complexity regions and N-terminal tails/propeptides; can extend into short, flexible helices in small proteins, and is also seen in localized patches within folded domains | 8.00 | 5 |
| #4900 | Arginine-rich segments and residues | Arginine residue identity/basic-tract feature: primarily marks arginine (R) residues—and dense arginine-rich, low-complexity segments—independent of protein family, fold, or function; shows a strong bias for R over other residues (occasional weak lysine signal), appearing both in disordered/repetitive regions and on individual R within structured domains or near transmembrane boundaries. | 8.00 | 8 |
| #678 | Sparse terminal residue peaks | Sparse residue-level activations distributed across short proteins and protein segments, with frequent firing in disordered N-terminal/C-terminal tails as well as in some short transmembrane or compact domains | 7.66 | 16 |
| #14493 | Short hydrophobic helical runs | Short hydrophobic helical segments rich in I/V/F (sometimes M), recognized in both membrane-targeting/insertion contexts (signal peptides, signal-anchors, transmembrane helices) and in hydrophobic helical patches within otherwise soluble small proteins. The feature flags contiguous aliphatic/hydrophobic runs across diverse proteins, including viral membrane/accessory proteins, secreted peptide/effector precursors, small ORFs/microproteins, and short basic/uncharacterized human proteins. | 7.42 | 3 |
| #11429 | G/S-rich disordered low-complexity | Intrinsically disordered, low‑complexity regions enriched in glycine and serine (with frequent threonine and Q/N tracts)—i.e., SG‑repeat and polar low‑complexity segments—common in eukaryotic regulators and some viral proteins, and largely absent from folded catalytic cores. | 7.32 | 29 |
| #14056 | Q/H-rich low-complexity IDRs | Intrinsically disordered, low‑complexity regions enriched for glutamine and histidine (often with proline/glycine runs)—flexible activation/linker segments that frequently flank structured cores (e.g., DNA‑binding or catalytic domains) in eukaryotic proteins. They are especially common in plant transcription factors, but also occur in secreted peptide precursors and in short terminal tails of enzymes (e.g., Fe(II)/2OG dioxygenases). The feature avoids folded domains and metal‑binding motifs and highlights long polar, compositionally biased coils (including poly‑Q/His patches). | 7.27 | 8 |
| #6356 | Ser/Gly/Pro-rich low-complexity IDRs | Often fires on intrinsically disordered, low-complexity regions (IDRs) enriched in Ser/Gly/Pro and basic (Lys/Arg) residues, frequently in secreted/viral proteins, but also activates on short peptides and small folded proteins with no clear sequence motif. Includes Ser-rich PTM-prone segments within unstructured regions. | 7.25 | 25 |
| #11159 | Glycine residue identity detector | Residue-identity detector for glycine (G), largely context-agnostic but with a tendency to fire on glycines in flexible/disordered, often N-terminal or membrane-proximal regions; also active on glycines within transmembrane helices; broadly taxon- and function-independent. | 7.19 | 25 |
| #2692 | Compositionally biased disordered LCRs | Feature targets compositionally biased, intrinsically disordered low‑complexity regions with long contiguous runs strongly enriched in small/polar (Gly/Ser/Asn/Thr) or acidic (Asp/Glu) residues; occasional activation on highly basic protamine‑like LCRs. Mere disorder, generic tails, or coiled‑coils are insufficient without such compositional bias. | 7.07 | 25 |
| #14534 | Unknown generic feature | Unknown generic feature | 24.86 | 65 |
| #1803 | Unknown generic feature | Unknown generic feature | 23.66 | 65 |
| #9005 | Unknown generic feature | Unknown generic feature | 21.06 | 65 |
| #14895 | Unknown generic feature | Unknown generic feature | 20.17 | 65 |
| #9214 | Unknown generic feature | Unknown generic feature | 15.84 | 65 |
| #9194 | Unknown generic feature | Unknown generic feature | 13.92 | 65 |
Top-K sparse-autoencoder activations on the ESM-C residual stream. Activation is a feature's peak value over its region; Prevalence is the number of residues it fires on; Δ activation is isoform − canonical. Only the top 30 features per category are saved; generic / "unknown" features are hidden until you expand a table. Read more about SAE features →
Clinical variants
Differential region — N-terminal extension (isoform-unique)
| Pos (iso) | AA change | Consequence | Source | Clin. sig. | AF (gnomAD) | Impact | AlphaMissense | ESM-C ΔLLR | Link |
|---|---|---|---|---|---|---|---|---|---|
| 0 | T→K | missense_variant | gnomAD | — | 1.03e-04 | — | N/A | — | chr19-58558575-G-T |
| 2 | A→A | synonymous_variant | gnomAD | — | 7.91e-05 | — | N/A | 0.00 | chr19-58558568-G-A |
| 4 | A→V | missense_variant | gnomAD | — | 5.55e-05 | — | N/A | -1.86 | chr19-58558563-G-A |
| 4 | A→G | missense_variant | gnomAD | — | 5.55e-05 | — | N/A | -0.02 | chr19-58558563-G-C |
| 7 | A→A | synonymous_variant | gnomAD | — | 3.49e-05 | — | N/A | 0.00 | chr19-58558553-C-T |
| 14 | G→G | synonymous_variant | gnomAD | — | 1.85e-05 | — | N/A | 0.00 | chr19-58558532-T-G |
| 14 | — | frameshift_variant | gnomAD | — | 1.85e-05 | LoF | — | — | chr19-58558532-TCCGCTCCTCTCCGCGCCCA-T |
| 15 | — | inframe_deletion | gnomAD | — | 1.35e-05 | — | N/A | — | chr19-58558529-CGCTCCGCTCCTCTCCGCGCCCACCGCGCGGCCCGCCTCG-C |
| 16 | A→T | missense_variant | gnomAD | — | 1.75e-04 | — | N/A | -2.84 | chr19-58558528-C-T |
| 18 | Q→E | missense_variant | gnomAD | — | 9.74e-06 | — | N/A | 1.02 | chr19-58558522-G-C |
| 23 | A→E | missense_variant | gnomAD | — | 5.45e-06 | — | N/A | -1.92 | chr19-58558506-G-T |
| 24 | A→T | missense_variant | gnomAD | — | 1.63e-05 | — | N/A | -3.77 | chr19-58558504-C-T |
| 25 | A→A | synonymous_variant | gnomAD | — | 4.53e-06 | — | N/A | 0.00 | chr19-58558499-T-C |
| 26 | A→P | missense_variant | gnomAD | — | 4.25e-06 | — | N/A | -2.91 | chr19-58558498-C-G |
| 28 | E→K | missense_variant | gnomAD | — | 3.55e-06 | — | N/A | -2.48 | chr19-58558492-C-T |
| 29 | A→V | missense_variant | gnomAD | — | 3.20e-06 | — | N/A | -2.54 | chr19-58558488-G-A |
| 30 | A→T | missense_variant | gnomAD | — | 9.24e-06 | — | N/A | -3.92 | chr19-58558486-C-T |
| 31 | A→V | missense_variant | gnomAD | — | 3.71e-05 | — | N/A | -2.69 | chr19-58558482-G-A |
| 32 | A→A | synonymous_variant | gnomAD | — | 2.33e-06 | — | N/A | 0.00 | chr19-58558478-C-A |
| 32 | A→E | missense_variant | gnomAD | — | 2.48e-06 | — | N/A | -2.22 | chr19-58558479-G-T |
| 34 | P→R | missense_variant | gnomAD | — | 2.06e-06 | — | N/A | 1.43 | chr19-58558473-G-C |
| 35 | R→K | missense_variant | gnomAD | — | 7.28e-05 | — | N/A | -3.53 | chr19-58558470-C-T |
| 36 | S→R | missense_variant | gnomAD | — | 7.21e-06 | — | N/A | 1.12 | chr19-58558466-G-C |
| 37 | G→G | synonymous_variant | gnomAD | — | 1.70e-06 | — | N/A | 0.00 | chr19-58558463-T-A |
| 38 | G→S | missense_variant | gnomAD | — | 2.45e-05 | — | N/A | -1.64 | chr19-58558462-C-T |
| 40 | A→A | synonymous_variant | gnomAD | — | 1.46e-06 | — | N/A | 0.00 | chr19-58558454-C-G |
| 40 | A→V | missense_variant | gnomAD | — | 1.48e-06 | — | N/A | -1.44 | chr19-58558455-G-A |
| 40 | A→E | missense_variant | gnomAD | — | 1.48e-06 | — | N/A | -1.34 | chr19-58558455-G-T |
| 41 | G→D | missense_variant | gnomAD | — | 1.01e-05 | — | N/A | -3.19 | chr19-58558452-C-T |
| 43 | G→D | missense_variant | gnomAD | — | 1.33e-06 | — | N/A | -3.21 | chr19-58558446-C-T |
| 44 | G→A | missense_variant | gnomAD | — | 1.34e-06 | — | N/A | -0.41 | chr19-58558443-C-G |
| 45 | G→G | synonymous_variant | gnomAD | — | 1.55e-05 | — | N/A | 0.00 | chr19-58558439-C-T |
| 45 | G→V | missense_variant | gnomAD | — | 1.28e-06 | — | N/A | -2.59 | chr19-58558440-C-A |
| 45 | G→E | missense_variant | gnomAD | — | 1.92e-05 | — | N/A | -2.25 | chr19-58558440-C-T |
| 46 | P→L | missense_variant | gnomAD | — | 2.50e-06 | — | N/A | -1.36 | chr19-58558437-G-A |
| 46 | P→S | missense_variant | gnomAD | — | 1.29e-06 | — | N/A | 1.17 | chr19-58558438-G-A |
| 47 | G→D | missense_variant | gnomAD | — | 1.24e-06 | — | N/A | -3.70 | chr19-58558434-C-T |
| 48 | G→V | missense_variant | gnomAD | — | 1.22e-06 | — | N/A | -2.24 | chr19-58558431-C-A |
| 49 | R→R | synonymous_variant | gnomAD | — | 1.19e-06 | — | N/A | 0.00 | chr19-58558427-C-G |
| 49 | R→R | synonymous_variant | gnomAD | — | 1.19e-06 | — | N/A | 0.00 | chr19-58558427-C-T |
| 49 | R→Q | missense_variant | gnomAD | — | 9.50e-05 | — | N/A | -2.95 | chr19-58558428-C-T |
| 49 | R→G | missense_variant | gnomAD | — | 1.20e-06 | — | N/A | 1.34 | chr19-58558429-G-C |
| 49 | R→Q | missense_variant | COSMIC | — | — | — | N/A | -2.95 | COSV53043558 |
| 50 | — | frameshift_variant | gnomAD | — | 1.17e-06 | LoF | — | — | chr19-58558424-ACCCCGGCCACCCGGCC-A |
| 50 | G→D | missense_variant | gnomAD | — | 2.33e-06 | — | N/A | -2.92 | chr19-58558425-C-T |
| 50 | — | frameshift_variant | gnomAD | — | 4.65e-06 | LoF | — | — | chr19-58558425-CCCCGGCCA-C |
| 51 | P→P | synonymous_variant | gnomAD | — | 1.84e-05 | — | N/A | 0.00 | chr19-58558421-G-A |
| 51 | P→S | missense_variant | gnomAD | — | 1.12e-06 | — | N/A | -0.16 | chr19-58558423-G-A |
| 52 | G→G | synonymous_variant | gnomAD | — | 3.25e-06 | — | N/A | 0.00 | chr19-58558418-C-A |
| 52 | G→G | synonymous_variant | COSMIC | — | — | — | N/A | 0.00 | COSV53042756 |
| 53 | P→H | missense_variant | gnomAD | — | 3.08e-06 | — | N/A | -2.28 | chr19-58558416-G-T |
| 54 | R→H | missense_variant | gnomAD | — | 1.88e-06 | — | N/A | -3.27 | chr19-58558413-C-T |
| 55 | G→D | missense_variant | gnomAD | — | 9.21e-07 | — | N/A | -3.35 | chr19-58558410-C-T |
| 55 | G→S | missense_variant | gnomAD | — | 5.56e-06 | — | N/A | -0.93 | chr19-58558411-C-T |
| 56 | G→G | synonymous_variant | gnomAD | — | 2.06e-05 | — | N/A | 0.00 | chr19-58558406-G-A |
| 56 | G→D | missense_variant | gnomAD | — | 1.78e-06 | — | N/A | -3.48 | chr19-58558407-C-T |
| 56 | G→S | missense_variant | gnomAD | — | 6.32e-06 | — | N/A | -1.23 | chr19-58558408-C-T |
| 57 | G→G | synonymous_variant | gnomAD | — | 8.77e-07 | — | N/A | 0.00 | chr19-58558403-G-T |
| 57 | — | inframe_deletion | gnomAD | — | 1.75e-06 | — | N/A | — | chr19-58558403-GCCGCCGCCGCGGGGCCCGGGACCCCGGCCACCCGGCCCC-G |
| 57 | G→V | missense_variant | gnomAD | — | 1.75e-06 | — | N/A | -3.30 | chr19-58558404-C-A |
| 57 | G→C | missense_variant | gnomAD | — | 1.77e-06 | — | N/A | -3.39 | chr19-58558405-C-A |
| 57 | G→R | missense_variant | gnomAD | — | 2.65e-06 | — | N/A | -1.29 | chr19-58558405-C-G |
| 57 | G→S | missense_variant | gnomAD | — | 2.14e-04 | — | N/A | -1.64 | chr19-58558405-C-T |
| 58 | S→I | missense_variant | gnomAD | — | 8.60e-07 | — | N/A | -4.58 | chr19-58558401-C-A |
| 58 | S→G | missense_variant | gnomAD | — | 3.43e-05 | — | N/A | 3.43 | chr19-58558402-T-C |
| 58 | — | inframe_insertion | gnomAD | — | 1.32e-05 | — | N/A | — | chr19-58558402-T-TGCC |
| 58 | — | inframe_insertion | gnomAD | — | 1.76e-06 | — | N/A | — | chr19-58558402-T-TGCCGCC |
| 58 | — | inframe_insertion | gnomAD | — | 1.76e-06 | — | N/A | — | chr19-58558402-T-TGCCGCCGCC |
| 58 | — | inframe_deletion | gnomAD | — | 1.23e-05 | — | N/A | — | chr19-58558402-TGCCGCC-T |
| 58 | — | inframe_deletion | gnomAD | — | 1.76e-06 | — | N/A | — | chr19-58558402-TGCCGCCGCC-T |
| 58 | S→G | missense_variant | COSMIC | — | — | — | N/A | 3.43 | COSV53040858 |
| 59 | G→G | synonymous_variant | gnomAD | — | 8.52e-07 | — | N/A | 0.00 | chr19-58558397-G-T |
| 59 | — | inframe_deletion | gnomAD | — | 5.16e-05 | — | N/A | — | chr19-58558399-CGCT-C |
| 59 | G→S | missense_variant | COSMIC | — | — | — | N/A | -2.09 | COSV53040494 |
| 60 | G→C | missense_variant | gnomAD | — | 1.70e-06 | — | N/A | -3.83 | chr19-58558396-C-A |
| 60 | — | inframe_deletion | gnomAD | — | 6.80e-06 | — | N/A | — | chr19-58558396-CGCCGCT-C |
| 61 | G→G | synonymous_variant | gnomAD | — | 8.18e-07 | — | N/A | 0.00 | chr19-58558391-G-T |
| 61 | G→S | missense_variant | gnomAD | — | 8.20e-07 | — | N/A | -1.97 | chr19-58558393-C-T |
| 61 | — | inframe_deletion | gnomAD | — | 1.56e-05 | — | N/A | — | chr19-58558393-CGCCGCCGCT-C |
| 61 | G→S | missense_variant | COSMIC | — | — | — | N/A | -1.97 | COSV107228768 |
| 62 | G→G | synonymous_variant | gnomAD | — | 8.16e-07 | — | N/A | 0.00 | chr19-58558388-G-T |
| 62 | — | frameshift_variant | gnomAD | — | 8.16e-07 | LoF | — | — | chr19-58558388-GC-G |
| 62 | G→C | missense_variant | gnomAD | — | 8.14e-07 | — | N/A | -4.17 | chr19-58558390-C-A |
| 62 | — | inframe_insertion | gnomAD | — | 2.44e-06 | — | N/A | — | chr19-58558390-C-CGCCGCCGCCGCT |
| 62 | — | inframe_deletion | gnomAD | — | 8.96e-06 | — | N/A | — | chr19-58558390-CGCCGCCGCCGCT-C |
| 63 | G→V | missense_variant | gnomAD | — | 8.44e-07 | — | N/A | -3.26 | chr19-58558386-C-A |
| 64 | — | inframe_insertion | gnomAD | — | 1.56e-06 | — | N/A | — | chr19-58558382-C-CCTGCCG |
| 64 | R→K | missense_variant | gnomAD | — | 7.83e-07 | — | N/A | 0.12 | chr19-58558383-C-T |
| 64 | R→W | missense_variant | gnomAD | — | 7.99e-07 | — | N/A | -5.89 | chr19-58558384-T-A |
| 64 | R→G | missense_variant | gnomAD | — | 7.99e-07 | — | N/A | 0.89 | chr19-58558384-T-C |
| 64 | — | inframe_insertion | gnomAD | — | 8.78e-05 | — | N/A | — | chr19-58558384-T-TGCC |
| 64 | — | inframe_insertion | gnomAD | — | 7.99e-06 | — | N/A | — | chr19-58558384-T-TGCCGCC |
| 64 | — | inframe_insertion | gnomAD | — | 5.59e-06 | — | N/A | — | chr19-58558384-T-TGCCGCCGCC |
| 64 | — | inframe_deletion | gnomAD | — | 5.56e-04 | — | N/A | — | chr19-58558384-TGCC-T |
| 64 | — | inframe_deletion | gnomAD | — | 1.92e-05 | — | N/A | — | chr19-58558384-TGCCGCC-T |
| 64 | — | inframe_deletion | gnomAD | — | 1.12e-05 | — | N/A | — | chr19-58558384-TGCCGCCGCC-T |
96 variants in the differential region.
Shared canonical core — sequence common to canonical and isoform; AlphaMissense applies here
| Pos (iso) | AA change | Consequence | Source | Clin. sig. | AF (gnomAD) | Impact | AlphaMissense | ESM-C ΔLLR | Link |
|---|---|---|---|---|---|---|---|---|---|
| 65 | — | inframe_deletion | gnomAD | — | 7.86e-07 | — | — | — | chr19-58558381-TCCTGCCGCCGCCGCCGCC-T |
| 66 | I→M | missense_variant | gnomAD | — | 7.54e-07 | damaging | likely_pathogenic (0.66) | -9.12 | chr19-58558376-G-C |
| 68 | L→L | synonymous_variant | gnomAD | — | 7.42e-07 | — | — | 0.00 | chr19-58558370-C-A |
| 68 | L→L | synonymous_variant | gnomAD | — | 1.48e-06 | — | — | 0.00 | chr19-58558370-C-T |
| 69 | F→L | missense_variant | gnomAD | — | 7.16e-07 | damaging | likely_pathogenic (0.99) | -7.19 | chr19-58558367-G-C |
| 70 | S→S | synonymous_variant | gnomAD | — | 6.38e-06 | — | — | 0.00 | chr19-58558364-C-T |
| 71 | L→L | synonymous_variant | gnomAD | — | 1.42e-06 | — | — | 0.00 | chr19-58558361-C-T |
| 71 | L→L | synonymous_variant | gnomAD | — | 6.38e-06 | — | — | 0.00 | chr19-58558363-G-A |
| 71 | L→P | missense_variant | COSMIC | — | — | damaging | likely_pathogenic (0.98) | -12.25 | COSV104570910 |
| 72 | K→K | synonymous_variant | gnomAD | — | 6.28e-06 | — | — | 0.00 | chr19-58558358-C-T |
| 73 | Q→H | missense_variant | gnomAD | — | 1.39e-06 | damaging | ambiguous (0.51) | -7.81 | chr19-58558355-C-G |
| 73 | Q→Q | synonymous_variant | gnomAD | — | 5.56e-06 | — | — | 0.00 | chr19-58558355-C-T |
| 73 | Q→K | missense_variant | gnomAD | — | 6.98e-07 | damaging | likely_benign (0.28) | -9.68 | chr19-58558357-G-T |
| 74 | Q→Q | synonymous_variant | COSMIC | — | — | — | — | 0.00 | COSV107228773 |
| 75 | K→K | synonymous_variant | gnomAD | — | 6.94e-07 | — | — | 0.00 | chr19-58558349-C-T |
| 75 | K→Q | missense_variant | gnomAD | — | 8.65e-05 | damaging | likely_benign (0.19) | -9.87 | chr19-58558351-T-G |
| 75 | — | inframe_deletion | gnomAD | — | 1.40e-06 | — | — | — | chr19-58558351-TCTG-T |
| 75 | K→Q | missense_variant | ClinVar | Uncertain significance | — | damaging | likely_benign (0.19) | -9.87 | ClinVar:3465079 |
| 75 | — | inframe_deletion | COSMIC | — | — | — | — | — | COSV99285296 |
| 75 | K→* | stop_gained | COSMIC | — | — | LoF | — | — | COSV108753444 |
| 75 | K→E | missense_variant | COSMIC | — | — | damaging | ambiguous (0.56) | -10.75 | COSV99285037 |
| 76 | K→N | missense_variant | gnomAD | — | 1.39e-06 | damaging | likely_pathogenic (0.72) | -9.06 | chr19-58558346-C-A |
| 76 | K→K | synonymous_variant | gnomAD | — | 6.94e-07 | — | — | 0.00 | chr19-58558346-C-T |
| 76 | K→E | missense_variant | gnomAD | — | 5.65e-05 | damaging | ambiguous (0.45) | -8.94 | chr19-58558348-T-C |
| 76 | K→Q | missense_variant | gnomAD | — | 6.97e-07 | damaging | likely_benign (0.22) | -8.25 | chr19-58558348-T-G |
| 76 | K→E | missense_variant | ClinVar | Uncertain significance | — | damaging | ambiguous (0.45) | -8.94 | ClinVar:3185534 |
| 76 | K→E | missense_variant | COSMIC | — | — | damaging | ambiguous (0.45) | -8.94 | COSV99285294 |
| 77 | E→E | synonymous_variant | gnomAD | — | 6.93e-07 | — | — | 0.00 | chr19-58558343-C-T |
| 77 | — | inframe_deletion | gnomAD | — | 7.63e-06 | — | — | — | chr19-58558345-CCTT-C |
| 77 | — | inframe_deletion | COSMIC | — | — | — | — | — | COSV53040527 |
| 77 | E→* | stop_gained | COSMIC | — | — | LoF | — | — | COSV53041176 |
| 78 | E→D | missense_variant | gnomAD | — | 6.92e-07 | — | likely_benign (0.09) | -6.03 | chr19-58558340-C-G |
| 78 | E→Q | missense_variant | gnomAD | — | 6.93e-07 | damaging | likely_benign (0.24) | -8.31 | chr19-58558342-C-G |
| 78 | E→K | missense_variant | gnomAD | — | 6.93e-07 | damaging | ambiguous (0.42) | -8.62 | chr19-58558342-C-T |
| 79 | E→E | synonymous_variant | gnomAD | — | 6.92e-07 | — | — | 0.00 | chr19-58558337-C-T |
| 79 | E→A | missense_variant | gnomAD | — | 1.60e-05 | damaging | likely_benign (0.17) | -7.81 | chr19-58558338-T-G |
| 79 | E→* | stop_gained | gnomAD | — | 1.59e-05 | LoF | — | — | chr19-58558339-C-A |
| 80 | S→S | synonymous_variant | gnomAD | — | 2.07e-06 | — | — | 0.00 | chr19-58558334-C-G |
| 80 | S→L | missense_variant | gnomAD | — | 2.07e-06 | — | likely_benign (0.13) | -7.11 | chr19-58558335-G-A |
| 80 | — | inframe_insertion | gnomAD | — | 1.39e-06 | — | — | — | chr19-58558336-A-ACTC |
| 80 | S→P | missense_variant | gnomAD | — | 2.78e-06 | damaging | likely_benign (0.23) | -8.26 | chr19-58558336-A-G |
| 80 | S→A | missense_variant | COSMIC | — | — | — | likely_benign (0.06) | -5.17 | COSV53042996 |
| 81 | A→V | missense_variant | gnomAD | — | 4.15e-06 | — | likely_benign (0.10) | -5.07 | chr19-58558332-G-A |
| 81 | A→V | missense_variant | ClinVar | — | — | — | likely_benign (0.10) | -5.07 | ClinVar:4446113 |
| 81 | A→V | missense_variant | COSMIC | — | — | — | likely_benign (0.10) | -5.07 | COSV53041239 |
| 82 | G→D | missense_variant | COSMIC | — | — | — | likely_benign (0.29) | -7.07 | COSV53040529 |
| 83 | G→R | missense_variant | gnomAD | — | 6.91e-07 | — | ambiguous (0.46) | -3.38 | chr19-58558327-C-G |
| 83 | G→S | missense_variant | gnomAD | — | 6.91e-07 | — | likely_benign (0.08) | -4.44 | chr19-58558327-C-T |
| 84 | T→T | synonymous_variant | gnomAD | — | 6.91e-07 | — | — | 0.00 | chr19-58558322-G-C |
| 84 | T→S | missense_variant | gnomAD | — | 6.91e-07 | — | likely_benign (0.07) | -4.42 | chr19-58558324-T-A |
| 85 | K→K | synonymous_variant | gnomAD | — | 1.38e-06 | — | — | 0.00 | chr19-58558319-C-T |
| 86 | G→G | synonymous_variant | gnomAD | — | 2.76e-06 | — | — | 0.00 | chr19-58558316-G-A |
| 86 | G→G | synonymous_variant | gnomAD | — | 6.90e-07 | — | — | 0.00 | chr19-58558316-G-C |
| 86 | G→S | missense_variant | gnomAD | — | 4.14e-06 | — | likely_benign (0.08) | -1.31 | chr19-58558318-C-T |
| 87 | S→G | missense_variant | COSMIC | — | — | — | likely_benign (0.06) | -3.12 | COSV99285000 |
| 88 | S→N | missense_variant | gnomAD | — | 6.89e-07 | — | likely_benign (0.13) | -4.49 | chr19-58558311-C-T |
| 88 | S→G | missense_variant | gnomAD | — | 6.21e-06 | — | likely_benign (0.06) | -3.87 | chr19-58558312-T-C |
| 88 | S→R | missense_variant | COSMIC | — | — | damaging | likely_pathogenic (0.69) | -9.52 | COSV53042721 |
| 89 | — | inframe_deletion | gnomAD | — | 4.36e-05 | — | — | — | chr19-58558308-TTGC-T |
| 90 | K→K | synonymous_variant | gnomAD | — | 1.38e-06 | — | — | 0.00 | chr19-58558304-C-T |
| 91 | A→A | synonymous_variant | gnomAD | — | 6.89e-07 | — | — | 0.00 | chr19-58558301-C-T |
| 91 | A→V | missense_variant | gnomAD | — | 2.07e-06 | — | ambiguous (0.39) | -6.93 | chr19-58558302-G-A |
| 91 | — | inframe_deletion | gnomAD | — | 6.89e-07 | — | — | — | chr19-58558303-CCTT-C |
| 93 | A→A | synonymous_variant | gnomAD | — | 2.07e-06 | — | — | 0.00 | chr19-58558295-C-T |
| 93 | A→A | synonymous_variant | COSMIC | — | — | — | — | 0.00 | COSV99285258 |
| 93 | A→V | missense_variant | COSMIC | — | — | damaging | likely_pathogenic (0.95) | -9.19 | COSV99285056 |
| 94 | A→A | synonymous_variant | gnomAD | — | 9.65e-06 | — | — | 0.00 | chr19-58558292-C-T |
| 96 | L→L | synonymous_variant | gnomAD | — | 6.89e-07 | — | — | 0.00 | chr19-58558286-C-A |
| 96 | L→P | missense_variant | gnomAD | — | 6.90e-07 | damaging | likely_pathogenic (1.00) | -13.31 | chr19-58558287-A-G |
| 96 | L→L | synonymous_variant | gnomAD | — | 6.89e-07 | — | — | 0.00 | chr19-58558288-G-A |
| 97 | R→R | synonymous_variant | gnomAD | — | 6.89e-07 | — | — | 0.00 | chr19-58558283-C-G |
| 97 | R→W | missense_variant | gnomAD | — | 6.90e-07 | damaging | likely_pathogenic (1.00) | -11.81 | chr19-58558285-G-A |
| 97 | R→R | synonymous_variant | gnomAD | — | 5.52e-06 | — | — | 0.00 | chr19-58558285-G-T |
| 97 | R→W | missense_variant | COSMIC | — | — | damaging | likely_pathogenic (1.00) | -11.81 | COSV108753443 |
| 98 | I→F | missense_variant | gnomAD | — | 6.90e-07 | damaging | likely_pathogenic (0.92) | -11.00 | chr19-58558282-T-A |
| 99 | Q→Q | synonymous_variant | COSMIC | — | — | — | — | 0.00 | COSV53043369 |
| 99 | Q→L | missense_variant | COSMIC | — | — | damaging | likely_pathogenic (0.60) | -10.37 | COSV53042340 |
| 100 | K→E | missense_variant | gnomAD | — | 6.92e-07 | damaging | likely_pathogenic (0.99) | -10.44 | chr19-58558276-T-C |
| 100 | K→N | missense_variant | COSMIC | — | — | damaging | likely_pathogenic (1.00) | -10.00 | COSV53041909 |
| 100 | K→R | missense_variant | COSMIC | — | — | damaging | likely_benign (0.29) | -8.50 | COSV99285214 |
| 101 | D→D | synonymous_variant | gnomAD | — | 6.84e-07 | — | — | 0.00 | chr19-58557156-G-A |
| 102 | I→V | missense_variant | gnomAD | — | 1.27e-04 | — | likely_benign (0.16) | -7.12 | chr19-58557155-T-C |
| 102 | I→K | missense_variant | COSMIC | — | — | damaging | likely_pathogenic (0.99) | -12.75 | COSV53041425 |
| 104 | E→* | stop_gained | COSMIC | — | — | LoF | — | — | COSV99285104 |
| 105 | L→L | synonymous_variant | gnomAD | — | 2.05e-06 | — | — | 0.00 | chr19-58557144-C-T |
| 105 | L→L | synonymous_variant | gnomAD | — | 6.84e-07 | — | — | 0.00 | chr19-58557146-G-A |
| 105 | L→L | synonymous_variant | COSMIC | — | — | — | — | 0.00 | COSV53043101 |
| 106 | N→N | synonymous_variant | gnomAD | — | 1.37e-06 | — | — | 0.00 | chr19-58557141-G-A |
| 107 | L→L | synonymous_variant | gnomAD | — | 6.16e-06 | — | — | 0.00 | chr19-58557138-C-A |
| 107 | L→L | synonymous_variant | gnomAD | — | 6.84e-07 | — | — | 0.00 | chr19-58557138-C-T |
| 110 | T→T | synonymous_variant | gnomAD | — | 1.37e-06 | — | — | 0.00 | chr19-58557129-C-A |
| 110 | T→T | synonymous_variant | gnomAD | — | 3.42e-06 | — | — | 0.00 | chr19-58557129-C-T |
| 110 | T→M | missense_variant | COSMIC | — | — | damaging | likely_pathogenic (0.87) | -8.67 | COSV99285409 |
| 112 | D→G | missense_variant | gnomAD | — | 6.84e-07 | damaging | likely_benign (0.25) | -7.86 | chr19-58557124-T-C |
| 112 | D→H | missense_variant | gnomAD | — | 6.84e-07 | damaging | likely_benign (0.24) | -8.55 | chr19-58557125-C-G |
| 112 | D→N | missense_variant | gnomAD | — | 6.84e-07 | — | likely_benign (0.10) | -6.49 | chr19-58557125-C-T |
| 113 | I→T | missense_variant | gnomAD | — | 1.37e-06 | damaging | likely_pathogenic (0.74) | -6.62 | chr19-58557121-A-G |
| 113 | I→F | missense_variant | gnomAD | — | 6.84e-07 | damaging | ambiguous (0.50) | -10.50 | chr19-58557122-T-A |
| 113 | — | frameshift_variant | COSMIC | — | — | LoF | — | — | COSV53043033 |
| 113 | I→V | missense_variant | COSMIC | — | — | — | likely_benign (0.11) | -6.81 | COSV99284991 |
| 114 | S→S | synonymous_variant | gnomAD | — | 1.37e-06 | — | — | 0.00 | chr19-58557117-G-A |
| 114 | S→I | missense_variant | gnomAD | — | 2.84e-03 | — | likely_benign (0.11) | -6.87 | chr19-58557118-C-A |
| 114 | S→T | missense_variant | gnomAD | — | 1.37e-06 | — | likely_benign (0.06) | -3.13 | chr19-58557118-C-G |
| 114 | S→N | missense_variant | gnomAD | — | 2.74e-06 | — | likely_benign (0.11) | -3.19 | chr19-58557118-C-T |
| 114 | S→I | missense_variant | ClinVar | — | — | — | likely_benign (0.11) | -6.87 | ClinVar:4446110 |
| 114 | S→R | missense_variant | COSMIC | — | — | damaging | ambiguous (0.36) | -7.77 | COSV53042718 |
| 115 | F→L | missense_variant | gnomAD | — | 1.37e-06 | damaging | likely_pathogenic (1.00) | -9.19 | chr19-58557116-A-G |
| 115 | F→L | missense_variant | COSMIC | — | — | damaging | likely_pathogenic (1.00) | -9.19 | COSV105860227 |
| 116 | S→S | synonymous_variant | gnomAD | — | 6.84e-07 | — | — | 0.00 | chr19-58557111-T-C |
| 117 | D→E | missense_variant | gnomAD | — | 6.84e-07 | — | likely_benign (0.21) | -5.30 | chr19-58557108-A-C |
| 117 | D→D | synonymous_variant | gnomAD | — | 2.74e-06 | — | — | 0.00 | chr19-58557108-A-G |
| 117 | D→H | missense_variant | gnomAD | — | 6.84e-07 | damaging | ambiguous (0.45) | -9.68 | chr19-58557110-C-G |
| 117 | D→N | missense_variant | gnomAD | — | 6.84e-07 | — | likely_benign (0.13) | -5.90 | chr19-58557110-C-T |
| 118 | P→A | missense_variant | gnomAD | — | 6.84e-07 | — | likely_benign (0.08) | -5.76 | chr19-58557107-G-C |
| 118 | P→T | missense_variant | gnomAD | — | 5.54e-05 | — | likely_benign (0.10) | -5.89 | chr19-58557107-G-T |
| 118 | P→T | missense_variant | ClinVar | Uncertain significance | — | — | likely_benign (0.10) | -5.89 | ClinVar:3185533 |
| 119 | D→D | synonymous_variant | gnomAD | — | 2.05e-06 | — | — | 0.00 | chr19-58557102-G-A |
| 119 | D→E | missense_variant | gnomAD | — | 6.84e-07 | — | likely_benign (0.17) | -5.87 | chr19-58557102-G-T |
| 119 | D→Y | missense_variant | COSMIC | — | — | damaging | likely_pathogenic (0.69) | -11.56 | COSV107228791 |
| 120 | D→D | synonymous_variant | gnomAD | — | 1.37e-06 | — | — | 0.00 | chr19-58557099-G-A |
| 120 | D→N | missense_variant | gnomAD | — | 1.37e-06 | — | likely_benign (0.23) | -7.31 | chr19-58557101-C-T |
| 120 | D→N | missense_variant | COSMIC | — | — | — | likely_benign (0.23) | -7.31 | COSV105017442 |
| 121 | L→L | synonymous_variant | gnomAD | — | 1.37e-05 | — | — | 0.00 | chr19-58557096-G-A |
| 121 | L→L | synonymous_variant | gnomAD | — | 1.37e-06 | — | — | 0.00 | chr19-58557096-G-C |
| 121 | L→F | missense_variant | gnomAD | — | 6.84e-07 | damaging | likely_pathogenic (0.59) | -7.75 | chr19-58557098-G-A |
| 121 | L→I | missense_variant | gnomAD | — | 6.84e-07 | damaging | likely_benign (0.20) | -8.56 | chr19-58557098-G-T |
| 122 | L→L | synonymous_variant | gnomAD | — | 6.84e-07 | — | — | 0.00 | chr19-58557093-G-T |
| 122 | L→I | missense_variant | COSMIC | — | — | damaging | likely_pathogenic (0.69) | -9.31 | COSV99285506 |
| 123 | N→T | missense_variant | gnomAD | — | 6.84e-07 | — | likely_benign (0.15) | -6.19 | chr19-58557091-T-G |
| 123 | N→S | missense_variant | COSMIC | — | — | — | likely_benign (0.09) | -5.69 | COSV99285507 |
| 124 | F→L | missense_variant | gnomAD | — | 6.84e-07 | damaging | likely_pathogenic (1.00) | -9.19 | chr19-58557087-G-C |
| 125 | K→N | missense_variant | gnomAD | — | 9.58e-06 | — | ambiguous (0.54) | -5.53 | chr19-58557084-C-A |
| 125 | K→K | synonymous_variant | gnomAD | — | 1.37e-06 | — | — | 0.00 | chr19-58557084-C-T |
| 126 | L→L | synonymous_variant | gnomAD | — | 6.16e-06 | — | — | 0.00 | chr19-58557083-G-A |
| 126 | L→L | synonymous_variant | COSMIC | — | — | — | — | 0.00 | COSV53043643 |
| 127 | V→V | synonymous_variant | gnomAD | — | 1.57e-05 | — | — | 0.00 | chr19-58557078-G-A |
| 128 | I→F | missense_variant | gnomAD | — | 6.84e-07 | damaging | likely_pathogenic (0.93) | -9.55 | chr19-58557077-T-A |
| 129 | C→C | synonymous_variant | gnomAD | — | 6.84e-07 | — | — | 0.00 | chr19-58557072-A-G |
| 129 | C→S | missense_variant | gnomAD | — | 6.84e-07 | — | likely_benign (0.12) | -2.29 | chr19-58557073-C-G |
| 129 | C→S | missense_variant | COSMIC | — | — | — | likely_benign (0.12) | -2.29 | COSV53040817 |
| 130 | P→P | synonymous_variant | gnomAD | — | 6.84e-07 | — | — | 0.00 | chr19-58557069-A-T |
| 133 | G→G | synonymous_variant | gnomAD | — | 6.16e-06 | — | — | 0.00 | chr19-58556928-G-A |
| 133 | G→D | missense_variant | COSMIC | — | — | damaging | likely_pathogenic (1.00) | -10.00 | COSV106089309 |
| 134 | F→L | missense_variant | gnomAD | — | 6.84e-07 | damaging | likely_pathogenic (0.95) | -5.01 | chr19-58556925-G-C |
| 135 | Y→Y | synonymous_variant | gnomAD | — | 1.37e-06 | — | — | 0.00 | chr19-58556922-G-A |
| 136 | K→E | missense_variant | COSMIC | — | — | damaging | likely_pathogenic (0.64) | -10.28 | COSV53397009 |
| 137 | S→N | missense_variant | gnomAD | — | 6.84e-07 | — | likely_benign (0.08) | -4.50 | chr19-58556917-C-T |
| 137 | S→N | missense_variant | ClinVar | Uncertain significance | — | — | likely_benign (0.08) | -4.50 | ClinVar:4601254 |
| 138 | G→G | synonymous_variant | gnomAD | — | 1.37e-06 | — | — | 0.00 | chr19-58556913-C-G |
| 139 | K→K | synonymous_variant | gnomAD | — | 6.84e-07 | — | — | 0.00 | chr19-58556910-C-T |
| 139 | K→K | synonymous_variant | COSMIC | — | — | — | — | 0.00 | COSV99448181 |
| 141 | V→V | synonymous_variant | gnomAD | — | 8.21e-06 | — | — | 0.00 | chr19-58556904-C-T |
| 141 | V→A | missense_variant | gnomAD | — | 4.79e-06 | — | likely_benign (0.23) | -4.70 | chr19-58556905-A-G |
| 141 | V→L | missense_variant | gnomAD | — | 6.84e-07 | — | likely_benign (0.17) | -4.49 | chr19-58556906-C-G |
| 141 | — | frameshift_variant | COSMIC | — | — | LoF | — | — | COSV53396861 |
| 143 | S→T | missense_variant | ClinVar | Uncertain significance | — | — | likely_benign (0.08) | -3.66 | ClinVar:2277650 |
| 143 | S→G | missense_variant | COSMIC | — | — | — | likely_benign (0.20) | -6.64 | COSV53396524 |
| 145 | K→K | synonymous_variant | gnomAD | — | 6.84e-07 | — | — | 0.00 | chr19-58556892-C-T |
| 145 | K→Q | missense_variant | gnomAD | — | 6.84e-07 | — | likely_benign (0.19) | -7.17 | chr19-58556894-T-G |
| 145 | K→Q | missense_variant | COSMIC | — | — | — | likely_benign (0.19) | -7.17 | COSV104386669 |
| 148 | Q→R | missense_variant | gnomAD | — | 6.92e-07 | damaging | ambiguous (0.41) | -8.73 | chr19-58556783-T-C |
| 148 | Q→E | missense_variant | gnomAD | — | 6.91e-07 | — | likely_benign (0.17) | -7.14 | chr19-58556784-G-C |
| 149 | G→D | missense_variant | gnomAD | — | 6.92e-07 | damaging | likely_pathogenic (0.92) | -9.43 | chr19-58556780-C-T |
| 150 | Y→Y | synonymous_variant | gnomAD | — | 3.48e-06 | — | — | 0.00 | chr19-58556776-G-A |
| 150 | Y→C | missense_variant | COSMIC | — | — | damaging | likely_pathogenic (1.00) | -11.94 | COSV53397310 |
| 151 | P→P | synonymous_variant | gnomAD | — | 6.99e-07 | — | — | 0.00 | chr19-58556773-C-G |
| 151 | P→P | synonymous_variant | gnomAD | — | 5.31e-05 | — | — | 0.00 | chr19-58556773-C-T |
| 151 | P→L | missense_variant | COSMIC | — | — | damaging | likely_pathogenic (1.00) | -11.81 | COSV105017440 |
| 154 | P→P | synonymous_variant | gnomAD | — | 3.51e-06 | — | — | 0.00 | chr19-58556764-G-T |
| 154 | P→P | synonymous_variant | COSMIC | — | — | — | — | 0.00 | COSV53397969 |
| 154 | P→L | missense_variant | COSMIC | — | — | damaging | likely_pathogenic (0.99) | -11.56 | COSV104582061 |
| 155 | P→H | missense_variant | COSMIC | — | — | damaging | likely_pathogenic (1.00) | -13.19 | COSV53396325 |
| 156 | K→K | synonymous_variant | gnomAD | — | 7.03e-07 | — | — | 0.00 | chr19-58556758-C-T |
| 156 | K→N | missense_variant | COSMIC | — | — | damaging | likely_pathogenic (0.99) | -10.37 | COSV104582067 |
| 157 | V→V | synonymous_variant | COSMIC | — | — | — | — | 0.00 | COSV99448318 |
| 160 | E→E | synonymous_variant | gnomAD | — | 2.13e-06 | — | — | 0.00 | chr19-58556746-C-T |
| 161 | T→S | missense_variant | gnomAD | — | 1.42e-06 | — | ambiguous (0.45) | -6.94 | chr19-58556745-T-A |
| 162 | M→I | missense_variant | gnomAD | — | 7.15e-07 | — | ambiguous (0.45) | -6.16 | chr19-58556740-C-T |
| 162 | M→V | missense_variant | gnomAD | — | 2.86e-06 | — | likely_benign (0.10) | -6.38 | chr19-58556742-T-C |
| 163 | V→I | missense_variant | gnomAD | — | 7.15e-07 | — | likely_benign (0.20) | -6.56 | chr19-58556739-C-T |
| 163 | V→V | synonymous_variant | COSMIC | — | — | — | — | 0.00 | COSV53397937 |
| 163 | V→V | synonymous_variant | COSMIC | — | — | — | — | 0.00 | COSV99447899 |
| 164 | — | frameshift_variant | COSMIC | — | — | LoF | — | — | COSV108778103 |
| 165 | H→H | synonymous_variant | gnomAD | — | 3.58e-06 | — | — | 0.00 | chr19-58556731-G-A |
| 167 | N→N | synonymous_variant | gnomAD | — | 7.19e-07 | — | — | 0.00 | chr19-58556725-G-A |
| 167 | N→S | missense_variant | COSMIC | — | — | damaging | likely_pathogenic (0.95) | -9.87 | COSV53397279 |
| 171 | E→E | synonymous_variant | gnomAD | — | 4.35e-06 | — | — | 0.00 | chr19-58556713-C-T |
| 171 | E→Q | missense_variant | gnomAD | — | 7.24e-07 | damaging | ambiguous (0.36) | -7.90 | chr19-58556715-C-G |
| 171 | E→K | missense_variant | COSMIC | — | — | damaging | likely_pathogenic (0.79) | -9.99 | COSV99448054 |
| 172 | G→G | synonymous_variant | gnomAD | — | 1.45e-06 | — | — | 0.00 | chr19-58556710-G-A |
| 173 | N→N | synonymous_variant | gnomAD | — | 2.90e-06 | — | — | 0.00 | chr19-58556707-G-A |
| 173 | N→D | missense_variant | gnomAD | — | 7.25e-07 | damaging | likely_pathogenic (0.94) | -11.25 | chr19-58556709-T-C |
| 174 | V→V | synonymous_variant | gnomAD | — | 2.90e-06 | — | — | 0.00 | chr19-58556704-G-T |
| 174 | V→V | synonymous_variant | COSMIC | — | — | — | — | 0.00 | COSV53396988 |
| 174 | V→I | missense_variant | COSMIC | — | — | — | likely_benign (0.24) | -5.68 | COSV53398436 |
| 176 | L→L | synonymous_variant | gnomAD | — | 2.91e-05 | — | — | 0.00 | chr19-58556698-G-A |
| 176 | L→L | synonymous_variant | gnomAD | — | 7.26e-07 | — | — | 0.00 | chr19-58556698-G-C |
| 176 | L→P | missense_variant | COSMIC | — | — | damaging | likely_pathogenic (1.00) | -11.62 | COSV53396431 |
| 177 | N→N | synonymous_variant | COSMIC | — | — | — | — | 0.00 | COSV53396949 |
| 177 | N→I | missense_variant | COSMIC | — | — | damaging | likely_pathogenic (1.00) | -14.37 | COSV99448459 |
| 178 | I→I | synonymous_variant | gnomAD | — | 6.90e-05 | — | — | 0.00 | chr19-58556692-G-A |
| 178 | I→I | synonymous_variant | COSMIC | — | — | — | — | 0.00 | COSV99448409 |
| 179 | L→L | synonymous_variant | gnomAD | — | 2.18e-06 | — | — | 0.00 | chr19-58556689-G-A |
| 179 | L→V | missense_variant | gnomAD | — | 7.27e-07 | damaging | likely_pathogenic (0.99) | -11.19 | chr19-58556691-G-C |
| 180 | R→K | missense_variant | COSMIC | — | — | damaging | likely_pathogenic (0.98) | -10.19 | COSV105017435 |
| 180 | R→G | missense_variant | COSMIC | — | — | damaging | likely_pathogenic (1.00) | -10.94 | COSV53397088 |
| 181 | E→K | missense_variant | COSMIC | — | — | damaging | likely_pathogenic (0.99) | -10.12 | COSV53398088 |
| 182 | D→N | missense_variant | COSMIC | — | — | damaging | likely_pathogenic (0.99) | -10.19 | COSV99448338 |
| 183 | W→* | stop_gained | COSMIC | — | — | LoF | — | — | COSV99448402 |
| 184 | K→T | missense_variant | COSMIC | — | — | damaging | likely_pathogenic (0.96) | -12.50 | COSV53396567 |
| 185 | P→P | synonymous_variant | gnomAD | — | 6.84e-07 | — | — | 0.00 | chr19-58556364-T-A |
| 185 | P→P | synonymous_variant | gnomAD | — | 6.84e-07 | — | — | 0.00 | chr19-58556364-T-G |
| 185 | — | frameshift_variant | COSMIC | — | — | LoF | — | — | COSV108778102 |
| 185 | P→T | missense_variant | COSMIC | — | — | damaging | likely_pathogenic (1.00) | -12.31 | COSV99448416 |
| 186 | V→V | synonymous_variant | gnomAD | — | 6.84e-07 | — | — | 0.00 | chr19-58556361-G-T |
| 187 | L→R | missense_variant | COSMIC | — | — | damaging | likely_pathogenic (1.00) | -14.44 | COSV108778100 |
| 188 | T→T | synonymous_variant | gnomAD | — | 6.84e-07 | — | — | 0.00 | chr19-58556355-C-G |
| 188 | T→T | synonymous_variant | gnomAD | — | 4.10e-06 | — | — | 0.00 | chr19-58556355-C-T |
| 188 | T→T | synonymous_variant | COSMIC | — | — | — | — | 0.00 | COSV99447934 |
| 189 | I→V | missense_variant | gnomAD | — | 1.37e-06 | — | likely_benign (0.11) | -7.31 | chr19-58556354-T-C |
| 190 | N→N | synonymous_variant | gnomAD | — | 6.84e-07 | — | — | 0.00 | chr19-58556349-G-A |
| 191 | S→Y | missense_variant | COSMIC | — | — | damaging | likely_pathogenic (0.99) | -12.00 | COSV53396807 |
| 192 | I→V | missense_variant | ClinVar | Uncertain significance | — | — | likely_benign (0.13) | -5.06 | ClinVar:4601253 |
| 192 | I→V | missense_variant | COSMIC | — | — | — | likely_benign (0.13) | -5.06 | COSV105872211 |
| 193 | I→M | missense_variant | ClinVar | — | — | damaging | likely_pathogenic (0.69) | -8.56 | ClinVar:4446107 |
| 195 | G→G | synonymous_variant | gnomAD | — | 2.74e-06 | — | — | 0.00 | chr19-58556334-G-A |
| 196 | L→L | synonymous_variant | gnomAD | — | 6.84e-07 | — | — | 0.00 | chr19-58556331-C-T |
| 196 | — | frameshift_variant | COSMIC | — | — | LoF | — | — | COSV108055283 |
| 197 | Q→H | missense_variant | COSMIC | — | — | damaging | likely_pathogenic (0.97) | -11.31 | COSV99448490 |
| 197 | Q→* | stop_gained | COSMIC | — | — | LoF | — | — | COSV99447996 |
| 198 | Y→* | stop_gained | COSMIC | — | — | LoF | — | — | COSV108055273 |
| 199 | L→L | synonymous_variant | gnomAD | — | 6.84e-07 | — | — | 0.00 | chr19-58556322-G-A |
| 199 | L→L | synonymous_variant | COSMIC | — | — | — | — | 0.00 | COSV53396965 |
| 202 | E→E | synonymous_variant | gnomAD | — | 6.84e-07 | — | — | 0.00 | chr19-58556227-C-T |
| 202 | E→K | missense_variant | gnomAD | — | 6.84e-07 | damaging | likely_pathogenic (0.99) | -10.94 | chr19-58556229-C-T |
| 202 | E→* | stop_gained | ClinVar | — | — | LoF | — | — | ClinVar:4446103 |
| 202 | E→Q | missense_variant | COSMIC | — | — | damaging | likely_pathogenic (0.90) | -11.12 | COSV53396619 |
| 203 | P→P | synonymous_variant | gnomAD | — | 1.37e-06 | — | — | 0.00 | chr19-58556224-G-A |
| 204 | N→N | synonymous_variant | COSMIC | — | — | — | — | 0.00 | COSV53397780 |
| 205 | P→P | synonymous_variant | gnomAD | — | 8.26e-03 | — | — | 0.00 | chr19-58556218-G-A |
| 205 | P→P | synonymous_variant | gnomAD | — | 2.74e-06 | — | — | 0.00 | chr19-58556218-G-C |
| 206 | E→* | stop_gained | gnomAD | — | 6.84e-07 | LoF | — | — | chr19-58556217-C-A |
| 206 | E→K | missense_variant | gnomAD | — | 1.37e-06 | damaging | likely_pathogenic (0.91) | -9.81 | chr19-58556217-C-T |
| 208 | P→P | synonymous_variant | gnomAD | — | 6.84e-07 | — | — | 0.00 | chr19-58556209-T-G |
| 209 | L→L | synonymous_variant | gnomAD | — | 6.84e-07 | — | — | 0.00 | chr19-58556206-C-T |
| 210 | N→N | synonymous_variant | gnomAD | — | 4.10e-06 | — | — | 0.00 | chr19-58556203-G-A |
| 211 | K→K | synonymous_variant | gnomAD | — | 6.84e-07 | — | — | 0.00 | chr19-58556200-C-T |
| 211 | K→T | missense_variant | gnomAD | — | 6.84e-07 | damaging | likely_pathogenic (0.95) | -9.50 | chr19-58556201-T-G |
| 212 | E→E | synonymous_variant | gnomAD | — | 6.84e-07 | — | — | 0.00 | chr19-58556197-C-T |
| 212 | E→D | missense_variant | COSMIC | — | — | — | likely_benign (0.13) | -4.80 | COSV53398288 |
| 213 | A→A | synonymous_variant | gnomAD | — | 2.46e-05 | — | — | 0.00 | chr19-58556194-G-A |
| 214 | A→T | missense_variant | gnomAD | — | 6.84e-07 | damaging | likely_pathogenic (0.99) | -9.75 | chr19-58556193-C-T |
| 215 | E→D | missense_variant | gnomAD | — | 7.52e-06 | — | likely_benign (0.11) | -5.28 | chr19-58556188-C-A |
| 215 | E→K | missense_variant | COSMIC | — | — | damaging | likely_pathogenic (0.64) | -10.56 | COSV53397401 |
| 216 | V→V | synonymous_variant | gnomAD | — | 2.05e-06 | — | — | 0.00 | chr19-58556185-G-A |
| 216 | V→V | synonymous_variant | gnomAD | — | 1.37e-06 | — | — | 0.00 | chr19-58556185-G-T |
| 216 | V→F | missense_variant | gnomAD | — | 1.71e-05 | damaging | likely_pathogenic (0.75) | -10.36 | chr19-58556187-C-A |
| 217 | L→L | synonymous_variant | gnomAD | — | 1.37e-06 | — | — | 0.00 | chr19-58556182-C-G |
| 218 | Q→R | missense_variant | gnomAD | — | 2.74e-06 | — | likely_benign (0.24) | -5.49 | chr19-58556180-T-C |
| 218 | Q→* | stop_gained | COSMIC | — | — | LoF | — | — | COSV53398029 |
| 219 | N→S | missense_variant | gnomAD | — | 6.84e-07 | — | likely_benign (0.07) | -1.12 | chr19-58556177-T-C |
| 220 | N→S | missense_variant | gnomAD | — | 6.84e-07 | — | likely_benign (0.17) | -5.62 | chr19-58556174-T-C |
| 220 | N→D | missense_variant | gnomAD | — | 6.84e-07 | damaging | likely_benign (0.26) | -7.56 | chr19-58556175-T-C |
| 221 | R→Q | missense_variant | gnomAD | — | 1.37e-06 | — | ambiguous (0.49) | -7.12 | chr19-58556171-C-T |
| 221 | R→Q | missense_variant | COSMIC | — | — | — | ambiguous (0.49) | -7.12 | COSV53398017 |
| 221 | R→W | missense_variant | COSMIC | — | — | damaging | likely_pathogenic (0.83) | -8.50 | COSV53397113 |
| 222 | R→R | synonymous_variant | gnomAD | — | 6.84e-07 | — | — | 0.00 | chr19-58556167-C-T |
| 222 | R→Q | missense_variant | gnomAD | — | 2.74e-06 | damaging | likely_benign (0.30) | -8.62 | chr19-58556168-C-T |
| 222 | R→W | missense_variant | gnomAD | — | 1.37e-06 | damaging | likely_pathogenic (0.68) | -8.37 | chr19-58556169-G-A |
| 222 | R→W | missense_variant | COSMIC | — | — | damaging | likely_pathogenic (0.68) | -8.37 | COSV53398430 |
| 223 | L→M | missense_variant | COSMIC | — | — | — | likely_benign (0.11) | -4.24 | COSV99448365 |
| 224 | F→Y | missense_variant | gnomAD | — | 6.84e-07 | damaging | likely_pathogenic (0.76) | -8.37 | chr19-58556162-A-T |
| 224 | F→L | missense_variant | COSMIC | — | — | damaging | likely_pathogenic (1.00) | -9.00 | COSV53397334 |
| 225 | E→K | missense_variant | COSMIC | — | — | damaging | likely_pathogenic (0.87) | -10.99 | COSV99448026 |
| 225 | E→Q | missense_variant | COSMIC | — | — | — | ambiguous (0.39) | -7.06 | COSV99447931 |
| 226 | Q→Q | synonymous_variant | gnomAD | — | 3.42e-06 | — | — | 0.00 | chr19-58556155-C-T |
| 226 | Q→L | missense_variant | gnomAD | — | 3.42e-06 | — | likely_benign (0.21) | -7.43 | chr19-58556156-T-A |
| 226 | Q→H | missense_variant | COSMIC | — | — | — | ambiguous (0.40) | -6.68 | COSV99447946 |
| 226 | Q→L | missense_variant | COSMIC | — | — | — | likely_benign (0.21) | -7.43 | COSV53397103 |
| 227 | N→N | synonymous_variant | gnomAD | — | 3.41e-04 | — | — | 0.00 | chr19-58556152-G-A |
| 228 | V→V | synonymous_variant | gnomAD | — | 6.84e-07 | — | — | 0.00 | chr19-58556149-C-T |
| 229 | Q→H | missense_variant | gnomAD | — | 4.52e-05 | — | likely_benign (0.15) | -4.12 | chr19-58556146-C-A |
| 229 | Q→Q | synonymous_variant | gnomAD | — | 2.05e-06 | — | — | 0.00 | chr19-58556146-C-T |
| 229 | Q→H | missense_variant | ClinVar | Uncertain significance | — | — | likely_benign (0.15) | -4.12 | ClinVar:2512439 |
| 230 | R→R | synonymous_variant | gnomAD | — | 1.37e-06 | — | — | 0.00 | chr19-58556143-G-A |
| 230 | R→H | missense_variant | gnomAD | — | 1.37e-06 | damaging | ambiguous (0.40) | -9.21 | chr19-58556144-C-T |
| 230 | R→H | missense_variant | ClinVar | Uncertain significance | — | damaging | ambiguous (0.40) | -9.21 | ClinVar:3976736 |
| 230 | R→H | missense_variant | COSMIC | — | — | damaging | ambiguous (0.40) | -9.21 | COSV53397913 |
| 230 | R→G | missense_variant | COSMIC | — | — | damaging | likely_pathogenic (0.73) | -11.58 | COSV53397160 |
| 231 | S→F | missense_variant | gnomAD | — | 6.84e-07 | damaging | likely_pathogenic (0.99) | -10.50 | chr19-58556141-G-A |
| 232 | M→T | missense_variant | gnomAD | — | 6.84e-07 | damaging | likely_pathogenic (0.95) | -7.23 | chr19-58556138-A-G |
| 232 | M→V | missense_variant | gnomAD | — | 6.84e-07 | — | likely_benign (0.30) | -5.77 | chr19-58556139-T-C |
| 232 | M→L | missense_variant | gnomAD | — | 6.84e-07 | — | likely_benign (0.16) | -2.62 | chr19-58556139-T-G |
| 233 | R→P | missense_variant | gnomAD | — | 6.84e-07 | damaging | likely_pathogenic (0.99) | -13.93 | chr19-58556135-C-G |
| 233 | R→Q | missense_variant | gnomAD | — | 6.84e-07 | damaging | likely_benign (0.22) | -8.25 | chr19-58556135-C-T |
| 233 | R→W | missense_variant | gnomAD | — | 2.05e-06 | damaging | likely_pathogenic (0.68) | -11.12 | chr19-58556136-G-A |
| 233 | R→G | missense_variant | gnomAD | — | 6.84e-07 | damaging | ambiguous (0.52) | -9.56 | chr19-58556136-G-C |
| 233 | R→R | synonymous_variant | gnomAD | — | 6.84e-07 | — | — | 0.00 | chr19-58556136-G-T |
| 233 | R→W | missense_variant | COSMIC | — | — | damaging | likely_pathogenic (0.68) | -11.12 | COSV53397232 |
| 234 | G→G | synonymous_variant | gnomAD | — | 9.17e-05 | — | — | 0.00 | chr19-58556131-A-G |
| 234 | G→V | missense_variant | gnomAD | — | 6.84e-07 | damaging | likely_pathogenic (1.00) | -10.44 | chr19-58556132-C-A |
| 234 | G→V | missense_variant | COSMIC | — | — | damaging | likely_pathogenic (1.00) | -10.44 | COSV53396668 |
| 235 | G→A | missense_variant | gnomAD | — | 6.84e-07 | damaging | likely_pathogenic (0.60) | -6.74 | chr19-58556129-C-G |
| 235 | G→S | missense_variant | gnomAD | — | 6.84e-07 | — | ambiguous (0.45) | -5.80 | chr19-58556130-C-T |
| 235 | G→V | missense_variant | COSMIC | — | — | damaging | likely_pathogenic (0.93) | -7.87 | COSV53397000 |
| 235 | G→D | missense_variant | COSMIC | — | — | damaging | likely_pathogenic (0.97) | -10.05 | COSV53397761 |
| 236 | Y→Y | synonymous_variant | gnomAD | — | 9.58e-06 | — | — | 0.00 | chr19-58556125-G-A |
| 237 | I→I | synonymous_variant | gnomAD | — | 1.23e-05 | — | — | 0.00 | chr19-58556122-G-A |
| 237 | I→M | missense_variant | gnomAD | — | 2.74e-06 | — | likely_benign (0.30) | -5.11 | chr19-58556122-G-C |
| 237 | I→I | synonymous_variant | gnomAD | — | 1.37e-06 | — | — | 0.00 | chr19-58556122-G-T |
| 237 | I→V | missense_variant | gnomAD | — | 6.85e-06 | — | likely_benign (0.06) | -0.06 | chr19-58556124-T-C |
| 237 | I→V | missense_variant | ClinVar | Uncertain significance | — | — | likely_benign (0.06) | -0.06 | ClinVar:4601252 |
| 238 | G→A | missense_variant | gnomAD | — | 2.05e-06 | — | ambiguous (0.47) | -7.12 | chr19-58556120-C-G |
| 238 | G→S | missense_variant | gnomAD | — | 1.37e-06 | — | likely_benign (0.31) | -5.56 | chr19-58556121-C-T |
| 238 | G→D | missense_variant | COSMIC | — | — | damaging | likely_pathogenic (0.76) | -7.22 | COSV99448418 |
| 239 | S→C | missense_variant | COSMIC | — | — | damaging | likely_benign (0.20) | -9.38 | COSV99448201 |
| 240 | T→N | missense_variant | gnomAD | — | 2.74e-06 | — | likely_benign (0.15) | -5.77 | chr19-58556114-G-T |
| 240 | T→P | missense_variant | COSMIC | — | — | damaging | likely_pathogenic (0.86) | -12.71 | COSV99447856 |
| 242 | F→F | synonymous_variant | gnomAD | — | 6.85e-07 | — | — | 0.00 | chr19-58556107-A-G |
| 243 | E→D | missense_variant | gnomAD | — | 3.43e-06 | — | likely_benign (0.08) | -4.14 | chr19-58556104-C-G |
| 243 | E→E | synonymous_variant | gnomAD | — | 2.06e-06 | — | — | 0.00 | chr19-58556104-C-T |
| 244 | R→H | missense_variant | gnomAD | — | 2.06e-06 | — | ambiguous (0.38) | -7.50 | chr19-58556102-C-T |
| 244 | R→C | missense_variant | gnomAD | — | 6.86e-07 | damaging | likely_pathogenic (0.60) | -8.18 | chr19-58556103-G-A |
| 244 | R→C | missense_variant | COSMIC | — | — | damaging | likely_pathogenic (0.60) | -8.18 | COSV99448464 |
| 246 | L→L | synonymous_variant | gnomAD | — | 1.72e-05 | — | — | 0.00 | chr19-58556095-C-G |
| 247 | K→R | missense_variant | gnomAD | — | 6.87e-07 | — | likely_benign (0.06) | -5.43 | chr19-58556093-T-C |
| 247 | K→E | missense_variant | gnomAD | — | 6.87e-07 | damaging | ambiguous (0.38) | -8.87 | chr19-58556094-T-C |
| 247 | K→N | missense_variant | COSMIC | — | — | damaging | likely_pathogenic (0.67) | -7.02 | COSV53397039 |
328 variants in the shared canonical core.
LoF = frameshift / stop-gain / splice-disrupting — inherently loss-of-function, flagged by consequence (AlphaMissense and ESM-C score only missense/substitutions, so they are blank here by design, not by absence of impact). AlphaMissense is computed in the canonical reading frame, so it scores the shared core but reads N/A across an isoform-unique extension — use ESM-C ΔLLR there.