UBE2D2
EXTENDED 178 aa (canonical 147 aa) · UniProt P62837 · CDLMPS
chr5:139561698:+:ATC:ENST00000398733.8
AI summary N-terminal extension folds confidently alone but sits structurally disconnected from the UBC catalytic core, with no domain, localization, or targeting change.
The added 32-aa N-terminus is highly disordered and proline/arginine-rich (whole-protein disorder shift, S2) yet also folds with high per-residue pLDDT in isolation — but PAE between the extension and the shared UBC body is very high (~18 Å) with only a handful of contacts confined to the junction, meaning it does not integrate as a real structural addition to the catalytic domain. No InterPro domain is gained or lost, DeepLoc/SignalP/TargetP calls are unchanged, and the shared UBC-like fold (which starts downstream of the extension) is essentially untouched (RMSD 0.23 Å at high pTM).
UBE2D2's known function depends entirely on its compact UBC catalytic core (E1 charging, backside Ub engagement, RING/HECT E3 cooperation) and its cytosolic/endosomal localization — none of which this extension's mechanism findings disturb: the core fold, domain content, and predicted compartment are all preserved. The extension reads as a disordered, poorly-integrated N-terminal appendage rather than a functional module, so on the mechanism evidence alone it has no clear bearing on UBE2D2's established E2-conjugase role.
No mass-spec peptide validation of the extension, and the P1 'structured' pLDDT reading is undercut by very high inter-region PAE and minimal contacts, so the biological reality of this extension as a folded or functional appendage remains uncertain.
Folding
Evidence — click any tile for the differential-region detail
LLM reasoning
How well is the isoform-unique region conserved at the amino-acid level across primates (mean percent identity)? High AA identity in close relatives argues the alternative protein is real and translated, not a sequencing or annotation artifact. (Reading-frame intactness is shown as context.)
Comparison
Sequence conservation · primate
mean AA identity over the differential ORF is the score basis; frame intactness across the clade is context
| Conservation metric | Canonical ORF | Differential ORF |
|---|---|---|
| Mean AA identity | 100% | 97% |
| Frame intact (fraction of species) | 88% | 87% |
| Species aligned | 25 | 23 |
| Species frame-intact | 22 | 20 |
| Start codon conserved | 96% | 96% |
| Deepest intact species | Microcebus_murinus | Microcebus_murinus |
| Phylo depth (MRCA) | 7 | 7 |
The same amino-acid identity test across mammals. Conservation over deeper evolutionary distance is stronger evidence the ORF is under selection to be translated. (Reading-frame intactness is shown as context.)
Comparison
Sequence conservation · mammalian
mean AA identity over the differential ORF is the score basis; frame intactness across the clade is context
| Conservation metric | Canonical ORF | Differential ORF |
|---|---|---|
| Mean AA identity | 100% | 82% |
| Frame intact (fraction of species) | 75% | 44% |
| Species aligned | 20 | 18 |
| Species frame-intact | 15 | 8 |
| Start codon conserved | 94% | 71% |
| Deepest intact species | Loxodonta_africana | Canis_lupus_familiaris |
| Phylo depth (MRCA) | 12 | 11 |
phyloP / phastCons measure per-base evolutionary constraint. A high absolute phyloP over the differential region means the sequence itself is under strong purifying selection — evidence it is coding and does something. (Enrichment over the shared core is shown as context only.)
Comparison
Per-base conservation · differential region
absolute mean phyloP over the differential region is the score basis (high = strong purifying selection); the shared column and enrichment are context, not the claim
| Track | Differential | Shared | vs shared |
|---|---|---|---|
| phyloP mean | 3.82 | 5.77 | 0.663 |
| phastCons mean | 0.838 | 0.976 | — |
Start site · canonical vs isoform
initiation context + per-base conservation at each start codon
| Property | Canonical | Isoform |
|---|---|---|
| Start codon | ATG | ATC |
| Kozak context (−9..+4) | CCTCCCACCATGG | CCGACCGAGATCG |
| phyloP at start codon | 6.79 | 1.83 |
| phastCons at start codon | 1 | 0.03 |
| phyloP over Kozak window | 4.77 | 1.36 |
| phastCons over Kozak window | 0.858 | 0.162 |
| Kozak mismatch — full consensus | 3 | 6 |
| Kozak window GC content | 0.692 | 0.692 |
LLM reasoning
Is the alternative start used in more than one cell line? Reproducible initiation across independent samples argues against a one-off ribosome artifact.
Comparison
Per-cell-line usage · canonical vs isoform
Fisher q = 2.3e-16
| Cell line | Canonical | Isoform | p-value |
|---|---|---|---|
| HeLa | 27.5 | 12.4 | 1.56e-09 |
| K562 | 9.88 | 8.07 | 2.3e-16 |
| U2OS | 2.8 | 2.14 | 1.09e-08 |
| RPE1 Async | 2.96 | 2.03 | 6.63e-07 |
| RPE1 Que | 1.71 | 0.947 | 2.74e-06 |
How efficiently ribosomes initiate at this start (TIS) relative to background, per cell line. Strong, reproducible initiation supports a genuine translation event.
Comparison
Start-Site Usage · canonical vs isoform
ribosome initiation efficiency at the canonical start vs this alternative start, per cell line
| Cell line | Canonical | Isoform |
|---|---|---|
| HeLa | 0.382 | 0.172 |
| K562 | 0.113 | 0.0926 |
| U2OS | 0.0429 | 0.0328 |
| RPE1 Async | 0.0371 | 0.0254 |
| RPE1 Que | 0.0343 | 0.019 |
Were tryptic peptides unique to the differential region detected by mass spec? Direct peptide evidence is the strongest proof the isoform protein exists.
Comparison
Peptide Evidence (canonical vs isoform)
isoform-unique peptides are direct evidence the alternative protein exists
| Feature | Canonical | Isoform |
|---|---|---|
| Tryptic peptides (in-silico) | 10 | 7 |
| Validated by mass-spec | 0 | 0 |
| Isoform-unique peptides | — | 7 |
Details
Peptide Evidence (canonical vs isoform)
- peptide MAEAASGSLAPSPSLPRPRPRPGGR 0–25
- peptide AEAASGSLAPSPSLPRPRPRPGGR 1–25
- peptide HPPPTMALK 26–35
- peptide MAEAASGSLAPSPSLPRPRPRPGGRR 0–26
- peptide AEAASGSLAPSPSLPRPRPRPGGRR 1–26
- peptide RHPPPTMALK 25–35
- peptide HPPPTMALKR 26–36
LLM reasoning
Do the predicted localization features (DeepLoc prediction, sorting signals, or membrane association) differ between the canonical and the isoform? A protein with changed localization features acts in a different cellular context.
Comparison
Subcellular localization (DeepLoc)
localization features changed (prediction/signals/membrane)
| Property | Canonical | Isoform |
|---|---|---|
| Predicted location | Cytoplasm|Nucleus | Cytoplasm|Nucleus |
| Sorting signals | Nuclear localization signal|Nuclear export signal | Nuclear localization signal|Nuclear export signal |
| Membrane | Peripheral|Soluble | Soluble |
Do N-terminal targeting signals (SignalP secretion, TargetP mitochondrial/chloroplast) differ between canonical and isoform? N-terminal changes most directly add or remove targeting peptides.
LLM reasoning
Does healthy human germline variation (gnomAD) avoid this region (depletion ratio < 1×), and is it intrinsically constrained (ESM-C constraint delta > 0)? Depletion of population variation plus high sequence constraint mean the region resists change — it is functionally important. Scored only where the unique region is canonical coding sequence, i.e. on truncations. gnomAD is a tolerance catalogue, not a disease one; disease/cancer variants (ClinVar / COSMIC) live in M2.
Comparison
Germline Variants · differential vs shared region
gnomAD (population) variant density per nucleotide — depletion ratio < 1× means healthy human variation avoids the region (constrained), the M1 basis alongside ESM-C constraint
| Variant set | Differential | Shared | Depletion ratio |
|---|---|---|---|
| gnomAD variants | 136 | 109 | 5.9× |
Sequence constraint (ESM-C) · differential vs shared region
lower LLR = more constrained; enrichment = differential / conserved
| Property | Differential | Shared | Enrichment |
|---|---|---|---|
| ESM-C mean LLR | -1.92 | -0.0173 | — |
| Constrained positions | 0 | 0 | — |
Predicted-damaging variants · germline (gnomAD)
AlphaMissense / ESM-C / LoF calls per region (length-normalized)
| Variant set | Differential | Shared | Enrichment |
|---|---|---|---|
| Scorable variants | 131 | 109 | 5.7× |
| Damaging variants | 30 | 31 | 4.6× |
| — of which loss-of-function | 30 | 2 | 71× |
| AlphaMissense-pathogenic | 0 | 21 | 0× |
Predictor scores · germline (gnomAD)
scored: 265 ESM-C · 89 AlphaMissense
| Property | Differential | Shared |
|---|---|---|
| Mean ΔLLR (ESM-C) | -0.681 | -3.36 |
| Min ΔLLR (ESM-C) | -2.97 | -11.4 |
| Mean AlphaMissense | — | 0.56 |
Are disease (ClinVar / COSMIC) variants enriched per nucleotide in the differential region versus the shared core (disease enrichment ratio ≥ 1×)? Disease variants concentrating in the unique region tie it to phenotype.
Comparison
Clinical-variant burden · differential vs shared region
counts per region; ratio is density-normalized (variants per nucleotide, differential ÷ shared — ≥1× = disease variants concentrate in the differential region, the M2 basis)
| Variant set | Differential | Shared | Enrichment |
|---|---|---|---|
| Disease variants | 3 | 58 | 0.25× |
| Pathogenic | 0 | 0 | — |
Predicted-damaging variants · disease (ClinVar/COSMIC)
AlphaMissense / ESM-C / LoF calls per region (length-normalized)
| Variant set | Differential | Shared | Enrichment |
|---|---|---|---|
| Scorable variants | 3 | 58 | 0.25× |
| Damaging variants | 1 | 41 | 0.12× |
| — of which loss-of-function | 1 | 3 | 1.6× |
| AlphaMissense-pathogenic | 0 | 34 | 0× |
Predictor scores · disease (ClinVar/COSMIC)
scored: 265 ESM-C · 89 AlphaMissense
| Property | Differential | Shared |
|---|---|---|
| Mean ΔLLR (ESM-C) | -0.621 | -7.34 |
| Min ΔLLR (ESM-C) | -1.24 | -11.9 |
| Mean AlphaMissense | — | 0.787 |
LLM reasoning
Does the gained/lost region fold into ordered structure (pLDDT) and shift the protein's biophysical character (charge, hydropathy, disorder)? Structured, biophysically distinct regions are more likely to be functional.
Comparison
Structure (ESMFold2) · canonical vs isoform
TM-score 0.911 · RMSD 0.231 Å · 7 interface contacts
| Metric | Canonical | Isoform |
|---|---|---|
| pLDDT (whole protein) | 0.885 | 0.961 |
Fold confidence · differential vs shared region
mean ESMFold2 pLDDT in each region
| Metric | Differential | Shared | Enrichment |
|---|---|---|---|
| pLDDT | 0.909 | 0.823 | 1.1 |
The shared region is the stretch of protein identical in both the isoform and the canonical (the canonical body for an extension; the post-truncation body for a truncation). Folded in both contexts it normally comes out nearly identical, so its Cα backbone RMSD ≈ 0. A high shared-region RMSD (superposed on the shared residues only, from the ESMFold2 structures) means the extension or truncation reorganizes how the retained region folds — a rare, high-interest functional signal. TM-score is a length-normalized companion; the check only fires when both structures are confidently folded, and uORF/altORF isoforms (no shared region) are not evaluable.
Comparison
Shared-region structural change
Cα RMSD superposed on the shared residues only; TM-score is length-normalized · shared-region Cα RMSD 0.231 Å · shared TM-score 0.997 · shared region 147 aa · min shared pLDDT 0.885 · global TM-score 0.911 · global RMSD 0.231 Å
| Property | Canonical | Isoform |
|---|---|---|
| Shared-region pLDDT | 0.885 | 0.971 |
Does the differential region contain actual secondary structure — a helix or strand — rather than coil? ESMFold2 predicts coordinates but no secondary structure, so elements are assigned from those coordinates (P-SEA, Cα geometry). Because they are derived from a prediction, each element carries its own mean pLDDT: a geometrically clean helix running through a disordered stretch is geometry fitted to a guess, so BOTH length and confidence are required to score. The direction differs by ORF type — an extension GAINS the element (a candidate functional addition), a truncation LOSES one, and there the element is read off the canonical structure because the removed segment exists only in it. This says nothing about whether the element is integrated with the rest of the fold; the contact and PAE evidence in this same category answers that.
Comparison
Secondary structure · differential vs shared region
helices and strands assigned from the predicted coordinates (P-SEA)
| Metric | Differential | Shared |
|---|---|---|
| Alpha helices | 0 | 4 |
| Beta strands | 0 | 2 |
| Longest element (aa) | 0 | 15 |
| Mean pLDDT | — | 0.97 |
Elements and coordinates
0 in the differential region, 6 in the shared core — residue numbering is 1-based on the protein holding the region
| Isoform-unique | Shared core |
|---|---|
| — | alpha helix 32–46 15 aa · pLDDT 0.96 |
| — | beta strand 63–68 6 aa · pLDDT 0.96 |
| — | beta strand 79–87 9 aa · pLDDT 0.97 |
| — | alpha helix 130–142 13 aa · pLDDT 0.98 |
| — | alpha helix 152–159 8 aa · pLDDT 0.97 |
| — | alpha helix 162–174 13 aa · pLDDT 0.97 |
Below threshold
1 in the differential region, 2 in the shared core — shorter than 6 aa or below pLDDT 0.70, so not counted above
| Isoform-unique | Shared core |
|---|---|
| beta strand 2–6 5 aa · pLDDT 0.86 | beta strand 52–55 4 aa · pLDDT 0.98 |
| — | beta strand 97–100 4 aa · pLDDT 0.98 |
LLM reasoning
Does the isoform gain or lose a real InterPro functional domain in the differential region (disorder/structural-only signatures excluded)? Gaining or losing a domain changes function directly.
Comparison
Domains & motifs (canonical vs isoform)
gained = only in the isoform; lost = only in the canonical
| Feature | Canonical | Isoform |
|---|---|---|
| InterPro domains | 9 | 10 |
| Short linear motifs | 3 | 3 |
Details
Domains & motifs (canonical vs isoform)
- gained Consensus disorder prediction domain
Biophysical character of the isoform-differential region versus the shared canonical core — pI, hydropathy, charge, disorder and related properties. Descriptive; the folding (P1) score keys off the GRAVY / charge / disorder deltas.
Comparison
Biophysics · differential vs shared region
highlighted rows are enriched in the differential region
| Property | Differential | Shared | Enrichment |
|---|---|---|---|
| Isoelectric point (pI) | 12.7 | 7.73 | 1.64 |
| Hydropathy (GRAVY) | -1.03 | -0.348 | 2.96 |
| Fraction charged | 0.226 | 0.238 | 0.948 |
| Disorder fraction | 0.385 | 0.0821 | 4.69 |
| Disorder-promoting | 0.839 | 0.517 | 1.62 |
| Low-complexity fraction | 0.903 | 0 | — |
| Prion-like fraction | 0.226 | 0.225 | 1.01 |
| LLPS score | 0.327 | 0.13 | 2.52 |
| π–π propensity | 0.194 | 0.265 | 0.729 |
| Aromaticity | 0 | 0.102 | 0 |
| Instability index | 104 | 42.9 | 2.43 |
| Shannon entropy | 2.93 | 4.19 | 0.698 |
| Normalized complexity | 0.677 | 0.969 | 0.698 |
Sparse-autoencoder (SAE) interpretability features on the ESM-C residual stream, comparing the isoform against the canonical protein. This is the S3 criterion — a presence check that fires when any interpretable feature is gained or lost. Only the top 30 features per category (ranked by activation, prevalence ≥ 2) are saved; each table shows the top 10 interpretable features, and generic / "unknown" features are hidden unless you expand it. What are SAE features? →
Part 1 · Global feature comparison
Which SAE features fire in the isoform vs the canonical protein, across the whole sequence.
Isoform-only features — 153 total (prevalence ≥ 2)
| Feature | Label | Description | Activation | Prevalence |
|---|---|---|---|---|
| #448 | N-terminal leader/targeting segments | N-terminal leader/targeting segments: the feature activates on the extreme N-terminus (first 2–5 residues) of proteins, especially the N-region of signal peptides and organelle transit peptides, and more generally on short, disordered/low-complexity N-terminal tails; strongest at residue ~3 and largely independent of amino-acid identity, occurring across all taxa and functions. | 16.61 | 2 |
| #4992 | Small-residue disordered N-terminus | Context-dependent free N-terminus signature: short, disordered amino-terminal stretches beginning with the initiator Met followed by residues compatible with MAP cleavage or N-acetylation (e.g., M–G/A/S/T/V, and other contexts including M–D/E/L); strongest emphasis on small residues at positions 2–3; activation is absent when the native N-terminus is structured or starts with a signal peptide/transmembrane segment. | 15.46 | 5 |
| #7996 | Generic extreme N-terminus detector | Generic extreme N-terminus detector: marks residues around position 5 (occasionally 6–7) at the very start of the polypeptide, most often within the N-region of signal/transit peptides of secreted or organelle-targeted precursors, but also in generally disordered N-termini of cytosolic/nuclear proteins; not motif- or residue-specific | 10.80 | 2 |
| #3609 | Position-2 N-terminus residue detector | Detector of the identity of the second residue at the extreme N-terminus of proteins, with strongest preference for small residues (especially Ser, Ala, and Gly; also Thr/Pro), i.e., an N-terminal position-2 signal commonly associated with generic, disordered starts and co-translational N-terminus processing | 10.78 | 2 |
| #11504 | Initiator methionine at M1 | Initiator methionine at the very start of the polypeptide chain (M1), i.e., the translation start residue, independent of protein family, taxonomy, membrane association, or presence of signal/leader/propeptide regions; often annotated as post-translationally removed and typically situated in a flexible, coil-like N-terminus. | 10.62 | 2 |
| #5241 | Early N-terminal positional signal | Generic early N-terminus positional signal peaking at residue ~5–7, found broadly within the disordered N-tail of proteins, including both cleavable leader/targeting segments (signal peptides, propeptides) of precursor proteins and disordered N-terminal regions of mature proteins; residue identity is not specific | 9.59 | 3 |
| #7986 | Intrinsically disordered low-complexity regions | Compositionally biased, intrinsically disordered low‑complexity regions used as flexible interaction/nucleic‑acid–binding modules—enriched in Pro/Ser/Lys/Glu/Gln/Thr/Gly/Arg and encompassing poly‑Pro tracts, Lys/Arg‑rich basic stretches, and acidic Asp/Glu clusters—common in transcription/RNA‑processing factors, viral basic regulators, and the C‑terminal scaffolds of RNase E | 6.68 | 22 |
| #7026 | Catalytic domain helix/loop boundary motif | A localized internal-domain motif that marks a specific short helix/loop boundary within a catalytic domain, situated downstream of the catalytic proton-donor residue in α/β hydrolase-like folds. | 6.38 | 3 |
| #4180 | Hinge residues; C-terminal linker motifs | Secondary-structure transition/hinge residues—particularly helix termini and adjacent coil/turn positions—and conserved C-terminal/linker sequence motifs in nucleotide-binding domains | 6.32 | 2 |
| #6178 | Nuclear basic/polar low-complexity IDRs | Intrinsically disordered, low‑complexity regulatory segments of nuclear proteins—especially transcription factors/cofactors, chromatin regulators, RNA/splicing factors, and nuclear RING‑type E3 ligases—enriched in basic (Lys/Arg) and polar (Ser/Pro/Gln/Glu/Gly/Thr) tracts that serve as activation/repression, partner‑binding, or localization (NLS‑like) regions; typically outside folded DNA‑binding/catalytic domains and often overlapping phosphorylation sites or simple helices within IDRs. | 6.05 | 7 |
| #6484 | Prion-like low-complexity IDRs | Polar low-complexity intrinsically disordered regions (IDRs)—especially Q/N-rich (prion-like) segments and S/T- and glycine-rich tracts with simple SG/TG/PG repeats—typically located in inter-domain linkers and terminal tails of eukaryotic proteins and often harboring Ser/Thr phosphorylation sites. | 5.53 | 6 |
| #11753 | Cytosolic N-terminal regulatory tails | Cytosolic N-terminal tails and disordered/low-complexity regulatory segments at the extreme N-terminus of eukaryotic signaling proteins, including phosphatases (classical PTPs, PP2C family) and membrane receptors/ion channels. Sequence character is mixed: some tails are Ser/Gly-rich or basic/acidic low-complexity stretches, others are short hydrophobic/charged N-terminal helices preceding a catalytic domain. The feature avoids transmembrane cores, catalytic/active-site loops, metal-binding residues, and well-folded globular interiors. | 5.04 | 34 |
| #5482 | N-terminal targeting presequences | Universal eukaryotic N-terminal targeting presequences: the feature detects short, cleavable leader regions at the extreme N-terminus that direct proteins to organelles or the secretory pathway—especially chloroplast/apicoplast transit peptides and thylakoid lumen signals, but also mitochondrial targeting peptides and classical signal peptides. These segments are Ser/Thr- and small/hydrophobic–rich, enriched in Lys/Arg and depleted of acidic residues, typically low-structure/low-confidence and ending at the maturation cleavage site. | 4.85 | 2 |
| #3398 | Ser/Pro-rich disordered tails | Serine/proline–rich low-complexity intrinsically disordered segments, especially terminal tails, linkers, and N‑terminal signal-peptide regions of small or secreted proteins; structured catalytic cores are largely ignored. | 4.82 | 5 |
| #4012 | Proline-directed IDR phosphorylation | Intrinsically disordered Ser/Thr phosphorylation hotspots, with a strong preference for proline‑directed motifs (S/T‑P) characteristic of CDK/MAPK-like kinase targets across diverse eukaryotic regulators and viral phosphoproteins. | 4.77 | 3 |
| #9145 | Helix–loop boundary/linker signal | A generic helix–loop boundary/linker signal: the feature prefers short coil segments and helix-capping/hinge positions that connect or terminate alpha-helices, especially interhelical loops in multi-pass membrane proteins and structured active-site/cofactor-binding loops in soluble enzymes; it tolerates diverse amino acids but often includes small helix-modulating or interfacial aromatic residues. | 4.73 | 2 |
| #12620 | Disordered low-complexity regions | Intrinsically disordered, low‑complexity, compositionally biased regions/tails (IDRs), typically enriched in Ser/Gly/Pro/Ala/Thr and often occurring as acidic (Asp/Glu) or basic (Arg/Lys) tracts or Gln/Asn‑/Gln‑rich repeats; these segments are common in secreted precursors, viral proteins, micropeptides, and testis‑associated proteins, and can also occur as low‑complexity termini or surface loops appended to otherwise folded enzymes; they frequently coincide with low predicted structural confidence. | 4.67 | 9 |
| #12818 | N-terminal targeting and disorder | N-terminal targeting and processing segments of secreted/endomembrane and organelle-targeted proteins—cleavable signal peptides, N-terminal signal‑anchor helices, and mitochondrial/chloroplast transit peptides—and the immediately downstream low‑complexity/disordered propeptide or luminal/matrix-facing stem; more generally, a preference for intrinsically disordered N‑terminal tails (including Ser/Thr/Ala‑rich plant segments), with occasional hits on short flexible N-termini of soluble cytosolic proteins. | 4.66 | 29 |
| #2530 | LTV1/MPP10 motifs, TruD helix | Sequence-specific motif feature active in ribosome/snoRNP biogenesis factors and TruD-family pseudouridine synthases. In LTV1 homologs the feature recognizes a conserved internal "SSVxRRNEQL"-like motif, and in MPP10 homologs it recognizes a conserved "P(A/V)PVITEE" motif. In TruD/PUS7-family proteins it activates along an internal alpha-helix within the TRUD catalytic domain. | 4.66 | 3 |
| #7338 | Amphipathic coiled-coil helices | Coiled-coil–like amphipathic alpha-helices with heptad-repeat character (hydrophobic a/d layers often L/A and charged/polar e/g positions enriched in E/Q/K/R/S/T), as found in long dimeric/oligomeric scaffolds, adaptors, and motor tails; the signal also extends to shorter amphipathic helices in peptide precursors. | 4.51 | 10 |
| #2038 | Disordered low-complexity regulatory regions | Low-complexity, intrinsically disordered regulatory regions enriched for serine/threonine and glutamine/asparagine (often with glycine/proline/histidine runs), typically located in N- or C-terminal tails and flexible linkers of eukaryotic proteins; these segments include transcriptional activation/repression regions and phosphorylation-modulated interaction sites, while structured catalytic or DNA/RNA-binding domains are not targeted. | 4.43 | 19 |
| #11732 | GH18/GH13 catalytic core loops | Conserved sites within the catalytic core of GH18 chitinase-like proteins and GH13 α-glucosidase/maltase-family enzymes, with peaks typically falling on loop/helix regions internal to the catalytic domain rather than on terminal tails. | 4.43 | 3 |
| #13964 | Short amphipathic alpha-helices | Short amphipathic alpha-helical segments enriched in leucine (and other hydrophobes) with interspersed basic residues, most prominently the helices that compose HEAT/ARM-like alpha-solenoid repeats but also analogous isolated N-terminal helices in diverse proteins | 4.38 | 3 |
| #9395 | Short terminal targeting segment | Terminal or short internal targeting/interaction segment: a short, compositionally biased or hydrophobic patch near a protein terminus, or occasionally an internal low‑complexity loop, that functions as a signal peptide or signal‑anchor (Sec/ER/thylakoid/Tat entry) when hydrophobic, or as a basic/Ser/Lys/Arg/Pro‑rich low‑complexity segment that mediates localization or partner binding in soluble proteins; typically the single prominent terminal segment of short secreted, membrane, or regulatory proteins | 4.35 | 4 |
| #5021 | Polybasic intrinsically disordered regions | Low-complexity, often intrinsically disordered regions in short proteins, precursors, and microproteins across taxa, including viral accessory proteins, neuropeptide/hormone precursors, sperm nuclear proteins, flexible enzyme tails/inserts, and antisense-derived/uncharacterized microproteins. Activation is biased toward, but not restricted to, basic (Lys/Arg) clusters, with frequent peaks also at Pro/Ser/Gly residues in disordered context. | 4.11 | 7 |
| #8461 | Ciliary/meiotic charged coiled-coils | Signal for residues within long α-helical coiled-coil segments of cilia/flagella- and meiosis-associated structural proteins, firing on heptad-repeat positions in charged, Glu/Lys/Arg-rich stretches. | 4.08 | 3 |
| #12751 | Ser/Pro-biased low complexity regions | Compositionally biased regions (IDRs or short low-complexity segments) that are explicitly Ser/Pro-biased or enriched for simple dipeptide repeats (SS/PP, RS/SR, RG/RGG) and basic residues; these tracts occur in viral accessory proteins and micropeptides and as Ser/Pro-rich linkers/termini in diverse proteins. Generic hydrophobic signal peptides without such polarity/repeat bias are not targets. | 4.06 | 4 |
| #9962 | S/T/P-rich disordered regulatory tails | Intrinsically disordered, low‑complexity regions enriched in serine, threonine, proline and polar/charged residues—flexible regulatory linkers/tails and propeptide segments that often host short linear motifs (e.g., phosphorylation- and proline‑rich motifs) and proteolytic processing sites; signal is absent from well‑folded catalytic domains and is common across taxa | 4.05 | 5 |
| #11023 | Low-complexity disordered regions | Intrinsically disordered, low‑complexity segments—often N‑terminal tails or leader/signal regions—enriched in Ser/Pro/Gly/Arg and predicted as coils/low‑confidence structure; the feature marks flexible, poorly structured regions rather than a specific function | 3.97 | 4 |
| #11676 | Disordered Ser/Thr phosphorylation sites | Short linear motifs centered on serine/threonine within intrinsically disordered regions that correspond to eukaryotic Ser/Thr phosphorylation sites (often SP/TP or basic R/K-flanked motifs), used for regulatory control across diverse proteins including viral proteins | 3.97 | 4 |
Canonical-only features — 34 total (prevalence ≥ 2)
| Feature | Label | Description | Activation | Prevalence |
|---|---|---|---|---|
| #6362 | Kinase domain helix-capping residues | Residues at helix–coil transition (helix-capping/linker) positions within the conserved protein kinase catalytic domain, especially in the C-lobe around long helices and the activation-segment vicinity; a generic structural marker of kinase cores rather than active-site motifs, conserved across Ser/Thr, Tyr, and dual-specificity kinases (including eukaryote-like bacterial kinases) | 7.42 | 2 |
| #7265 | N-terminal initiator Met recognition | Recognition of the absolute protein N-terminus—specifically the initiator methionine (M1; often N-formylmethionine in bacteria/mitochondria) and the immediate +1 residue that becomes the mature N-terminus upon N-terminal methionine excision | 5.41 | 2 |
| #735 | Aromatic edge beta-strand motifs | Short edge beta-strand segments in beta-sheet–rich domains (especially jelly-roll/beta‑sandwich folds such as Ig-like, cupredoxin, velvet and LPMO), typically at terminal/edge strands or loop–strand junctions, often containing or flanked by aromatic residues (Y/F) and frequently adjacent to a cysteine that participates in a disulfide or metal‑ligand; common in extracellular/secreted proteins but not restricted to them. | 4.49 | 2 |
| #8428 | Noncatalytic beta-strand and loop hotspots | Residue-level signature within compact beta-rich domains, firing at sparse positions within or immediately adjacent to beta-strands and in short connecting loops, generally non-catalytic and common across diverse folds. | 3.77 | 2 |
| #10932 | Globular hydrophobic core packing | Hydrophobic packing within well-folded, soluble α/β domains—i.e., core-facing residues in helices and strands of globular regulatory/signaling domains (kinase catalytic cores, bacterial HTH/sigma DNA‑binding domains, etc.); a generic structural/fold signal rather than a motif- or active‑site–specific feature. | 3.21 | 2 |
| #1785 | RlmK/L dual-motif peaks | Short conserved sequence motifs within bacterial ribosomal RNA large-subunit methyltransferases (RlmK/RlmL family), with peaks on hydrophobic/aromatic residues embedded in two characteristic stretches of the enzyme. | 2.70 | 2 |
| #11070 | Aromatic-polar loop micro-motif | Short polar/aromatic micro-motif centered on H or N typically followed by an aromatic (F/Y) and embedded in coils, turns, or short beta-strands; often located in active-site or ligand-binding loops, prominently adjacent to catalytic/binding residues in N-acetylhexosaminidases (GH20) and chitinases, but reused across diverse enzyme folds and outer-membrane transporters. | 2.65 | 2 |
| #5114 | Short N-terminal targeting arm | N-terminal segments (typically the first ~10 residues) across a wide range of bacterial, viral, and eukaryotic proteins. These are often signal/targeting/assembly arms, including RNA/DNA-binding tails, secretion/chaperone-binding regions, and assembly linkers. | 2.32 | 5 |
| #6919 | Helix caps and turns | Residues at alpha-helix boundaries and short interhelical turns (helix caps/hinges), including the first/last few positions of soluble or transmembrane helices and the brief loops linking paired helices in helical repeat units; enriched for small residues (especially Gly/Ala, also Asn/Thr/Val/Pro) consistent with helix-capping and turn formation | 2.21 | 2 |
| #8838 | P450 heme-binding Cys pocket | Cytochrome P450 heme-binding Cys-pocket and its upstream “meander” loop—i.e., the conserved C‑terminal heme‑ligand region including and immediately preceding the FGxGxxxCxG signature that carries the axial heme‑coordinating cysteine; a fold-level element shared across diverse P450 families and taxa | 2.14 | 2 |
| #15589 | Mid-chain hydrophobic core | Mid‑chain, well‑packed structural core segment of small domains—a contiguous, hydrophobic (Gly/Val/Leu/Ile‑rich) secondary‑structure run that stabilizes the domain (most often a β‑strand block or β‑α‑β unit; in all‑α folds the analogous central helix pack; in membrane proteins a mid‑bundle TM helix), typically adjacent to but not itself the catalytic/ligand‑binding residues. | 2.13 | 3 |
| #12567 | N-terminal amphipathic leader helix | A short, compositionally biased N-terminal segment around positions ~18–40 that is enriched in charged/polar residues (often with glycine runs) and behaves as an intrinsically disordered patch with strong propensity to form a single short amphipathic alpha-helix; this generic “leader-like” interface is used across proteins as a targeting/secretion/assembly or interaction module (rather than a specific sequence motif). | 2.09 | 9 |
| #10622 | Amphipathic helix assembly motifs | Short amphipathic α‑helical interface segments (8–20 aa) enriched in Leu/Ile/Val and acidic residues with interspersed Lys/Arg and nearby aromatics, used as helix–helix packing/assembly or protein–protein interaction motifs. The feature recurs in helical scaffolds of retroviral/retrotransposon Gag/capsid and histone‑fold proteins, and more broadly in other all‑α domains. | 2.05 | 2 |
| #10752 | Short solvent-exposed connector loops | Generic structural signal for short, solvent-exposed loop/turn connectors between secondary structure elements—especially beta-beta hairpin loops and helix-strand junctions—enriched in small/charged residues (S/T, N, D/E, K/R); often adjacent to functional sites but not family-specific | 2.03 | 2 |
| #1888 | AAA+/Lon noncatalytic residue peaks | Residue-level detector active in AAA+ ATPase / Lon-protease family proteins, with discrete peaks on structured non-catalytic positions and minimal activation on catalytic, ATP-binding, or disordered regions. | 2.00 | 2 |
| #2960 | Helix N-cap pocket hotspots | Conserved strand-to-helix beginnings and helix N-cap segments that scaffold or border functional pockets—especially cofactor- and active-site regions—across many enzyme folds (PLP-, ThDP-, SAM-, and flavin-dependent) and, by extension, analogous short helical elements in non-enzymatic RNA/protein-binding proteins; these sites are enriched in small/hydrophobic residues with frequent glycine and acidic (Asp/Glu) capping positions. | 1.90 | 2 |
| #5818 | Short amphipathic helix propensity | Short alpha-helical segment signal that is most prominent in acyl-CoA dehydrogenase family enzymes and a few other helical proteins, capturing local helix propensity and typical amphipathic/hydrophobic helical composition rather than specific catalytic or ligand-binding motifs. | 1.79 | 2 |
| #4608 | BTB/POZ N-terminus beta-strand detector | Residue-level detector of beta-strand structural context in well-folded domains, with a strong preference for the very start of the BTB/POZ domain (the first beta-strand or its immediate N-terminal boundary). Outside BTB proteins it similarly marks isolated beta-strand (and adjacent inter-strand loop) positions within diverse domains rather than catalytic or ligand-coordination sites. | 1.78 | 2 |
| #2363 | Charged alpha-helical tracts | Polar/charged alpha-helical tracts (often low-complexity): E/D/K/R/S/T/N/Q-rich, low-aromatic segments occurring both in intrinsically disordered regions and as solvent-exposed helices within folded domains; these amphipathic/coiled-coil–like helices mediate oligomerization, nucleic-acid/protein binding, and assembly, especially in viral proteins. | 1.76 | 2 |
| #1197 | Exposed beta-loop binding rims | Short, solvent-exposed β-strand–to–loop segments that form ligand-recognition rims of β-rich domains—often glycan-binding in lectin/pentraxin folds and β-propeller CAZymes, but also protein-recognition loops in β-sandwich modules such as SPRY/PRY-SPRY, and analogous exposed loops near processing or substrate-binding sites in secreted virulence proteases—used in extracellular adhesion/innate-immunity proteins, secretory cargo lectins, and selected cytosolic adaptors; frequently adjacent to Ca2+-dependent sites where applicable. | 1.74 | 2 |
| #10995 | β-strand–rich extracellular interfaces | β-strand–rich interaction surfaces with strong enrichment in secreted/lumenal protein regions, but not exclusive to them; the feature targets β-strands and adjacent loops that are often post-translationally stabilized (N-glycans) in extracellular proteins, and analogous β-elements in certain intracellular trafficking and phosphoinositide-processing proteins. | 1.73 | 2 |
| #15784 | OmpR/PhoB signature motifs | Conserved sequence/structural elements of OmpR/PhoB-type response regulator transcription factors, marking signature positions in both the receiver domain and the C-terminal winged HTH DNA-binding domain. | 1.73 | 2 |
| #14603 | Charged amphipathic interaction helices | Charged, amphipathic alpha‑helical interaction segments—often N‑terminal and coiled‑coil‑like—that mediate protein–protein or protein–membrane contacts in small complex subunits and diverse enzymes/membrane proteins across taxa | 1.72 | 2 |
| #2494 | Structured-to-disordered boundary segments | A short, terminal or domain-edge coil/loop segment—most often a C‑terminal tail or a loop-to-helix/strand capping region—that marks the transition from a structured core to a flexible/low‑complexity extension. These segments frequently show reduced structural confidence (low pLDDT), are enriched in charged and glycine residues, and occasionally embed specific motifs (e.g., metal-binding cysteines), but typically lack discrete annotated active sites. | 1.71 | 3 |
| #13576 | Amphipathic recognition helices | A broad detector of short, well-ordered amphipathic alpha-helices that serve as recognition/interaction surfaces (e.g., DNA-binding basic helices, kinase docking helices, coiled-coil edges), rather than catalytic motifs; consistent with a mixed basic/acidic and hydrophobic residue composition and active across diverse taxa and functions. | 1.70 | 2 |
| #7296 | Bromodomains and flanking IDRs | Chromatin reader modules—especially bromodomains—and their immediately adjacent low-complexity, PTM-rich regulatory linkers/tails in nuclear chromatin regulators; the feature fires on the bromodomain core itself and extends into flanking acidic/Ser/Pro/Gly- and Lys/Arg-biased segments, with weaker signal on other reader domains (e.g., MBD and HMG-box) accompanied by similar IDRs; a minority of non-chromatin proteins with analogous acidic/basic low-complexity tracts can activate at lower scores (e.g., IAP E3 ligases, secreted phospholipases). | 1.70 | 2 |
| #7518 | Histidine kinase helix-start recognition | Sparse recognition of residues within the cytosolic domains (PAS, HisKA/DHp, HATPase_c, REC) of two-component histidine kinases and related signal-transduction proteins, with peaks frequently landing on Pro or acidic (Asp/Glu) residues but also on a variety of other residues at helix-start and adjacent coil positions. | 1.67 | 2 |
| #6737 | Cofactor adjacent functional segment | A contiguous, mid-protein "functional segment" used to position or interact with cofactors/ions or partner subunits—a loop/helix/strand adjacent to metal/ion or cofactor-binding sites (Fe–S, di-iron, Mg2+, Cu, SAM) or lying at assembly interfaces in redox/energy-metabolism complexes and related enzymes. | 1.66 | 3 |
| #8849 | Alpha-helical CAZyme scaffolds | Alpha-helical scaffold segments in carbohydrate-active enzymes (glycoside hydrolases/lyases/phosphorylases): the feature highlights short runs within well-ordered core helices, often enriched in hydrophobic/aromatic residues, and typically avoids catalytic motifs and signal/transit peptides. | 1.55 | 3 |
| #11906 | Post-signal N-terminal domain start | Order/disorder boundary and post-signal-peptide / domain-start segments in secreted and extracellular proteins: the feature activates on the immediately mature N-terminal portions of secreted/lumenal or extracellular proteins just after signal peptide cleavage, and more generally on flexible interdomain linkers or domain-start regions that transition from low-complexity/disordered sequence into structured elements (often initial beta strands), occasionally encompassing N-glycosylation sites. | 1.52 | 2 |
Shared features by |Δ| activation — 916 total (prevalence ≥ 2)
| Feature | Label | Description | Δ activation | Iso act. | Canon act. | Iso prev. | Canon prev. |
|---|---|---|---|---|---|---|---|
| #10704 | N-terminal PEST-like IDRs | Low-complexity, often acidic/proline-rich intrinsically disordered N-terminal tails of intracellular eukaryotic proteins, frequently overlapping annotated PEST-like segments. | +5.24 | 10.09 | 4.84 | 18 | 8 |
| #9998 | Weak N+3 positional cue | Absolute N-terminal positional cue centered near the fourth residue (N+3), independent of amino-acid identity; most often within signal/transit peptides of precursors but also present in non-secreted proteins; the effect is weak and often does not produce a discrete detected site. | +4.07 | 11.70 | 7.62 | 3 | 2 |
| #14234 | Ordered cores and transmembrane helices | Generic preference for ordered chain regions of folded proteins, with particularly strong responses in hydrophobic α-helices that span membranes; avoids signal peptides and low-confidence N-terminal precursor regions. | +3.91 | 8.46 | 4.55 | 17 | 17 |
| #12048 | Generic secondary-structure context marker | A near-ubiquitous, low-amplitude feature marking generic local secondary-structure context—residues within α-helices, β-strands, and the connecting loops—with a frequent emphasis on helix–coil junctions and domain/linker regions, independent of residue type, function, or taxonomy. | +3.50 | 8.96 | 5.46 | 20 | 18 |
| #8063 | Disordered N-terminal activation window | Short N-terminal disordered segment found in a variety of proteins, including plant chloroplast transit peptides, bacterial prokaryotic ubiquitin-like protein Pup N-termini, anti-sigma factor N-tails, and some small RiPP precursor leaders. The activating region is Ser/Thr/Ala/Pro/Gly/Gln-rich, low in acidic residues, and typically lies in a disordered N-terminal stretch; the feature does not generally mark all transit peptides or all RiPP leaders. | +3.44 | 7.48 | 4.05 | 11 | 8 |
| #13334 | N-terminal IDR boundary | Recognition of short, low-complexity, intrinsically disordered N-terminal tails immediately upstream of the first folded domain; these segments are often Ser/Thr/Pro/Gly/Ala-biased and sometimes extend a few residues into the initial helix or the very beginning of the first domain (domain boundary/disorder-to-order transition), especially in transcription regulators and co-regulators. | +3.33 | 7.79 | 4.46 | 42 | 17 |
| #6854 | Low complexity helical repeats | A general, composition-driven signal for non-globular sequence regions: it prefers low-complexity, compositionally biased tracts and repetitive helical polymers, emphasizing charged/polar (E, D, K, R, S, Q) or Gly/Pro-rich content, and residues at helix–coil transitions; it is largely muted in compact, well-packed helical cores. | +3.26 | 5.69 | 2.43 | 28 | 12 |
| #10977 | Early N-terminal disordered patch | Short N-terminal patch at the extreme N-terminus (around residues 5–10), capturing the first distinctive segment of preproteins/precursors and N-terminal disordered tails; residue composition is heterogeneous but frequently polar (Ser/Thr) and/or charged. | +3.13 | 7.84 | 4.71 | 7 | 3 |
| #13449 | Membrane-biased alpha-helix detector | Generic alpha-helix detector with strongest preference for long hydrophobic helices that associate with membranes (transmembrane segments and signal peptides), but also activating on extended coiled-coils and internal alpha-helices in soluble proteins | +2.77 | 8.12 | 5.35 | 15 | 12 |
| #8015 | Short helical docking motifs | Short protein–protein interaction segments in eukaryotic intracellular signaling scaffolds and regulators; includes amphipathic docking/dimerization helices as well as interaction patches embedded within small folded signaling domains (e.g., RA, DH, PH), exemplified by AKAP PKA regulatory subunit–binding helices and extending to BH3, bHLH/coiled-coil, and other helical docking elements across diverse signaling, cytoskeletal, and transcriptional complexes. | +2.69 | 6.73 | 4.05 | 12 | 8 |
| #5437 | Coiled-coils and propeptide spacers | Long α-helical coiled-coil regions of large scaffold/structural proteins, plus the propeptide/spacer regions of secreted neuropeptide and prohormone precursors. The feature responds to heptad-repeat coiled-coil sequence character (typically with R/K-X-X-L/E patterns and charged residues in a/d/e/g positions) and to short polar/charged propeptide segments flanking dibasic cleavage sites. | +2.52 | 4.11 | 1.59 | 13 | 2 |
| #2895 | Motif-and-repeat sequence detector | Sequence-pattern detector that responds most strongly to bacterial helix-turn-helix DNA-binding motifs, particularly the C-terminal HTH of sigma-54-dependent response regulators and related HTH-type transcription factors, and also to (i) dibasic Lys/Arg convertase cleavage motifs and residues flanking amidated peptides in secreted prohormone/peptide precursors, (ii) Gly–X–Y/proline-rich collagen-like repeats (including hydroxyproline sites), (iii) hydrophobic a/d positions in coiled-coil heptads of long scaffolds, and (iv) PTM-prone Ser/Thr embedded in Lys/Arg-rich disordered tails of chromatin/DNA-binding proteins. The unifying concept is recognition of short DNA-binding/processing/PTM motifs and repetitive/low-complexity sequence context used across diverse proteins. | +2.23 | 4.57 | 2.34 | 12 | 7 |
| #15815 | Basic N-terminal targeting patch | Short, basic/polar N‑terminal leader/transit segment immediately after the initiator methionine (positions ~3–6) — i.e., the positively charged/polar start of signal peptides and organelle transit peptides, and more generally low‑acidic, disordered N‑terminal patches (including occasional cysteine‑rich starts) used for targeting, export, rRNA binding, or membrane topology cues | +2.23 | 10.90 | 8.67 | 7 | 4 |
| #11526 | Glycine amidation motif detector | Residue-level detector for small/flexible residues—especially glycine—in short, low-structure linkers and proteolytic processing signals of peptide precursors, with a strong preference for the C‑terminal amidation context in which a glycine (amide donor) immediately precedes mono/di‑basic residues (G‑K/R); outside precursors it gives sparse hits on similar small-residue sites in intrinsically disordered regions and occasionally within structured domains across diverse taxa. | +2.17 | 4.33 | 2.17 | 13 | 7 |
| #4834 | DNA-binding and helical repeats | General helical secondary-structure elements, with strongest activation on short DNA-binding recognition helices (HTH/homeodomain/σ-factor region 4) and on extended helical repeats including collagen Gly–X–Y triple helices; also frequently engages signal/transmembrane helices in secreted and membrane proteins | +2.14 | 4.73 | 2.58 | 12 | 4 |
| #612 | Proline-rich disordered regulatory linkers | Intrinsically disordered regulatory linkers and targeting segments that flank signaling/catalytic domains—most prominently adjacent to RhoGAP domains—enriched in proline/serine/threonine/acidic/glycine residues and often containing SH3-binding PxxP motifs and/or polybasic membrane/focal-adhesion–targeting patches; in some cases the activation extends into the N-terminal edge of the RhoGAP catalytic domain (including the arginine-finger vicinity). Similar sequence/structural signatures are also recognized in flexible interdomain hinges of other eukaryotic signaling proteins. | +2.07 | 4.57 | 2.49 | 41 | 8 |
| #7565 | Hydrophobic helices and signals | Generic, low-specificity signal for short hydrophobic/alpha-helical stretches with a mild N-terminal bias, encompassing signal/transit peptides, single-pass transmembrane anchors, and coiled-coil helices; also yields weak single-residue responses to hydrophobic/aromatic positions within structured loops and peptide-hormone segments. | +2.06 | 7.83 | 5.77 | 16 | 13 |
| #1722 | S/T/Pro-rich disordered regions | Intrinsically disordered, low-complexity sequence elements enriched in Ser/Thr/Pro/polar residues, characteristic of flexible linkers, regulatory tails, and polar/low-complexity tracts across diverse taxa. | +2.04 | 5.66 | 3.62 | 12 | 5 |
| #5321 | N-terminal domain-start leader regions | N-terminal leader/capping segments at the start of a protein or of a new domain, including pre-/pro-peptides, signal-anchor transmembrane helices with their flanking helices, and the first β/α elements that initiate the structured core; the feature emphasizes positional “domain/start-of-mature-chain” regions rather than catalytic metal-binding cores. | +1.98 | 3.68 | 1.70 | 34 | 4 |
| #3604 | Acidic Pro/Ser-rich disordered regions | Proline/serine-rich low-complexity and disordered regions, frequently with acidic/PEST-like character; the feature also fires on small proteins and on internal stretches outside obvious low-complexity tracts when local Pro/Arg/Ser content is high. | +1.91 | 4.09 | 2.18 | 6 | 4 |
| #11395 | Spire KIND/WH2 with charged LCRs | Conserved motifs within the KIND domain and flanking the WH2 actin-binding motifs of Spire family actin nucleators, together with low-complexity, intrinsically disordered regions enriched in charged/polar residues (E/D/S/T/K/R) in other proteins; occasionally shows weaker signal within or flanking helical segments (e.g., transmembrane or coiled-coil boundaries). | +1.86 | 5.64 | 3.78 | 9 | 7 |
| #9044 | General helical structural elements | General helical structural elements: the feature marks residues within helical segments across many protein classes—alpha-helices in HTH/homeodomain DNA-binding regions, coiled-coils, transmembrane helices, and signal-peptide/propeptide helices—with a characteristic periodic emphasis on Gly in the collagen triple-helix (Gly–X–Y) repeat. | +1.76 | 3.86 | 2.09 | 11 | 7 |
| #2257 | Eukaryotic low-complexity IDRs | Intrinsic disorder/low-complexity signal: the feature marks long, compositionally biased IDRs in eukaryotic proteins—especially Lys/Arg-rich basic patches, Glu/Asp-rich acidic tracts, and Q/N/S/T-rich stretches—found in N-/C-terminal tails and inter-domain linkers, while avoiding well-folded domains and transmembrane helices; commonly seen in translation/RNA biogenesis factors and secretory/quality-control proteins. | +1.76 | 3.35 | 1.59 | 27 | 4 |
| #6452 | Chlorophyll/heme NGRAAM/WxWG motifs | A sequence-motif feature firing on short conserved patches in chlorophyll/heme-related proteins, including light-harvesting-like (LHC/HLI/OHP) proteins and ferrochelatases, with recurrent context around an "NGRAAM"-like motif and an aromatic "WxWG" patch; also activates on short patches in small membrane-regulator peptides. | +1.60 | 4.86 | 3.26 | 19 | 17 |
| #1370 | Short low-complexity N-terminal segment | Short, low‑complexity, intrinsically disordered N‑terminal segments (~10–20 aa) that frequently correspond to leader/transit/propeptide presequences (e.g., chloroplast transit peptides, RiPP/bacteriocin leaders) or generic unstructured N‑tails; typically non‑hydrophobic and enriched in Ser/Thr/Pro but tolerant of basic (poly‑Arg/Lys) or Cys‑rich compositions | +1.58 | 4.16 | 2.58 | 12 | 10 |
| #12004 | N-terminal Pro/His helical motifs | Short N-terminal sequence motifs in metabolic enzymes, characterized by Pro- and His-containing patches embedded in early α-helical segments (e.g. S-P-x-P-Y/F-H in aspartyl aminopeptidases; G-H-P-D in S-adenosylmethionine synthases) | +1.57 | 3.73 | 2.16 | 6 | 2 |
| #14712 | SPS N-terminal regulatory motif | A conserved N-terminal regulatory motif in plant sucrose-phosphate synthases (SPS), centered on a short ATR(N/S/P)xRERS / SPQERN-type sequence upstream of the catalytic glycosyltransferase core. | +1.55 | 5.53 | 3.98 | 16 | 11 |
| #1803 | Unknown generic feature | Unknown generic feature | +6.04 | 17.75 | 11.71 | 174 | 143 |
| #14534 | Unknown generic feature | Unknown generic feature | +3.75 | 15.23 | 11.48 | 173 | 143 |
| #9005 | Unknown generic feature | Unknown generic feature | +2.38 | 12.15 | 9.77 | 142 | 124 |
Part 2 · Differential coordinates
Features firing on just the isoform-unique (extension / alt-frame) residues.
Unique-region features — 241 distinct (prevalence ≥ 2)
| Feature | Label | Description | Activation | Prevalence |
|---|---|---|---|---|
| #5725 | UBC/E2-like fold recognition | UBC/E2-like fold recognition across ubiquitin and ubiquitin-like conjugation systems, capturing catalytically active E2s and inactive homologs (UEV, RWD) and the cognate ESCRT components that harbor them; the signal corresponds to the conserved UBC-core surface, with a bias toward the N-terminal α-helix, and can extend to physicochemically similar acidic, α-helical blocks that mimic this surface | 17.52 | 31 |
| #4992 | Small-residue disordered N-terminus | Context-dependent free N-terminus signature: short, disordered amino-terminal stretches beginning with the initiator Met followed by residues compatible with MAP cleavage or N-acetylation (e.g., M–G/A/S/T/V, and other contexts including M–D/E/L); strongest emphasis on small residues at positions 2–3; activation is absent when the native N-terminus is structured or starts with a signal peptide/transmembrane segment. | 15.46 | 4 |
| #15815 | Basic N-terminal targeting patch | Short, basic/polar N‑terminal leader/transit segment immediately after the initiator methionine (positions ~3–6) — i.e., the positively charged/polar start of signal peptides and organelle transit peptides, and more generally low‑acidic, disordered N‑terminal patches (including occasional cysteine‑rich starts) used for targeting, export, rRNA binding, or membrane topology cues | 10.90 | 4 |
| #10704 | N-terminal PEST-like IDRs | Low-complexity, often acidic/proline-rich intrinsically disordered N-terminal tails of intracellular eukaryotic proteins, frequently overlapping annotated PEST-like segments. | 10.09 | 13 |
| #5241 | Early N-terminal positional signal | Generic early N-terminus positional signal peaking at residue ~5–7, found broadly within the disordered N-tail of proteins, including both cleavable leader/targeting segments (signal peptides, propeptides) of precursor proteins and disordered N-terminal regions of mature proteins; residue identity is not specific | 9.59 | 2 |
| #1066 | UBC/UEV HPN-loop interface | Strand-helix-loop interaction patch of compact alpha/beta domains, most prominently the UBC/UEV (ubiquitin-conjugating-like) fold where it captures the conserved HPN-loop and adjacent beta-strand/helix that form the partner-binding surface; secondarily, the feature recognizes geometrically analogous beta-strand–helix patches in other interaction/DNA-binding domains (e.g., T-box and EVH1-like). | 9.14 | 30 |
| #12048 | Generic secondary-structure context marker | A near-ubiquitous, low-amplitude feature marking generic local secondary-structure context—residues within α-helices, β-strands, and the connecting loops—with a frequent emphasis on helix–coil junctions and domain/linker regions, independent of residue type, function, or taxonomy. | 8.96 | 4 |
| #14234 | Ordered cores and transmembrane helices | Generic preference for ordered chain regions of folded proteins, with particularly strong responses in hydrophobic α-helices that span membranes; avoids signal peptides and low-confidence N-terminal precursor regions. | 8.46 | 3 |
| #13449 | Membrane-biased alpha-helix detector | Generic alpha-helix detector with strongest preference for long hydrophobic helices that associate with membranes (transmembrane segments and signal peptides), but also activating on extended coiled-coils and internal alpha-helices in soluble proteins | 8.12 | 4 |
| #10977 | Early N-terminal disordered patch | Short N-terminal patch at the extreme N-terminus (around residues 5–10), capturing the first distinctive segment of preproteins/precursors and N-terminal disordered tails; residue composition is heterogeneous but frequently polar (Ser/Thr) and/or charged. | 7.84 | 4 |
| #7565 | Hydrophobic helices and signals | Generic, low-specificity signal for short hydrophobic/alpha-helical stretches with a mild N-terminal bias, encompassing signal/transit peptides, single-pass transmembrane anchors, and coiled-coil helices; also yields weak single-residue responses to hydrophobic/aromatic positions within structured loops and peptide-hormone segments. | 7.83 | 4 |
| #13334 | N-terminal IDR boundary | Recognition of short, low-complexity, intrinsically disordered N-terminal tails immediately upstream of the first folded domain; these segments are often Ser/Thr/Pro/Gly/Ala-biased and sometimes extend a few residues into the initial helix or the very beginning of the first domain (domain boundary/disorder-to-order transition), especially in transcription regulators and co-regulators. | 7.79 | 31 |
| #8063 | Disordered N-terminal activation window | Short N-terminal disordered segment found in a variety of proteins, including plant chloroplast transit peptides, bacterial prokaryotic ubiquitin-like protein Pup N-termini, anti-sigma factor N-tails, and some small RiPP precursor leaders. The activating region is Ser/Thr/Ala/Pro/Gly/Gln-rich, low in acidic residues, and typically lies in a disordered N-terminal stretch; the feature does not generally mark all transit peptides or all RiPP leaders. | 7.48 | 9 |
| #13818 | N-terminal domain start signal | Start-of-domain signal: the feature detects the transition into the first ordered secondary-structure elements at the N-terminus of a conserved globular domain (often immediately after a signal/transit peptide, lipidation site, or disordered tail), with strongest matches for cyclophilin-type PPIase folds and weaker matches for many other domain starts. | 7.44 | 31 |
| #8015 | Short helical docking motifs | Short protein–protein interaction segments in eukaryotic intracellular signaling scaffolds and regulators; includes amphipathic docking/dimerization helices as well as interaction patches embedded within small folded signaling domains (e.g., RA, DH, PH), exemplified by AKAP PKA regulatory subunit–binding helices and extending to BH3, bHLH/coiled-coil, and other helical docking elements across diverse signaling, cytoskeletal, and transcriptional complexes. | 6.73 | 4 |
| #7986 | Intrinsically disordered low-complexity regions | Compositionally biased, intrinsically disordered low‑complexity regions used as flexible interaction/nucleic‑acid–binding modules—enriched in Pro/Ser/Lys/Glu/Gln/Thr/Gly/Arg and encompassing poly‑Pro tracts, Lys/Arg‑rich basic stretches, and acidic Asp/Glu clusters—common in transcription/RNA‑processing factors, viral basic regulators, and the C‑terminal scaffolds of RNase E | 6.68 | 22 |
| #16243 | Structured-region leucine detector | Generic detector of leucine side chains in structured regions of folded domains, with occasional weaker responses to other hydrophobics; largely indifferent to specific functional motifs or protein class. | 6.68 | 2 |
| #8237 | Conserved structured domain peaks | A feature that activates within structured partner-recognition/DNA-binding domains and adjacent ordered segments of eukaryotic regulatory proteins, with peaks falling on conserved residues inside these domains rather than on flanking low-complexity stretches. | 6.67 | 2 |
| #14927 | Structural and domain boundary residues | Residues at structural and domain junctions: starts/ends of secondary-structure elements (especially helix starts and transmembrane-helix termini) and positions immediately preceding or following annotated domains, often adjacent to ligand/cofactor-binding regions | 6.45 | 2 |
| #7026 | Catalytic domain helix/loop boundary motif | A localized internal-domain motif that marks a specific short helix/loop boundary within a catalytic domain, situated downstream of the catalytic proton-donor residue in α/β hydrolase-like folds. | 6.38 | 2 |
| #934 | SprT metalloprotease dual-motif detector | Detector of conserved sequence motifs within SprT-like metalloprotease domains, particularly in SPRTN orthologs and Germ cell nuclear acidic protein (GCNA) family proteins. | 6.32 | 2 |
| #9237 | Glycine detector in disordered regions | Residue-identity detector for glycine (G), with a mild bias toward intrinsically disordered/low-complexity segments (often including very N-terminal residues); pan-taxonomic and function-agnostic, but activation is sparse and frequently absent, becoming detectable mainly in glycine-rich low-complexity contexts; occasional weak responses to other small/polar residues. | 6.14 | 3 |
| #13480 | Alanine and small-residue enrichment | A sequence-composition feature that detects small, non-aromatic residues—especially alanine—and their enrichment, lighting up individual Ala residues and stretches enriched in A/G/S/T/P/V (alanine-rich, low-complexity segments) irrespective of protein function or secondary structure. | 6.07 | 4 |
| #627 | Generic serine detector | Generic serine detector: the feature marks serine residues, with weaker affinity for threonine (and occasionally proline), and is especially prominent in low-complexity S/T-rich stretches; it is context- and structure-agnostic and does not correspond to a specific functional motif. | 6.05 | 4 |
| #14895 | Unknown generic feature | Unknown generic feature | 21.36 | 31 |
| #9214 | Unknown generic feature | Unknown generic feature | 18.11 | 31 |
| #1803 | Unknown generic feature | Unknown generic feature | 17.75 | 31 |
| #9194 | Unknown generic feature | Unknown generic feature | 15.70 | 31 |
| #14534 | Unknown generic feature | Unknown generic feature | 15.23 | 31 |
| #9005 | Unknown generic feature | Unknown generic feature | 12.15 | 31 |
Top-K sparse-autoencoder activations on the ESM-C residual stream. Activation is a feature's peak value over its region; Prevalence is the number of residues it fires on; Δ activation is isoform − canonical. Only the top 30 features per category are saved; generic / "unknown" features are hidden until you expand a table. Read more about SAE features →
Clinical variants
Differential region — N-terminal extension (isoform-unique)
| Pos (iso) | AA change | Consequence | Source | Clin. sig. | AF (gnomAD) | Impact | AlphaMissense | ESM-C ΔLLR | Link |
|---|---|---|---|---|---|---|---|---|---|
| 0 | I→I | synonymous_variant | gnomAD | — | 3.35e-06 | — | N/A | — | chr5-139561701-C-T |
| 2 | E→A | missense_variant | gnomAD | — | 1.07e-06 | — | N/A | 2.60 | chr5-139561706-A-C |
| 2 | E→D | missense_variant | gnomAD | — | 1.05e-06 | — | N/A | -0.89 | chr5-139561707-G-C |
| 2 | E→D | missense_variant | gnomAD | — | 1.05e-06 | — | N/A | -0.89 | chr5-139561707-G-T |
| 3 | A→G | missense_variant | gnomAD | — | 1.04e-06 | — | N/A | -0.27 | chr5-139561709-C-G |
| 4 | A→T | missense_variant | gnomAD | — | 6.06e-06 | — | N/A | -1.62 | chr5-139561711-G-A |
| 4 | A→D | missense_variant | gnomAD | — | 9.89e-07 | — | N/A | -2.11 | chr5-139561712-C-A |
| 4 | A→A | synonymous_variant | gnomAD | — | 1.92e-06 | — | N/A | 0.00 | chr5-139561713-C-G |
| 4 | A→A | synonymous_variant | gnomAD | — | 1.63e-05 | — | N/A | 0.00 | chr5-139561713-C-T |
| 5 | S→P | missense_variant | gnomAD | — | 1.91e-06 | — | N/A | 0.70 | chr5-139561714-T-C |
| 5 | S→S | synonymous_variant | gnomAD | — | 9.50e-07 | — | N/A | 0.00 | chr5-139561716-A-C |
| 5 | S→S | synonymous_variant | gnomAD | — | 8.55e-06 | — | N/A | 0.00 | chr5-139561716-A-G |
| 6 | G→S | missense_variant | gnomAD | — | 9.29e-07 | — | N/A | 0.45 | chr5-139561717-G-A |
| 6 | G→D | missense_variant | gnomAD | — | 5.02e-05 | — | N/A | -2.19 | chr5-139561718-G-A |
| 6 | G→V | missense_variant | gnomAD | — | 9.30e-07 | — | N/A | -0.69 | chr5-139561718-G-T |
| 7 | — | frameshift_variant | gnomAD | — | 9.16e-07 | LoF | — | — | chr5-139561720-T-TC |
| 7 | S→F | missense_variant | gnomAD | — | 7.59e-05 | — | N/A | -1.64 | chr5-139561721-C-T |
| 7 | S→S | synonymous_variant | gnomAD | — | 8.80e-07 | — | N/A | 0.00 | chr5-139561722-C-T |
| 8 | L→L | synonymous_variant | gnomAD | — | 3.27e-05 | — | N/A | 0.00 | chr5-139561725-A-G |
| 9 | A→T | missense_variant | gnomAD | — | 1.74e-06 | — | N/A | -1.50 | chr5-139561726-G-A |
| 9 | A→S | missense_variant | gnomAD | — | 6.10e-06 | — | N/A | 0.91 | chr5-139561726-G-T |
| 9 | — | frameshift_variant | gnomAD | — | 8.65e-07 | LoF | — | — | chr5-139561727-C-CA |
| 9 | — | frameshift_variant | gnomAD | — | 2.42e-05 | LoF | — | — | chr5-139561727-C-CCCCTTCCCCGT |
| 9 | A→V | missense_variant | gnomAD | — | 3.46e-06 | — | N/A | -1.64 | chr5-139561727-C-T |
| 9 | — | frameshift_variant | gnomAD | — | 2.68e-05 | LoF | — | — | chr5-139561727-CCCCTTCCCCGT-C |
| 9 | A→A | synonymous_variant | gnomAD | — | 8.58e-07 | — | N/A | 0.00 | chr5-139561728-C-A |
| 9 | A→A | synonymous_variant | gnomAD | — | 1.72e-06 | — | N/A | 0.00 | chr5-139561728-C-G |
| 10 | P→T | missense_variant | gnomAD | — | 2.55e-06 | — | N/A | -2.18 | chr5-139561729-C-A |
| 10 | P→S | missense_variant | gnomAD | — | 1.70e-06 | — | N/A | -0.71 | chr5-139561729-C-T |
| 10 | P→P | synonymous_variant | gnomAD | — | 1.69e-06 | — | N/A | 0.00 | chr5-139561731-T-C |
| 11 | — | frameshift_variant | gnomAD | — | 2.51e-06 | LoF | — | — | chr5-139561732-TCCCCGTCCCTTCCCCGC-T |
| 11 | S→S | synonymous_variant | gnomAD | — | 3.30e-06 | — | N/A | 0.00 | chr5-139561734-C-T |
| 12 | P→A | missense_variant | gnomAD | — | 2.46e-06 | — | N/A | -0.88 | chr5-139561735-C-G |
| 12 | P→S | missense_variant | gnomAD | — | 8.21e-07 | — | N/A | -0.86 | chr5-139561735-C-T |
| 12 | P→R | missense_variant | gnomAD | — | 3.29e-06 | — | N/A | -0.72 | chr5-139561736-C-G |
| 12 | P→L | missense_variant | gnomAD | — | 8.05e-05 | — | N/A | -0.50 | chr5-139561736-C-T |
| 12 | P→P | synonymous_variant | gnomAD | — | 1.82e-06 | — | N/A | 0.00 | chr5-139561737-G-A |
| 12 | P→P | synonymous_variant | gnomAD | — | 9.11e-07 | — | N/A | 0.00 | chr5-139561737-G-T |
| 13 | S→T | missense_variant | gnomAD | — | 8.67e-07 | — | N/A | -2.00 | chr5-139561738-T-A |
| 13 | S→A | missense_variant | gnomAD | — | 9.11e-05 | — | N/A | 0.20 | chr5-139561738-T-G |
| 13 | — | frameshift_variant | gnomAD | — | 3.47e-06 | LoF | — | — | chr5-139561738-T-TC |
| 13 | — | frameshift_variant | gnomAD | — | 8.67e-07 | LoF | — | — | chr5-139561738-T-TG |
| 13 | — | inframe_deletion | gnomAD | — | 7.81e-06 | — | N/A | — | chr5-139561738-TCCCTTCCCCGCC-T |
| 13 | S→F | missense_variant | gnomAD | — | 8.39e-07 | — | N/A | -1.56 | chr5-139561739-C-T |
| 13 | S→S | synonymous_variant | COSMIC | — | — | — | N/A | 0.00 | COSV54077651 |
| 14 | — | frameshift_variant | gnomAD | — | 3.11e-05 | LoF | — | — | chr5-139561741-C-CA |
| 14 | — | frameshift_variant | gnomAD | — | 3.75e-05 | LoF | — | — | chr5-139561741-C-CG |
| 14 | — | frameshift_variant | gnomAD | — | 7.98e-07 | LoF | — | — | chr5-139561741-C-CT |
| 14 | L→F | missense_variant | gnomAD | — | 7.98e-07 | — | N/A | -1.80 | chr5-139561741-C-T |
| 14 | L→P | missense_variant | gnomAD | — | 2.80e-06 | — | N/A | 1.44 | chr5-139561742-T-C |
| 14 | L→R | missense_variant | gnomAD | — | 1.11e-04 | — | N/A | 1.66 | chr5-139561742-T-G |
| 14 | — | frameshift_variant | gnomAD | — | 1.77e-05 | LoF | — | — | chr5-139561742-TTCCCCGCCCCC-T |
| 14 | — | inframe_deletion | gnomAD | — | 1.01e-04 | — | N/A | — | chr5-139561742-TTCCCCGCCCCCG-T |
| 14 | L→L | synonymous_variant | gnomAD | — | 1.00e-05 | — | N/A | 0.00 | chr5-139561743-T-A |
| 14 | L→L | synonymous_variant | gnomAD | — | 2.25e-05 | — | N/A | 0.00 | chr5-139561743-T-C |
| 14 | — | frameshift_variant | gnomAD | — | 2.39e-04 | LoF | — | — | chr5-139561743-T-TA |
| 14 | — | frameshift_variant | gnomAD | — | 2.35e-04 | LoF | — | — | chr5-139561743-T-TC |
| 14 | — | frameshift_variant | gnomAD | — | 1.79e-04 | LoF | — | — | chr5-139561743-T-TG |
| 14 | — | frameshift_variant | gnomAD | — | 8.34e-07 | LoF | — | — | chr5-139561743-T-TTCCC |
| 14 | — | frameshift_variant | gnomAD | — | 5.84e-06 | LoF | — | — | chr5-139561743-TCCCC-T |
| 14 | — | inframe_deletion | gnomAD | — | 6.67e-06 | — | N/A | — | chr5-139561743-TCCCCGC-T |
| 14 | — | frameshift_variant | gnomAD | — | 8.34e-07 | LoF | — | — | chr5-139561743-TCCCCGCCCCCGTCCCCG-T |
| 15 | P→S | missense_variant | gnomAD | — | 9.21e-06 | — | N/A | -0.91 | chr5-139561744-C-T |
| 15 | P→H | missense_variant | gnomAD | — | 1.59e-06 | — | N/A | -2.97 | chr5-139561745-C-A |
| 15 | P→R | missense_variant | gnomAD | — | 1.59e-06 | — | N/A | -0.58 | chr5-139561745-C-G |
| 15 | P→L | missense_variant | gnomAD | — | 2.38e-04 | — | N/A | -1.36 | chr5-139561745-C-T |
| 15 | P→P | synonymous_variant | gnomAD | — | 1.57e-06 | — | N/A | 0.00 | chr5-139561746-C-A |
| 15 | P→P | synonymous_variant | gnomAD | — | 7.87e-07 | — | N/A | 0.00 | chr5-139561746-C-G |
| 15 | P→P | synonymous_variant | gnomAD | — | 3.54e-05 | — | N/A | 0.00 | chr5-139561746-C-T |
| 16 | R→S | missense_variant | gnomAD | — | 1.58e-06 | — | N/A | -0.88 | chr5-139561747-C-A |
| 16 | R→C | missense_variant | gnomAD | — | 1.58e-06 | — | N/A | -2.38 | chr5-139561747-C-T |
| 16 | — | frameshift_variant | gnomAD | — | 8.06e-05 | LoF | — | — | chr5-139561747-CG-C |
| 16 | R→H | missense_variant | gnomAD | — | 5.08e-06 | — | N/A | -2.27 | chr5-139561748-G-A |
| 16 | — | frameshift_variant | gnomAD | — | 2.03e-05 | LoF | — | — | chr5-139561748-G-GC |
| 16 | R→L | missense_variant | gnomAD | — | 2.03e-05 | — | N/A | -0.65 | chr5-139561748-G-T |
| 16 | — | frameshift_variant | gnomAD | — | 2.54e-06 | LoF | — | — | chr5-139561748-GC-G |
| 16 | — | frameshift_variant | gnomAD | — | 2.54e-06 | LoF | — | — | chr5-139561748-GCC-G |
| 16 | R→R | synonymous_variant | gnomAD | — | 7.87e-07 | — | N/A | 0.00 | chr5-139561749-C-A |
| 16 | R→R | synonymous_variant | gnomAD | — | 1.42e-05 | — | N/A | 0.00 | chr5-139561749-C-T |
| 16 | — | frameshift_variant | COSMIC | — | — | LoF | — | — | COSV99551267 |
| 17 | P→A | missense_variant | gnomAD | — | 7.79e-07 | — | N/A | -1.25 | chr5-139561750-C-G |
| 17 | P→S | missense_variant | gnomAD | — | 3.97e-05 | — | N/A | -1.87 | chr5-139561750-C-T |
| 17 | P→R | missense_variant | gnomAD | — | 1.55e-06 | — | N/A | -0.91 | chr5-139561751-C-G |
| 17 | P→P | synonymous_variant | gnomAD | — | 1.54e-06 | — | N/A | 0.00 | chr5-139561752-C-G |
| 18 | R→S | missense_variant | gnomAD | — | 7.68e-07 | — | N/A | -1.22 | chr5-139561753-C-A |
| 18 | R→G | missense_variant | gnomAD | — | 7.68e-07 | — | N/A | -0.88 | chr5-139561753-C-G |
| 18 | R→C | missense_variant | gnomAD | — | 1.54e-06 | — | N/A | -2.36 | chr5-139561753-C-T |
| 18 | — | frameshift_variant | gnomAD | — | 1.54e-06 | LoF | — | — | chr5-139561753-CG-C |
| 18 | — | inframe_deletion | gnomAD | — | 7.68e-07 | — | N/A | — | chr5-139561753-CGTCCCCGCCCCG-C |
| 18 | R→H | missense_variant | gnomAD | — | 2.33e-06 | — | N/A | -2.33 | chr5-139561754-G-A |
| 18 | R→L | missense_variant | gnomAD | — | 1.40e-05 | — | N/A | -1.20 | chr5-139561754-G-T |
| 18 | — | frameshift_variant | gnomAD | — | 6.98e-06 | LoF | — | — | chr5-139561754-GT-G |
| 18 | R→R | synonymous_variant | gnomAD | — | 3.18e-06 | — | N/A | 0.00 | chr5-139561755-T-A |
| 18 | R→R | synonymous_variant | gnomAD | — | 2.13e-04 | — | N/A | 0.00 | chr5-139561755-T-C |
| 18 | R→R | synonymous_variant | gnomAD | — | 4.52e-04 | — | N/A | 0.00 | chr5-139561755-T-G |
| 18 | — | frameshift_variant | gnomAD | — | 2.33e-05 | LoF | — | — | chr5-139561755-T-TC |
| 18 | — | frameshift_variant | gnomAD | — | 2.12e-06 | LoF | — | — | chr5-139561755-TC-T |
| 19 | P→S | missense_variant | gnomAD | — | 1.56e-06 | — | N/A | -1.66 | chr5-139561756-C-T |
| 19 | P→P | synonymous_variant | gnomAD | — | 7.48e-07 | — | N/A | 0.00 | chr5-139561758-C-T |
| 20 | R→G | missense_variant | gnomAD | — | 1.50e-06 | — | N/A | -0.47 | chr5-139561759-C-G |
| 20 | R→C | missense_variant | gnomAD | — | 7.49e-07 | — | N/A | -2.55 | chr5-139561759-C-T |
| 20 | R→H | missense_variant | gnomAD | — | 4.60e-05 | — | N/A | -2.43 | chr5-139561760-G-A |
| 20 | R→L | missense_variant | gnomAD | — | 7.66e-06 | — | N/A | -1.40 | chr5-139561760-G-T |
| 20 | — | frameshift_variant | gnomAD | — | 2.55e-06 | LoF | — | — | chr5-139561760-GC-G |
| 20 | R→R | synonymous_variant | gnomAD | — | 7.48e-07 | — | N/A | 0.00 | chr5-139561761-C-A |
| 20 | R→R | synonymous_variant | gnomAD | — | 1.50e-06 | — | N/A | 0.00 | chr5-139561761-C-G |
| 20 | R→R | synonymous_variant | gnomAD | — | 1.50e-06 | — | N/A | 0.00 | chr5-139561761-C-T |
| 21 | P→S | missense_variant | gnomAD | — | 1.49e-06 | — | N/A | -1.27 | chr5-139561762-C-T |
| 21 | — | frameshift_variant | gnomAD | — | 2.98e-06 | LoF | — | — | chr5-139561764-C-CG |
| 21 | P→P | synonymous_variant | gnomAD | — | 2.24e-06 | — | N/A | 0.00 | chr5-139561764-C-G |
| 21 | P→P | synonymous_variant | gnomAD | — | 7.45e-07 | — | N/A | 0.00 | chr5-139561764-C-T |
| 21 | — | frameshift_variant | gnomAD | — | 3.73e-06 | LoF | — | — | chr5-139561764-CG-C |
| 22 | G→E | missense_variant | gnomAD | — | 1.20e-04 | — | N/A | -1.92 | chr5-139561766-G-A |
| 23 | G→D | missense_variant | gnomAD | — | 2.24e-06 | — | N/A | -1.42 | chr5-139561769-G-A |
| 23 | G→G | synonymous_variant | gnomAD | — | 3.02e-06 | — | N/A | 0.00 | chr5-139561770-C-A |
| 24 | R→H | missense_variant | gnomAD | — | 1.04e-06 | — | N/A | -1.71 | chr5-139561772-G-A |
| 24 | R→R | synonymous_variant | gnomAD | — | 7.44e-07 | — | N/A | 0.00 | chr5-139561773-C-T |
| 25 | R→C | missense_variant | gnomAD | — | 7.44e-07 | — | N/A | -2.36 | chr5-139561774-C-T |
| 25 | R→H | missense_variant | gnomAD | — | 1.01e-06 | — | N/A | -2.17 | chr5-139561775-G-A |
| 25 | — | frameshift_variant | gnomAD | — | 2.03e-06 | LoF | — | — | chr5-139561775-G-GCCACCCGCCTC |
| 25 | R→L | missense_variant | gnomAD | — | 1.01e-06 | — | N/A | -1.19 | chr5-139561775-G-T |
| 26 | — | frameshift_variant | gnomAD | — | 9.97e-07 | LoF | — | — | chr5-139561778-AC-A |
| 26 | H→Q | missense_variant | gnomAD | — | 7.42e-06 | — | N/A | 0.39 | chr5-139561779-C-G |
| 26 | H→H | synonymous_variant | gnomAD | — | 9.65e-06 | — | N/A | 0.00 | chr5-139561779-C-T |
| 27 | P→A | missense_variant | gnomAD | — | 7.42e-07 | — | N/A | -1.67 | chr5-139561780-C-G |
| 27 | P→S | missense_variant | gnomAD | — | 1.48e-06 | — | N/A | -1.33 | chr5-139561780-C-T |
| 27 | P→R | missense_variant | gnomAD | — | 1.48e-06 | — | N/A | -0.09 | chr5-139561781-C-G |
| 27 | P→P | synonymous_variant | gnomAD | — | 2.18e-06 | — | N/A | 0.00 | chr5-139561782-G-A |
| 27 | P→P | synonymous_variant | gnomAD | — | 1.09e-06 | — | N/A | 0.00 | chr5-139561782-G-T |
| 28 | P→S | missense_variant | gnomAD | — | 1.48e-06 | — | N/A | -1.09 | chr5-139561783-C-T |
| 28 | P→R | missense_variant | gnomAD | — | 4.44e-05 | — | N/A | -0.31 | chr5-139561784-C-G |
| 28 | P→L | missense_variant | gnomAD | — | 1.48e-06 | — | N/A | -1.26 | chr5-139561784-C-T |
| 29 | P→S | missense_variant | gnomAD | — | 1.48e-06 | — | N/A | -1.24 | chr5-139561786-C-T |
| 29 | P→R | missense_variant | gnomAD | — | 7.40e-07 | — | N/A | -0.68 | chr5-139561787-C-G |
| 29 | P→L | missense_variant | gnomAD | — | 7.40e-07 | — | N/A | -1.54 | chr5-139561787-C-T |
| 29 | P→P | synonymous_variant | gnomAD | — | 1.33e-05 | — | N/A | 0.00 | chr5-139561788-C-T |
| 29 | P→S | missense_variant | COSMIC | — | — | — | N/A | -1.24 | COSV99551256 |
| 30 | T→A | missense_variant | gnomAD | — | 8.87e-07 | — | N/A | 1.28 | chr5-139561789-A-G |
| 30 | T→T | synonymous_variant | gnomAD | — | 2.95e-06 | — | N/A | 0.00 | chr5-139561791-C-T |
139 variants in the differential region.
Shared canonical core — sequence common to canonical and isoform; AlphaMissense applies here
| Pos (iso) | AA change | Consequence | Source | Clin. sig. | AF (gnomAD) | Impact | AlphaMissense | ESM-C ΔLLR | Link |
|---|---|---|---|---|---|---|---|---|---|
| 31 | M→L | missense_variant | gnomAD | — | 3.29e-06 | damaging | — | -9.87 | chr5-139561792-A-C |
| 31 | M→V | missense_variant | gnomAD | — | 1.64e-06 | damaging | — | -10.50 | chr5-139561792-A-G |
| 31 | — | frameshift_variant | gnomAD | — | 8.96e-07 | LoF | — | — | chr5-139561793-TG-T |
| 32 | A→S | missense_variant | gnomAD | — | 7.37e-07 | damaging | ambiguous (0.45) | -7.84 | chr5-139561795-G-T |
| 32 | A→A | synonymous_variant | gnomAD | — | 8.36e-06 | — | — | 0.00 | chr5-139561797-T-G |
| 33 | L→L | synonymous_variant | gnomAD | — | 3.61e-05 | — | — | 0.00 | chr5-139561798-C-T |
| 34 | K→E | missense_variant | gnomAD | — | 7.41e-07 | damaging | likely_pathogenic (0.94) | -9.31 | chr5-139561801-A-G |
| 34 | K→K | synonymous_variant | gnomAD | — | 7.44e-07 | — | — | 0.00 | chr5-139561803-G-A |
| 35 | R→R | synonymous_variant | gnomAD | — | 7.40e-07 | — | — | 0.00 | chr5-139561804-A-C |
| 36 | I→I | synonymous_variant | gnomAD | — | 7.54e-07 | — | — | 0.00 | chr5-139561809-C-T |
| 37 | H→H | synonymous_variant | gnomAD | — | 7.39e-07 | — | — | 0.00 | chr5-139561812-C-T |
| 37 | H→P | missense_variant | COSMIC | — | — | damaging | likely_pathogenic (0.88) | -11.36 | COSV105026649 |
| 38 | K→E | missense_variant | COSMIC | — | — | damaging | likely_pathogenic (0.95) | -10.50 | COSV54077432 |
| 38 | K→K | synonymous_variant | COSMIC | — | — | — | — | 0.00 | COSV105026650 |
| 39 | E→Q | missense_variant | ClinVar | — | — | damaging | likely_pathogenic (0.99) | -11.25 | ClinVar:4311627 |
| 39 | E→Q | missense_variant | COSMIC | — | — | damaging | likely_pathogenic (0.99) | -11.25 | COSV99551389 |
| 41 | N→N | synonymous_variant | gnomAD | — | 6.84e-07 | — | — | 0.00 | chr5-139600380-T-C |
| 42 | D→D | synonymous_variant | gnomAD | — | 1.37e-06 | — | — | 0.00 | chr5-139600383-T-C |
| 43 | L→L | synonymous_variant | gnomAD | — | 3.28e-05 | — | — | 0.00 | chr5-139600384-C-T |
| 43 | L→L | synonymous_variant | gnomAD | — | 6.84e-07 | — | — | 0.00 | chr5-139600386-G-C |
| 44 | A→A | synonymous_variant | gnomAD | — | 6.84e-07 | — | — | 0.00 | chr5-139600389-A-G |
| 45 | R→G | missense_variant | gnomAD | — | 6.84e-07 | damaging | likely_pathogenic (0.67) | -10.06 | chr5-139600390-C-G |
| 45 | R→W | missense_variant | gnomAD | — | 6.84e-07 | damaging | likely_pathogenic (0.82) | -11.43 | chr5-139600390-C-T |
| 45 | R→R | synonymous_variant | gnomAD | — | 2.74e-06 | — | — | 0.00 | chr5-139600392-G-A |
| 45 | R→Q | missense_variant | COSMIC | — | — | damaging | ambiguous (0.37) | -8.31 | COSV54079740 |
| 46 | D→D | synonymous_variant | gnomAD | — | 6.84e-07 | — | — | 0.00 | chr5-139600395-C-T |
| 48 | P→P | synonymous_variant | COSMIC | — | — | — | — | 0.00 | COSV54077679 |
| 49 | A→S | missense_variant | gnomAD | — | 6.84e-07 | — | likely_benign (0.15) | -5.78 | chr5-139600402-G-T |
| 49 | A→A | synonymous_variant | gnomAD | — | 6.84e-07 | — | — | 0.00 | chr5-139600404-A-G |
| 49 | A→A | synonymous_variant | gnomAD | — | 6.84e-07 | — | — | 0.00 | chr5-139600404-A-T |
| 50 | Q→L | missense_variant | gnomAD | — | 6.84e-07 | damaging | likely_benign (0.31) | -9.19 | chr5-139600406-A-T |
| 51 | C→C | synonymous_variant | gnomAD | — | 6.84e-07 | — | — | 0.00 | chr5-139600410-T-C |
| 52 | S→S | synonymous_variant | gnomAD | — | 2.05e-06 | — | — | 0.00 | chr5-139600413-A-T |
| 53 | A→S | missense_variant | COSMIC | — | — | damaging | likely_pathogenic (0.78) | -9.19 | COSV108064706 |
| 55 | P→P | synonymous_variant | gnomAD | — | 6.84e-07 | — | — | 0.00 | chr5-139600422-T-C |
| 58 | D→D | synonymous_variant | gnomAD | — | 4.10e-06 | — | — | 0.00 | chr5-139600431-T-C |
| 58 | D→E | missense_variant | gnomAD | — | 2.05e-06 | — | likely_benign (0.17) | -3.06 | chr5-139600431-T-G |
| 59 | D→D | synonymous_variant | gnomAD | — | 6.84e-07 | — | — | 0.00 | chr5-139600434-T-C |
| 59 | D→N | missense_variant | COSMIC | — | — | damaging | ambiguous (0.56) | -8.00 | COSV54079619 |
| 59 | D→D | synonymous_variant | COSMIC | — | — | — | — | 0.00 | COSV54080651 |
| 60 | M→T | missense_variant | gnomAD | — | 6.85e-07 | damaging | ambiguous (0.47) | -8.18 | chr5-139614586-T-C |
| 61 | F→F | synonymous_variant | COSMIC | — | — | — | — | 0.00 | COSV99551398 |
| 61 | F→L | missense_variant | COSMIC | — | — | damaging | likely_pathogenic (0.99) | -8.62 | COSV54080918 |
| 61 | F→L | missense_variant | COSMIC | — | — | damaging | likely_pathogenic (0.99) | -8.62 | COSV108064633 |
| 62 | H→H | synonymous_variant | gnomAD | — | 6.85e-07 | — | — | 0.00 | chr5-139614593-T-C |
| 65 | A→V | missense_variant | COSMIC | — | — | damaging | likely_pathogenic (0.98) | -9.94 | COSV54080672 |
| 66 | T→I | missense_variant | COSMIC | — | — | damaging | likely_pathogenic (0.94) | -10.31 | COSV54080777 |
| 67 | I→V | missense_variant | gnomAD | — | 6.84e-07 | — | likely_benign (0.25) | -6.97 | chr5-139614606-A-G |
| 67 | I→V | missense_variant | ClinVar | Uncertain significance | — | — | likely_benign (0.25) | -6.97 | ClinVar:2331085 |
| 70 | P→S | missense_variant | COSMIC | — | — | damaging | likely_pathogenic (0.92) | -8.37 | COSV54077336 |
| 72 | D→D | synonymous_variant | gnomAD | — | 6.71e-05 | — | — | 0.00 | chr5-139614702-C-T |
| 73 | S→N | missense_variant | gnomAD | — | 6.84e-07 | damaging | likely_pathogenic (0.92) | -9.43 | chr5-139614704-G-A |
| 74 | P→A | missense_variant | gnomAD | — | 2.05e-06 | — | ambiguous (0.37) | -6.47 | chr5-139614706-C-G |
| 74 | P→L | missense_variant | gnomAD | — | 6.84e-07 | damaging | likely_pathogenic (0.83) | -8.31 | chr5-139614707-C-T |
| 76 | Q→E | missense_variant | COSMIC | — | — | — | likely_benign (0.11) | -7.27 | COSV54077632 |
| 77 | G→S | missense_variant | gnomAD | — | 6.84e-07 | — | likely_benign (0.17) | -3.15 | chr5-139614715-G-A |
| 78 | G→G | synonymous_variant | gnomAD | — | 6.84e-07 | — | — | 0.00 | chr5-139614720-A-G |
| 82 | L→L | synonymous_variant | gnomAD | — | 6.84e-07 | — | — | 0.00 | chr5-139614730-T-C |
| 83 | T→I | missense_variant | gnomAD | — | 6.84e-07 | damaging | likely_pathogenic (0.74) | -7.86 | chr5-139614734-C-T |
| 85 | H→R | missense_variant | COSMIC | — | — | damaging | likely_pathogenic (0.68) | -8.25 | COSV54080726 |
| 86 | F→F | synonymous_variant | COSMIC | — | — | — | — | 0.00 | COSV107280747 |
| 87 | P→L | missense_variant | COSMIC | — | — | damaging | likely_pathogenic (0.97) | -9.50 | COSV108064632 |
| 91 | P→P | synonymous_variant | gnomAD | — | 6.84e-07 | — | — | 0.00 | chr5-139614759-C-A |
| 91 | P→P | synonymous_variant | gnomAD | — | 6.84e-07 | — | — | 0.00 | chr5-139614759-C-T |
| 91 | P→S | missense_variant | COSMIC | — | — | damaging | likely_pathogenic (1.00) | -10.94 | COSV54077205 |
| 93 | K→R | missense_variant | gnomAD | — | 6.84e-07 | damaging | likely_benign (0.28) | -8.75 | chr5-139614764-A-G |
| 93 | K→R | missense_variant | COSMIC | — | — | damaging | likely_benign (0.28) | -8.75 | COSV54078386 |
| 94 | P→L | missense_variant | COSMIC | — | — | damaging | likely_pathogenic (0.98) | -10.00 | COSV105026663 |
| 97 | V→V | synonymous_variant | gnomAD | — | 6.84e-07 | — | — | 0.00 | chr5-139614863-T-C |
| 100 | T→K | missense_variant | COSMIC | — | — | damaging | likely_pathogenic (0.70) | -9.87 | COSV54080505 |
| 101 | T→T | synonymous_variant | gnomAD | — | 6.84e-07 | — | — | 0.00 | chr5-139614875-A-T |
| 102 | R→R | synonymous_variant | gnomAD | — | 2.05e-06 | — | — | 0.00 | chr5-139614876-A-C |
| 102 | R→I | missense_variant | COSMIC | — | — | damaging | likely_pathogenic (0.97) | -11.93 | COSV54077275 |
| 104 | Y→C | missense_variant | gnomAD | — | 1.37e-06 | damaging | likely_pathogenic (0.86) | -9.56 | chr5-139614883-A-G |
| 104 | Y→Y | synonymous_variant | gnomAD | — | 1.37e-06 | — | — | 0.00 | chr5-139614884-T-C |
| 104 | Y→Y | synonymous_variant | COSMIC | — | — | — | — | 0.00 | COSV105026661 |
| 105 | H→H | synonymous_variant | gnomAD | — | 6.84e-07 | — | — | 0.00 | chr5-139614887-T-C |
| 105 | H→R | missense_variant | COSMIC | — | — | damaging | likely_pathogenic (1.00) | -9.56 | COSV99551272 |
| 105 | H→N | missense_variant | COSMIC | — | — | damaging | likely_pathogenic (0.99) | -8.62 | COSV99551494 |
| 105 | H→L | missense_variant | COSMIC | — | — | damaging | likely_pathogenic (0.98) | -10.56 | COSV99551495 |
| 106 | P→P | synonymous_variant | gnomAD | — | 9.92e-05 | — | — | 0.00 | chr5-139614890-A-G |
| 107 | N→N | synonymous_variant | gnomAD | — | 6.84e-07 | — | — | 0.00 | chr5-139614893-T-C |
| 111 | N→N | synonymous_variant | gnomAD | — | 6.16e-06 | — | — | 0.00 | chr5-139614905-T-C |
| 112 | G→S | missense_variant | COSMIC | — | — | damaging | likely_pathogenic (0.99) | -8.87 | COSV54080718 |
| 112 | G→D | missense_variant | COSMIC | — | — | damaging | likely_pathogenic (1.00) | -10.19 | COSV54077127 |
| 112 | G→G | synonymous_variant | COSMIC | — | — | — | — | 0.00 | COSV99551303 |
| 113 | S→N | missense_variant | gnomAD | — | 6.84e-07 | damaging | likely_pathogenic (0.74) | -8.12 | chr5-139614910-G-A |
| 113 | S→S | synonymous_variant | COSMIC | — | — | — | — | 0.00 | COSV54079119 |
| 119 | L→L | synonymous_variant | gnomAD | — | 6.84e-07 | — | — | 0.00 | chr5-139614927-C-T |
| 119 | L→L | synonymous_variant | gnomAD | — | 8.89e-06 | — | — | 0.00 | chr5-139614929-A-G |
| 120 | R→R | synonymous_variant | gnomAD | — | 7.53e-06 | — | — | 0.00 | chr5-139614932-A-G |
| 120 | R→* | stop_gained | COSMIC | — | — | LoF | — | — | COSV54077789 |
| 120 | R→Q | missense_variant | COSMIC | — | — | damaging | likely_pathogenic (0.88) | -9.87 | COSV54081074 |
| 121 | S→P | missense_variant | gnomAD | — | 6.84e-07 | damaging | likely_pathogenic (0.91) | -11.00 | chr5-139614933-T-C |
| 121 | S→S | synonymous_variant | gnomAD | — | 1.37e-06 | — | — | 0.00 | chr5-139614935-A-T |
| 122 | Q→Q | synonymous_variant | gnomAD | — | 1.37e-06 | — | — | 0.00 | chr5-139614938-G-A |
| 122 | Q→H | missense_variant | COSMIC | — | — | damaging | likely_pathogenic (0.92) | -7.87 | COSV99551304 |
| 123 | W→C | missense_variant | COSMIC | — | — | damaging | likely_pathogenic (1.00) | -9.81 | COSV54081180 |
| 124 | S→S | synonymous_variant | gnomAD | — | 1.37e-06 | — | — | 0.00 | chr5-139614944-T-C |
| 124 | S→F | missense_variant | COSMIC | — | — | damaging | likely_pathogenic (1.00) | -11.81 | COSV108786173 |
| 127 | L→L | synonymous_variant | gnomAD | — | 6.85e-07 | — | — | 0.00 | chr5-139614951-C-T |
| 127 | — | frameshift_variant | COSMIC | — | — | LoF | — | — | COSV54078283 |
| 128 | T→I | missense_variant | gnomAD | — | 6.85e-07 | damaging | likely_pathogenic (1.00) | -10.62 | chr5-139614955-C-T |
| 128 | T→I | missense_variant | COSMIC | — | — | damaging | likely_pathogenic (1.00) | -10.62 | COSV54079667 |
| 130 | S→L | missense_variant | gnomAD | — | 6.85e-07 | damaging | likely_pathogenic (0.96) | -9.62 | chr5-139614961-C-T |
| 130 | S→S | synonymous_variant | gnomAD | — | 3.43e-06 | — | — | 0.00 | chr5-139614962-A-C |
| 130 | S→S | synonymous_variant | gnomAD | — | 6.85e-07 | — | — | 0.00 | chr5-139614962-A-T |
| 134 | L→L | synonymous_variant | gnomAD | — | 2.09e-06 | — | — | 0.00 | chr5-139623373-T-C |
| 135 | — | frameshift_variant | gnomAD | — | 6.97e-07 | LoF | — | — | chr5-139623377-CCATCTGTT-C |
| 137 | C→C | synonymous_variant | gnomAD | — | 6.90e-07 | — | — | 0.00 | chr5-139623384-T-C |
| 138 | S→T | missense_variant | gnomAD | — | 6.90e-07 | damaging | likely_pathogenic (0.91) | -11.44 | chr5-139623385-T-A |
| 138 | S→F | missense_variant | COSMIC | — | — | damaging | likely_pathogenic (0.99) | -10.75 | COSV105026656 |
| 139 | L→L | synonymous_variant | gnomAD | — | 6.89e-07 | — | — | 0.00 | chr5-139623390-G-C |
| 139 | L→L | synonymous_variant | COSMIC | — | — | — | — | 0.00 | COSV54077579 |
| 140 | L→L | synonymous_variant | gnomAD | — | 8.91e-03 | — | — | 0.00 | chr5-139623391-T-C |
| 140 | L→L | synonymous_variant | COSMIC | — | — | — | — | 0.00 | COSV54079502 |
| 141 | C→C | synonymous_variant | gnomAD | — | 6.88e-07 | — | — | 0.00 | chr5-139623396-T-C |
| 143 | P→A | missense_variant | gnomAD | — | 6.88e-07 | damaging | likely_pathogenic (0.61) | -4.52 | chr5-139623400-C-G |
| 143 | P→P | synonymous_variant | gnomAD | — | 6.88e-07 | — | — | 0.00 | chr5-139623402-C-G |
| 144 | N→D | missense_variant | COSMIC | — | — | damaging | likely_pathogenic (0.98) | -10.12 | COSV54077287 |
| 146 | D→V | missense_variant | gnomAD | — | 6.88e-07 | damaging | likely_pathogenic (0.95) | -11.31 | chr5-139623410-A-T |
| 146 | D→D | synonymous_variant | gnomAD | — | 6.88e-07 | — | — | 0.00 | chr5-139623411-T-C |
| 147 | D→Y | missense_variant | COSMIC | — | — | damaging | likely_pathogenic (0.99) | -10.12 | COSV99551400 |
| 148 | P→P | synonymous_variant | gnomAD | — | 6.89e-07 | — | — | 0.00 | chr5-139623417-T-C |
| 148 | P→P | synonymous_variant | gnomAD | — | 6.89e-07 | — | — | 0.00 | chr5-139623417-T-G |
| 149 | L→L | synonymous_variant | gnomAD | — | 6.89e-07 | — | — | 0.00 | chr5-139623420-A-G |
| 149 | L→* | stop_gained | COSMIC | — | — | LoF | — | — | COSV99551309 |
| 151 | P→T | missense_variant | gnomAD | — | 6.89e-07 | damaging | likely_pathogenic (0.92) | -8.19 | chr5-139623424-C-A |
| 151 | P→P | synonymous_variant | gnomAD | — | 6.89e-07 | — | — | 0.00 | chr5-139623426-T-A |
| 151 | P→L | missense_variant | COSMIC | — | — | damaging | likely_pathogenic (0.93) | -8.00 | COSV54077732 |
| 152 | E→K | missense_variant | gnomAD | — | 6.89e-07 | damaging | likely_pathogenic (0.95) | -10.05 | chr5-139623427-G-A |
| 152 | E→E | synonymous_variant | gnomAD | — | 6.89e-07 | — | — | 0.00 | chr5-139623429-G-A |
| 153 | I→M | missense_variant | gnomAD | — | 6.89e-07 | damaging | likely_pathogenic (0.72) | -8.19 | chr5-139623432-T-G |
| 153 | I→V | missense_variant | ClinVar | Uncertain significance | — | — | ambiguous (0.37) | -6.94 | ClinVar:3976716 |
| 154 | A→T | missense_variant | gnomAD | — | 6.89e-07 | damaging | likely_pathogenic (0.98) | -8.75 | chr5-139623433-G-A |
| 154 | A→G | missense_variant | gnomAD | — | 6.89e-07 | — | ambiguous (0.45) | -7.15 | chr5-139623434-C-G |
| 154 | A→A | synonymous_variant | gnomAD | — | 6.89e-07 | — | — | 0.00 | chr5-139623435-T-C |
| 155 | R→R | synonymous_variant | gnomAD | — | 6.89e-07 | — | — | 0.00 | chr5-139623436-C-A |
| 155 | R→G | missense_variant | gnomAD | — | 6.89e-07 | damaging | likely_pathogenic (0.60) | -7.87 | chr5-139623436-C-G |
| 155 | R→Q | missense_variant | gnomAD | — | 3.45e-06 | — | ambiguous (0.34) | -6.62 | chr5-139623437-G-A |
| 155 | R→Q | missense_variant | COSMIC | — | — | — | ambiguous (0.34) | -6.62 | COSV54080979 |
| 156 | I→L | missense_variant | gnomAD | — | 6.88e-07 | — | likely_benign (0.14) | -5.20 | chr5-139623439-A-C |
| 157 | Y→C | missense_variant | gnomAD | — | 6.89e-07 | — | likely_benign (0.33) | -6.65 | chr5-139623443-A-G |
| 159 | T→T | synonymous_variant | gnomAD | — | 6.89e-07 | — | — | 0.00 | chr5-139623450-A-C |
| 159 | T→I | missense_variant | COSMIC | — | — | — | ambiguous (0.50) | -6.29 | COSV107280745 |
| 161 | R→G | missense_variant | gnomAD | — | 6.90e-07 | — | ambiguous (0.48) | -6.93 | chr5-139623454-A-G |
| 161 | R→T | missense_variant | gnomAD | — | 6.90e-07 | — | ambiguous (0.43) | -6.49 | chr5-139623455-G-C |
| 162 | E→Q | missense_variant | gnomAD | — | 6.90e-07 | damaging | likely_benign (0.21) | -7.67 | chr5-139623457-G-C |
| 162 | E→K | missense_variant | COSMIC | — | — | damaging | ambiguous (0.39) | -8.05 | COSV54077797 |
| 164 | Y→Y | synonymous_variant | gnomAD | — | 6.84e-07 | — | — | 0.00 | chr5-139626759-C-T |
| 166 | R→G | missense_variant | gnomAD | — | 1.37e-06 | — | likely_benign (0.21) | -7.18 | chr5-139626763-A-G |
| 167 | I→V | missense_variant | gnomAD | — | 6.84e-07 | — | likely_benign (0.09) | -3.55 | chr5-139626766-A-G |
| 168 | A→V | missense_variant | COSMIC | — | — | damaging | likely_pathogenic (0.87) | -5.52 | COSV54079876 |
| 169 | R→W | missense_variant | gnomAD | — | 2.05e-06 | damaging | ambiguous (0.49) | -8.05 | chr5-139626772-C-T |
| 169 | R→Q | missense_variant | gnomAD | — | 1.37e-06 | — | likely_benign (0.10) | -2.28 | chr5-139626773-G-A |
| 169 | R→L | missense_variant | gnomAD | — | 6.84e-07 | — | ambiguous (0.37) | -6.48 | chr5-139626773-G-T |
| 169 | R→P | missense_variant | ClinVar | Uncertain significance | — | damaging | likely_pathogenic (0.99) | -9.92 | ClinVar:3330554 |
| 169 | R→Q | missense_variant | COSMIC | — | — | — | likely_benign (0.10) | -2.28 | COSV99551584 |
| 170 | E→K | missense_variant | gnomAD | — | 6.84e-07 | damaging | likely_pathogenic (0.77) | -8.75 | chr5-139626775-G-A |
| 170 | E→K | missense_variant | COSMIC | — | — | damaging | likely_pathogenic (0.77) | -8.75 | COSV54080517 |
| 172 | T→A | missense_variant | COSMIC | — | — | damaging | likely_pathogenic (0.66) | -6.23 | COSV54077550 |
| 175 | Y→C | missense_variant | gnomAD | — | 6.84e-07 | damaging | likely_pathogenic (0.88) | -8.25 | chr5-139626791-A-G |
| 175 | Y→F | missense_variant | gnomAD | — | 1.37e-06 | — | likely_benign (0.20) | -5.93 | chr5-139626791-A-T |
| 175 | Y→Y | synonymous_variant | gnomAD | — | 6.84e-07 | — | — | 0.00 | chr5-139626792-T-C |
| 176 | A→A | synonymous_variant | gnomAD | — | 1.10e-05 | — | — | 0.00 | chr5-139626795-G-A |
| 177 | M→T | missense_variant | gnomAD | — | 6.85e-07 | — | ambiguous (0.51) | -7.30 | chr5-139626797-T-C |
| 177 | M→T | missense_variant | ClinVar | Uncertain significance | — | — | ambiguous (0.51) | -7.30 | ClinVar:4601227 |
167 variants in the shared canonical core.
LoF = frameshift / stop-gain / splice-disrupting — inherently loss-of-function, flagged by consequence (AlphaMissense and ESM-C score only missense/substitutions, so they are blank here by design, not by absence of impact). AlphaMissense is computed in the canonical reading frame, so it scores the shared core but reads N/A across an isoform-unique extension — use ESM-C ΔLLR there.