CBX1
EXTENDED 239 aa (canonical 185 aa) · UniProt P83916 · CDLMPS
chr17:48101392:-:GTG:ENST00000225603.9
AI summary A 55-aa N-terminal extension folds into a confident but unintegrated strand and shifts whole-protein hydropathy, without touching the chromodomain reader function.
The extension adds a 29-residue strand (pLDDT 0.83) that folds independently but shows essentially no contact with the shared CBX1 core (PAE ~30 Å to the body), so it reads as a dangling appendage rather than an integrated module; in parallel it shifts the whole protein toward higher hydropathy and lower fraction-charged, and a strong interpretable-feature shift accompanies it. Localization and domain architecture are unchanged: DeepLoc keeps both isoform and canonical nuclear (top_prob dropping from 0.93 to 0.80 but no compartment flip), and no real InterPro domain — including no chromodomain footprint — overlaps the added sequence. The large shared-region RMSD (9.73 Å) is not trustworthy as a core refold given low global pTM (0.34–0.39), so no genuine reorganization of the retained fold is supported.
CBX1's function rests entirely on its chromodomain reading H3K9me3 and its downstream partnerships (SUV39H1, PurB/Sp3, PRC2, KAP-1/Chk2), all located in the shared, unchanged C-terminal/core region — this extension neither gains nor loses a domain that would touch any of that machinery, and it does not relocalize the protein out of the nucleus. The added structured-but-disconnected strand and altered hydropathy/charge balance could plausibly tune solubility, aggregation propensity, or a novel unannotated interaction surface, but nothing in the differential findings ties this to H3K9me3 reading, heterochromatin nucleation, or DNA-damage-induced mobilization — the isoform's known biology is essentially untouched by this extension based on current evidence.
The shared-region RMSD (9.73 Å) is likely placement uncertainty given low pTM (0.34-0.39), not tagged as a core refold. Extension unique-region conservation is real (phyloP 2.79, primate 91%) supporting translation as a genuine ORF, but germline/disease-variant signals are uninformative here (extension was intronic) and were excluded from the readout per instructions.
Folding
Evidence — click any tile for the differential-region detail
LLM reasoning
How well is the isoform-unique region conserved at the amino-acid level across primates (mean percent identity)? High AA identity in close relatives argues the alternative protein is real and translated, not a sequencing or annotation artifact. (Reading-frame intactness is shown as context.)
Comparison
Sequence conservation · primate
mean AA identity over the differential ORF is the score basis; frame intactness across the clade is context
| Conservation metric | Canonical ORF | Differential ORF |
|---|---|---|
| Mean AA identity | 100% | 91% |
| Frame intact (fraction of species) | 84% | 80% |
| Species aligned | 25 | 25 |
| Species frame-intact | 21 | 20 |
| Start codon conserved | 100% | 72% |
| Deepest intact species | Propithecus_coquereli | Otolemur_garnettii |
| Phylo depth (MRCA) | 7 | 7 |
The same amino-acid identity test across mammals. Conservation over deeper evolutionary distance is stronger evidence the ORF is under selection to be translated. (Reading-frame intactness is shown as context.)
Comparison
Sequence conservation · mammalian
mean AA identity over the differential ORF is the score basis; frame intactness across the clade is context
| Conservation metric | Canonical ORF | Differential ORF |
|---|---|---|
| Mean AA identity | 98% | 84% |
| Frame intact (fraction of species) | 15% | 35% |
| Species aligned | 20 | 20 |
| Species frame-intact | 3 | 7 |
| Start codon conserved | 100% | 67% |
| Deepest intact species | Callithrix_jacchus | Loxodonta_africana |
| Phylo depth (MRCA) | 6 | 12 |
phyloP / phastCons measure per-base evolutionary constraint. A high absolute phyloP over the differential region means the sequence itself is under strong purifying selection — evidence it is coding and does something. (Enrichment over the shared core is shown as context only.)
Comparison
Per-base conservation · differential region
absolute mean phyloP over the differential region is the score basis (high = strong purifying selection); the shared column and enrichment are context, not the claim
| Track | Differential | Shared | vs shared |
|---|---|---|---|
| phyloP mean | 2.79 | 4.95 | 0.563 |
| phastCons mean | 0.692 | 0.92 | — |
Start site · canonical vs isoform
initiation context + per-base conservation at each start codon
| Property | Canonical | Isoform |
|---|---|---|
| Start codon | ATG | GTG |
| Kozak context (−9..+4) | GCGGGCACTATGG | GCCGCTTCAGTGA |
| phyloP at start codon | 6.43 | 1.04 |
| phastCons at start codon | 1 | 0.00533 |
| phyloP over Kozak window | 6.31 | 3.11 |
| phastCons over Kozak window | 1 | 0.671 |
| Kozak mismatch — full consensus | 3 | 5 |
| Kozak window GC content | 0.692 | 0.615 |
LLM reasoning
Is the alternative start used in more than one cell line? Reproducible initiation across independent samples argues against a one-off ribosome artifact.
Comparison
Per-cell-line usage · canonical vs isoform
Fisher q = 1.33e-12
| Cell line | Canonical | Isoform | p-value |
|---|---|---|---|
| HeLa | — | 3.1 | 0.000481 |
| K562 | — | 1.67 | 1.33e-12 |
| U2OS | — | 1.4 | 9.26e-08 |
| RPE1 Async | — | 1.88 | 1.29e-06 |
| RPE1 Que | — | 0.316 | 4.97e-09 |
How efficiently ribosomes initiate at this start (TIS) relative to background, per cell line. Strong, reproducible initiation supports a genuine translation event.
Comparison
Start-Site Usage · canonical vs isoform
ribosome initiation efficiency at the canonical start vs this alternative start, per cell line
| Cell line | Canonical | Isoform |
|---|---|---|
| HeLa | — | 0.0846 |
| K562 | — | 0.0767 |
| U2OS | — | 0.0215 |
| RPE1 Async | — | 0.0446 |
| RPE1 Que | — | 0.0155 |
Were tryptic peptides unique to the differential region detected by mass spec? Direct peptide evidence is the strongest proof the isoform protein exists.
Comparison
Peptide Evidence (canonical vs isoform)
isoform-unique peptides are direct evidence the alternative protein exists
| Feature | Canonical | Isoform |
|---|---|---|
| Tryptic peptides (in-silico) | 28 | 11 |
| Validated by mass-spec | 0 | 1 |
| Isoform-unique peptides | — | 11 |
Details
Peptide Evidence (canonical vs isoform)
- peptide DATAATR 2–9
- peptide LAAFLGATPPGDPTR 9–24
- peptide ASSAAPIPLGLLGAALSSVTLYTR 25–49
- peptide LAGTMGK 50–57
- peptide MRDATAATR 0–9
- validated DATAATRLAAFLGATPPGDPTR 2–24
- peptide LAAFLGATPPGDPTRR 9–25
- peptide RASSAAPIPLGLLGAALSSVTLYTR 24–49
- peptide ASSAAPIPLGLLGAALSSVTLYTRK 25–50
- peptide KLAGTMGK 49–57
- peptide LAGTMGKK 50–58
LLM reasoning
Do the predicted localization features (DeepLoc prediction, sorting signals, or membrane association) differ between the canonical and the isoform? A protein with changed localization features acts in a different cellular context.
Comparison
Subcellular localization (DeepLoc)
no localization-feature change
| Property | Canonical | Isoform |
|---|---|---|
| Predicted location | Nucleus | Nucleus |
| Sorting signals | Nuclear localization signal | Nuclear localization signal |
| Membrane | Soluble | Soluble |
Do N-terminal targeting signals (SignalP secretion, TargetP mitochondrial/chloroplast) differ between canonical and isoform? N-terminal changes most directly add or remove targeting peptides.
LLM reasoning
Does healthy human germline variation (gnomAD) avoid this region (depletion ratio < 1×), and is it intrinsically constrained (ESM-C constraint delta > 0)? Depletion of population variation plus high sequence constraint mean the region resists change — it is functionally important. Scored only where the unique region is canonical coding sequence, i.e. on truncations. gnomAD is a tolerance catalogue, not a disease one; disease/cancer variants (ClinVar / COSMIC) live in M2.
Comparison
Germline Variants · differential vs shared region
gnomAD (population) variant density per nucleotide — depletion ratio < 1× means healthy human variation avoids the region (constrained), the M1 basis alongside ESM-C constraint
| Variant set | Differential | Shared | Depletion ratio |
|---|---|---|---|
| gnomAD variants | 86 | 192 | 1.5× |
Sequence constraint (ESM-C) · differential vs shared region
lower LLR = more constrained; enrichment = differential / conserved
| Property | Differential | Shared | Enrichment |
|---|---|---|---|
| ESM-C mean LLR | -1.95 | -0.00473 | — |
| Constrained positions | 0 | 0 | — |
Predicted-damaging variants · germline (gnomAD)
AlphaMissense / ESM-C / LoF calls per region (length-normalized)
| Variant set | Differential | Shared | Enrichment |
|---|---|---|---|
| Scorable variants | 84 | 188 | 1.5× |
| Damaging variants | 6 | 103 | 0.2× |
| — of which loss-of-function | 6 | 13 | 1.6× |
| AlphaMissense-pathogenic | 0 | 58 | 0× |
Predictor scores · germline (gnomAD)
scored: 336 ESM-C · 168 AlphaMissense
| Property | Differential | Shared |
|---|---|---|
| Mean ΔLLR (ESM-C) | -0.723 | -5.64 |
| Min ΔLLR (ESM-C) | -6.18 | -13.1 |
| Mean AlphaMissense | — | 0.578 |
Are disease (ClinVar / COSMIC) variants enriched per nucleotide in the differential region versus the shared core (disease enrichment ratio ≥ 1×)? Disease variants concentrating in the unique region tie it to phenotype.
Comparison
Clinical-variant burden · differential vs shared region
counts per region; ratio is density-normalized (variants per nucleotide, differential ÷ shared — ≥1× = disease variants concentrate in the differential region, the M2 basis)
| Variant set | Differential | Shared | Enrichment |
|---|---|---|---|
| Disease variants | 4 | 89 | 0.15× |
| Pathogenic | 0 | 0 | — |
Predicted-damaging variants · disease (ClinVar/COSMIC)
AlphaMissense / ESM-C / LoF calls per region (length-normalized)
| Variant set | Differential | Shared | Enrichment |
|---|---|---|---|
| Scorable variants | 4 | 88 | 0.16× |
| Damaging variants | 0 | 66 | 0× |
| — of which loss-of-function | 0 | 9 | 0× |
| AlphaMissense-pathogenic | 0 | 43 | 0× |
Predictor scores · disease (ClinVar/COSMIC)
scored: 336 ESM-C · 168 AlphaMissense
| Property | Differential | Shared |
|---|---|---|
| Mean ΔLLR (ESM-C) | -2.33 | -7.53 |
| Min ΔLLR (ESM-C) | -5.7 | -13.9 |
| Mean AlphaMissense | — | 0.68 |
LLM reasoning
Does the gained/lost region fold into ordered structure (pLDDT) and shift the protein's biophysical character (charge, hydropathy, disorder)? Structured, biophysically distinct regions are more likely to be functional.
Comparison
Structure (ESMFold2) · canonical vs isoform
TM-score 0.515 · RMSD 4.79 Å · 29 interface contacts
| Metric | Canonical | Isoform |
|---|---|---|
| pLDDT (whole protein) | 0.822 | 0.841 |
Fold confidence · differential vs shared region
mean ESMFold2 pLDDT in each region
| Metric | Differential | Shared | Enrichment |
|---|---|---|---|
| pLDDT | 0.762 | 0.72 | 1.1 |
The shared region is the stretch of protein identical in both the isoform and the canonical (the canonical body for an extension; the post-truncation body for a truncation). Folded in both contexts it normally comes out nearly identical, so its Cα backbone RMSD ≈ 0. A high shared-region RMSD (superposed on the shared residues only, from the ESMFold2 structures) means the extension or truncation reorganizes how the retained region folds — a rare, high-interest functional signal. TM-score is a length-normalized companion; the check only fires when both structures are confidently folded, and uORF/altORF isoforms (no shared region) are not evaluable.
Comparison
Shared-region structural change
Cα RMSD superposed on the shared residues only; TM-score is length-normalized · shared-region Cα RMSD 9.73 Å · shared TM-score 0.565 · shared region 185 aa · min shared pLDDT 0.822 · global TM-score 0.515 · global RMSD 4.79 Å
| Property | Canonical | Isoform |
|---|---|---|
| Shared-region pLDDT | 0.822 | 0.864 |
Does the differential region contain actual secondary structure — a helix or strand — rather than coil? ESMFold2 predicts coordinates but no secondary structure, so elements are assigned from those coordinates (P-SEA, Cα geometry). Because they are derived from a prediction, each element carries its own mean pLDDT: a geometrically clean helix running through a disordered stretch is geometry fitted to a guess, so BOTH length and confidence are required to score. The direction differs by ORF type — an extension GAINS the element (a candidate functional addition), a truncation LOSES one, and there the element is read off the canonical structure because the removed segment exists only in it. This says nothing about whether the element is integrated with the rest of the fold; the contact and PAE evidence in this same category answers that.
Comparison
Secondary structure · differential vs shared region
helices and strands assigned from the predicted coordinates (P-SEA)
| Metric | Differential | Shared |
|---|---|---|
| Alpha helices | 0 | 3 |
| Beta strands | 1 | 2 |
| Longest element (aa) | 29 | 18 |
| Mean pLDDT | 0.83 | 0.90 |
Elements and coordinates
1 in the differential region, 5 in the shared core — residue numbering is 1-based on the protein holding the region
| Isoform-unique | Shared core |
|---|---|
| beta strand 3–31 29 aa · pLDDT 0.83 | alpha helix 115–132 18 aa · pLDDT 0.87 |
| — | beta strand 148–154 7 aa · pLDDT 0.81 |
| — | beta strand 197–202 6 aa · pLDDT 0.93 |
| — | alpha helix 203–209 7 aa · pLDDT 0.96 |
| — | alpha helix 211–221 11 aa · pLDDT 0.96 |
Below threshold
0 in the differential region, 5 in the shared core — shorter than 6 aa or below pLDDT 0.70, so not counted above
| Isoform-unique | Shared core |
|---|---|
| — | beta strand 55–65 11 aa · pLDDT 0.68 |
| — | beta strand 74–77 4 aa · pLDDT 0.92 |
| — | beta strand 89–93 5 aa · pLDDT 0.96 |
| — | beta strand 104–107 4 aa · pLDDT 0.96 |
| — | beta strand 186–189 4 aa · pLDDT 0.95 |
LLM reasoning
Does the isoform gain or lose a real InterPro functional domain in the differential region (disorder/structural-only signatures excluded)? Gaining or losing a domain changes function directly.
Comparison
Domains & motifs (canonical vs isoform)
gained = only in the isoform; lost = only in the canonical
| Feature | Canonical | Isoform |
|---|---|---|
| InterPro domains | 15 | 16 |
| Short linear motifs | 1 | 3 |
Details
Domains & motifs (canonical vs isoform)
- gained Coil domain
- gained CDK_Sites_SP_TP motif
- gained SH3_ClassII motif
Biophysical character of the isoform-differential region versus the shared canonical core — pI, hydropathy, charge, disorder and related properties. Descriptive; the folding (P1) score keys off the GRAVY / charge / disorder deltas.
Comparison
Biophysics · differential vs shared region
highlighted rows are enriched in the differential region
| Property | Differential | Shared | Enrichment |
|---|---|---|---|
| Isoelectric point (pI) | 11.9 | 4.54 | 2.63 |
| Hydropathy (GRAVY) | 0.233 | -1.28 | -0.182 |
| Fraction charged | 0.148 | 0.46 | 0.322 |
| Disorder fraction | 0.0984 | 0.232 | 0.425 |
| Disorder-promoting | 0.63 | 0.681 | 0.924 |
| Low-complexity fraction | 0.241 | 0.222 | 1.09 |
| Prion-like fraction | 0.185 | 0.2 | 0.926 |
| LLPS score | 0.133 | 0.192 | 0.693 |
| π–π propensity | 0.13 | 0.173 | 0.749 |
| Aromaticity | 0.037 | 0.0703 | 0.526 |
| Instability index | 23.3 | 53 | 0.44 |
| Shannon entropy | 3.32 | 3.9 | 0.852 |
| Normalized complexity | 0.768 | 0.902 | 0.852 |
Sparse-autoencoder (SAE) interpretability features on the ESM-C residual stream, comparing the isoform against the canonical protein. This is the S3 criterion — a presence check that fires when any interpretable feature is gained or lost. Only the top 30 features per category (ranked by activation, prevalence ≥ 2) are saved; each table shows the top 10 interpretable features, and generic / "unknown" features are hidden unless you expand it. What are SAE features? →
Part 1 · Global feature comparison
Which SAE features fire in the isoform vs the canonical protein, across the whole sequence.
Isoform-only features — 176 total (prevalence ≥ 2)
| Feature | Label | Description | Activation | Prevalence |
|---|---|---|---|---|
| #4992 | Small-residue disordered N-terminus | Context-dependent free N-terminus signature: short, disordered amino-terminal stretches beginning with the initiator Met followed by residues compatible with MAP cleavage or N-acetylation (e.g., M–G/A/S/T/V, and other contexts including M–D/E/L); strongest emphasis on small residues at positions 2–3; activation is absent when the native N-terminus is structured or starts with a signal peptide/transmembrane segment. | 12.53 | 2 |
| #448 | N-terminal leader/targeting segments | N-terminal leader/targeting segments: the feature activates on the extreme N-terminus (first 2–5 residues) of proteins, especially the N-region of signal peptides and organelle transit peptides, and more generally on short, disordered/low-complexity N-terminal tails; strongest at residue ~3 and largely independent of amino-acid identity, occurring across all taxa and functions. | 12.04 | 3 |
| #11504 | Initiator methionine at M1 | Initiator methionine at the very start of the polypeptide chain (M1), i.e., the translation start residue, independent of protein family, taxonomy, membrane association, or presence of signal/leader/propeptide regions; often annotated as post-translationally removed and typically situated in a flexible, coil-like N-terminus. | 11.12 | 2 |
| #8545 | N-terminus accessibility sensor | Detector of accessible peptide chain termini—primarily the extreme N-terminus (initiator methionine and immediate neighbors) in flexible, unstructured tails; position-specific rather than residue-specific—with occasional weak recognition of the C-terminus; common but not universal. | 10.51 | 2 |
| #9852 | N-terminus activation; hydrophobic helices secondary | Extreme N-terminal regions of proteins, starting at the initiator methionine and extending over a short stretch, with occasional secondary activation at internal hydrophobic/amphipathic helices. | 10.44 | 53 |
| #14493 | Short hydrophobic helical runs | Short hydrophobic helical segments rich in I/V/F (sometimes M), recognized in both membrane-targeting/insertion contexts (signal peptides, signal-anchors, transmembrane helices) and in hydrophobic helical patches within otherwise soluble small proteins. The feature flags contiguous aliphatic/hydrophobic runs across diverse proteins, including viral membrane/accessory proteins, secreted peptide/effector precursors, small ORFs/microproteins, and short basic/uncharacterized human proteins. | 10.31 | 4 |
| #3609 | Position-2 N-terminus residue detector | Detector of the identity of the second residue at the extreme N-terminus of proteins, with strongest preference for small residues (especially Ser, Ala, and Gly; also Thr/Pro), i.e., an N-terminal position-2 signal commonly associated with generic, disordered starts and co-translational N-terminus processing | 10.25 | 3 |
| #9998 | Weak N+3 positional cue | Absolute N-terminal positional cue centered near the fourth residue (N+3), independent of amino-acid identity; most often within signal/transit peptides of precursors but also present in non-secreted proteins; the effect is weak and often does not produce a discrete detected site. | 9.58 | 2 |
| #7903 | Arginine-rich polybasic patches | Basic polycationic patches enriched in arginine (often with lysine) within low-complexity/disordered regions, frequently N-terminal. These include classical monopartite NLSs, nucleic acid–binding basic tails, signal peptide n-regions and cytosolic juxtamembrane segments (positive-inside rule), and basic clusters in secreted precursors; typically depleted of acidic residues. | 9.52 | 7 |
| #12385 | Aromatic F/Y cluster motif | Aromatic (phenylalanine/tyrosine) cluster motif: short, hydrophobic low‑complexity segments enriched for F/Y that act as membrane‑interface anchors or aromatic “sticker” patches, most often at N‑termini or adjacent to transmembrane helices; recurrent in small viral/accessory proteins, secreted effectors and toxin/antimicrobial peptides, micropeptides from alt/lncRNA ORFs, and some nucleic‑acid–associated proteins. | 9.51 | 2 |
| #7384 | Leucine-biased disordered hydrophobic segments | Leucine-biased recognition of intrinsically disordered, low-complexity hydrophobic segments, including disordered insertions within otherwise structured proteins. | 9.24 | 8 |
| #11159 | Glycine residue identity detector | Residue-identity detector for glycine (G), largely context-agnostic but with a tendency to fire on glycines in flexible/disordered, often N-terminal or membrane-proximal regions; also active on glycines within transmembrane helices; broadly taxon- and function-independent. | 8.49 | 4 |
| #5337 | Pro/RGG-rich low-complexity IDRs | Intrinsically disordered, low‑complexity basic segments at termini and long loops, enriched in Pro/Gly and/or Arg/Ser (including RG/RGG or S/T clusters), common across eukaryotic and viral proteins. | 8.06 | 47 |
| #3670 | Terminal disordered peptide detector | Short unstructured peptides and N-terminal/C-terminal segments of larger proteins; occasional firing near N-terminal leader/signal sequences and basic targeting motifs (e.g., NLS); broadly residue-tolerant with bias toward hydrophobic, Gly/Pro, and basic residues; appears in viral accessory proteins, microproteins, and secreted precursors but also in diverse bacterial/archaeal enzymes, with sparse activation overall. | 7.86 | 11 |
| #14646 | Short N-terminal leader segment | Short N-terminal segments within the first ~10–40 amino acids, often Lys/Arg-enriched but sometimes purely hydrophobic. These commonly correspond to Sec-type signal peptides, signal anchors, or other short N-terminal segments preceding/overlapping the first hydrophobic helix, and also include disordered N-terminal tails in soluble proteins. | 7.34 | 37 |
| #5241 | Early N-terminal positional signal | Generic early N-terminus positional signal peaking at residue ~5–7, found broadly within the disordered N-tail of proteins, including both cleavable leader/targeting segments (signal peptides, propeptides) of precursor proteins and disordered N-terminal regions of mature proteins; residue identity is not specific | 7.30 | 2 |
| #4913 | Low-complexity disordered segments | Low-complexity and compositionally biased sequence stretches, frequently in disordered or flexible regions including N/C-terminal tails, but also extending to short low-complexity tracts within otherwise ordered proteins; peaks favor Ser and other polar residues, with additional hits on aromatic (Trp/Phe/Tyr) and basic-rich contexts, suggesting a composition/local-context signal rather than a specific fold or sharply defined motif | 6.97 | 2 |
| #6826 | Sparse internal aliphatic peaks | Sparse residue-level feature firing at scattered internal positions across diverse proteins, with peaks often landing on aliphatic/small residues (A, P, V, I, L) flanked by polar or charged context, frequently in loop-like or low-complexity stretches. | 6.91 | 3 |
| #8488 | Ala/Thr-rich N-terminal disorder | Ala/Thr-enriched composition feature with a bias for low-complexity intrinsically disordered regions and N-terminal prepro/signal-peptide segments; the feature reflects small-residue (A/T, secondarily S/P) composition often in flexible regions but also appears at A/T residues in some structured contexts. | 6.71 | 17 |
| #15324 | Low-complexity disordered segments | A detector of low‑complexity, intrinsically disordered, Ser/Thr/Gly/Pro/Ala‑rich segments (often N‑terminal), including proline‑rich stretches and occasional closely spaced cysteines; these regions occur in viral microproteins, secretory precursors/extracellular repeats, membrane‑proximal tails, and organelle transit peptides rather than in a specific folded domain. | 6.36 | 9 |
| #9234 | Charged low-complexity disordered regions | Intrinsically disordered, low-complexity segments enriched for Lys and acidic residues (E/D), plus N/Q, that include both polyampholyte and basic-residue tracts; typically flexible terminal tails and linker regions in diverse eukaryotic and viral proteins, including many small/secreted proteins and membrane-protein loops | 6.23 | 4 |
| #13063 | Serine-rich region detector | Serine residues across a wide range of structural contexts, with a bias toward serine-rich segments in intrinsically disordered or low-complexity regions (often N-terminal tails), but also firing on serines within transmembrane helices and other structured contexts; the feature behaves primarily as a serine detector with some preference for S-rich patches. | 6.19 | 4 |
| #3123 | Low-complexity IDRs and signal peptides | Low-complexity, intrinsically disordered segments enriched in Pro, Ser/Thr and basic residues: the feature prefers compositionally biased coils and polar/low-complexity stretches, and also fires on hydrophobic LVIA-rich segments such as the h-region of signal peptides; it activates broadly across taxa in secreted/membrane proteins, small regulatory peptides, and flexible disordered regions of soluble enzymes. | 5.86 | 8 |
| #10748 | Disordered alanine-rich tandem repeats | Low-complexity, intrinsically disordered tandem-repeat tracts enriched in small residues (Ala/Thr/Val with Pro/Gly), often alanine-rich, found in extracellular adhesin/mucin-like regions and in flexible termini/linkers; the feature marks periodic Ala positions within such repeats and is largely absent from folded catalytic cores. | 5.82 | 5 |
| #5021 | Polybasic intrinsically disordered regions | Low-complexity, often intrinsically disordered regions in short proteins, precursors, and microproteins across taxa, including viral accessory proteins, neuropeptide/hormone precursors, sperm nuclear proteins, flexible enzyme tails/inserts, and antisense-derived/uncharacterized microproteins. Activation is biased toward, but not restricted to, basic (Lys/Arg) clusters, with frequent peaks also at Pro/Ser/Gly residues in disordered context. | 5.68 | 6 |
| #15193 | Basic-rich disordered low-complexity tails | Short intrinsically disordered, low‑complexity segments—often terminal tails or propeptide-like stretches—enriched in small/polar residues (S/T/G/A) and Pro, frequently containing charged patches (Lys/Arg clusters) and occasional Cys; these regions occur in small/viral proteins and in secreted or membrane-associated proteins, while structured domains and transmembrane helices remain largely unactivated. | 5.61 | 19 |
| #4579 | Acidic polar disordered tails | Low-complexity, disorder-associated segments enriched in Asp/Asn/His/Tyr and depleted of Lys/Arg—frequently N- or C-terminal propeptides and regulatory tails, and occasionally short polar/acidic stretches within other contexts including amphipathic helices in small viral proteins. | 5.59 | 3 |
| #3278 | PRQH-rich disordered tails | Compositionally biased, intrinsically disordered low‑complexity segments enriched in Pro/Arg/Gln/His (frequent PR/PQ tracts, Arg- and His-clusters), with occasional sensitivity to Leu‑rich helical stretches (signal peptides or leucine zippers); typically terminal, widespread across taxa, and common in nucleic‑acid–binding and secreted proteins. | 5.47 | 8 |
| #787 | Basic disordered microprotein regions | The feature activates within small proteins, viral accessory/regulatory proteins, and short ORFs, often on disordered or low-complexity segments. Activation is frequently associated with basic-residue–containing patches (Lys/Arg) but is not restricted to a particular position in the protein. | 5.45 | 15 |
| #1579 | N-terminal amphipathic/hydrophobic helices | Generic short N-terminal amphipathic or hydrophobic helices: encompassing secretory signal peptides and signal‑anchor transmembrane helices (Sec/Tat), mitochondrial transit peptides, and similar amphipathic/basic N‑terminal helices found in soluble proteins; the feature can also fire on internal hydrophobic/amphipathic helices. | 5.41 | 2 |
Canonical-only features — 17 total (prevalence ≥ 2)
| Feature | Label | Description | Activation | Prevalence |
|---|---|---|---|---|
| #3476 | Broad-spectrum PPIase detector | Parvulin/PpiC-type peptidyl-prolyl cis-trans isomerase domain detector that recognizes bacterial PrsA-family foldase chaperones and related PPIase domains, including both catalytically active enzymes and inactive PPIase-like homologs (e.g., virulence-associated Mip, lumenal/periplasmic FKBPs) | 4.06 | 2 |
| #16313 | Catalytic-domain N-terminus motif | Short N-terminal segments at the start of the first catalytic domain, often coinciding with the first β-strand and active-site/substrate-binding residues, marking the beginning of the mature/folded core rather than a generic basic patch. | 2.31 | 3 |
| #1949 | PMT FxRFLDxxQY helix | A conserved internal alpha-helical motif in plant/protist phosphoethanolamine N-methyltransferases (PMTs), centered on the sequence "FQRFLDNVQY" (and close variants). | 2.01 | 2 |
| #9952 | Ordered helical assembly scaffold | A broad “ordered helical/assembly scaffold” signature: long, structured segments—often alpha-helical (TM helices, coiled-coils, helical hairpins) and stable cores of ligand/effector-binding domains—that mark architectural elements used for signal transfer and assembly across diverse molecular machines (signaling, secretion, transport, nucleotide/quinone metabolism, and viral replication), rather than specific catalytic micro-motifs. | 1.95 | 2 |
| #4487 | Sparse internal residue peaks | A diffuse, low-specificity feature with sparse residue-level activations distributed across a heterogeneous set of proteins, including bacterial hydrolases (amidases/Ntn-fold), acyl/D-alanyl carrier proteins, Sec translocon γ-subunits (bacterial, archaeal, and eukaryotic SSS1/SecE), and eukaryotic C2H2 zinc-finger transcription factors (ZIC family). | 1.93 | 2 |
| #8749 | RPA/SSB DBD-C zinc finger | Primarily single-stranded nucleic-acid–binding proteins of the RPA/SSB family and related DNA replication/repair auxiliaries, with the strongest residue-level signal concentrated in their C-terminal DNA-binding/zinc-finger module rather than in the N-terminal/central OB barrels; telomere-binding factors (POT1/TPP1/CST), exosome-cap OB subunits, plant organellar SSBs, and other OB-fold ssDNA/RNA-binding proteins also fall within the feature's scope. | 1.93 | 2 |
| #9414 | Peripheral docking interface segments | Peripheral docking/edge segments of cofactor- or intermediate-handling domains in multi-component transfer systems—most often C-terminal extensions or boundary β-strand/helix segments that mediate partner interactions and conformational gating around active sites (including corrinoid/B12, FAD/siroheme, CoA modules, and phosphoryl-transfer PTS domains; prominently observed in ThDP-system 2-oxoacid:ferredoxin oxidoreductase / pyruvate synthase alpha subunits). | 1.84 | 2 |
| #9937 | C-terminal accessory glycan region | C-terminal accessory region downstream of the catalytic core in glycan‑modifying enzymes—most prominently FGly‑dependent sulfatases but also other carbohydrate‑active hydrolases and lumenal/membrane glycan transferases. The feature targets the distal C‑terminal subdomain or tail (often a coil→alpha‑helix segment, or a flexible basic tail in SULFs) that contributes to substrate engagement/localization rather than the catalytic nucleophile/metal sites. | 1.82 | 2 |
| #11298 | Noncatalytic nucleotide-binding helical scaffold | An alpha-helical scaffold signature found inside nucleotide-binding enzyme cores—prominently in kinesin motor domains and in analogous helical subdomains of other NTPases, kinases, and TIR domains; the signal concentrates on non-catalytic helices adjacent to, but not coincident with, ATP/catalytic motifs, consistent with structural elements used for conformational coupling or interface formation. | 1.81 | 2 |
| #8679 | C-terminal domain boundary marker | Residue-level marker of structural boundaries—especially C‑terminal ends of compact domains or segments—often at helix/turn caps or immediately into linkers/disordered tails; occasionally a single interior residue within a long helix | 1.73 | 2 |
| #4397 | Secondary-structure boundary caps | Short 5–12 residue structural motifs at secondary-structure boundaries—often coil→strand or coil→helix transitions—acting as flexible caps or hinges; a general structural (not chemistry-specific) signal that recurs in oxidoreductases and glycosyl hydrolases. | 1.73 | 2 |
| #5824 | Pro/Gly-rich beta-strand caps | Short Pro/Gly-enriched β-strand edge/turn motifs at strand–loop (and loop–helix) junctions—often at the N-terminus of a domain—i.e., β-hairpin turns and strand N-caps used as structural connectors rather than catalytic sites | 1.71 | 2 |
| #5462 | RBP IDRs and coiled-coils | Extended regions outside the canonical RRM RNA-binding core in eukaryotic (often nuclear) RNA/RNP proteins. The feature lights up both long compositionally biased intrinsically disordered segments (Arg/Gly/Ser/Pro–rich low-complexity stretches, RG/RGG boxes, poly-G/P, and mixed basic/acidic runs) and conserved coiled-coil/helical dimerization regions adjacent to the RNA-binding modules (notably the DBHS-family coiled coil). | 1.70 | 2 |
| #7907 | Domain-edge mixed secondary surface segment | A broad surface segment at domain boundaries—often a region just after a signal peptide or N-terminal cap in secreted/periplasmic carbohydrate-active enzymes, and analogous regions in other proteins—that mixes disordered coils with short beta-strands/loops and lies outside catalytic residues but can host substrate-contact or modification sites. | 1.67 | 2 |
| #1914 | DE-rich acidic IDR tails | DE-rich, low-complexity intrinsically disordered acidic tracts—typically terminal tails—characterized by long Asp/Glu‑enriched stretches (often with G/S/P) in proteins from large macromolecular assemblies (e.g., transcription/translation/proteostasis complexes). | 1.67 | 2 |
| #7891 | Short localized secondary-structure segments | Short, contiguous local secondary-structure segments (~15-40 aa) within specific small RNA/protein-interaction domains, most prominently CpcD-like domains (in phycobilisome linker polypeptides and ferredoxin-NADP reductases) and R3H domains (in RNA-binding regulatory proteins). Peaks center on a β-strand/turn or short helix within the domain fold. | 1.65 | 2 |
| #3180 | Noncatalytic C-terminal interaction regions | Non-catalytic C-terminal interaction regions—long, charged/polar terminal domains or tails (often low-complexity or repeat-based) that mediate assembly, trafficking, and RNA/protein binding, rather than enzymatic catalysis; includes structured terminal modules (e.g., NTF2-like) and flexible extensions, with a strong bias toward the final domain/segment of the protein. | 1.48 | 2 |
Shared features by |Δ| activation — 938 total (prevalence ≥ 2)
| Feature | Label | Description | Δ activation | Iso act. | Canon act. | Iso prev. | Canon prev. |
|---|---|---|---|---|---|---|---|
| #220 | Phenylalanine-rich hydrophobic motif detector | Phenylalanine-focused residue identity feature: detects Phe (F) residues, with a preference for F-rich, hydrophobic stretches (e.g., signal peptides and transmembrane helices), but independent of secondary structure or specific function | +7.20 | 15.63 | 8.43 | 7 | 5 |
| #7985 | Signal peptides and disordered tails | Compositionally biased, hydrophobic/aromatic-rich segments—often in low-structure regions including N-terminal pre-sequences, flexible linkers/tails, and short exposed stretches within mature chains. These regions are enriched for Leu/Val/Ile/Phe/Tyr with Pro/Gly and often include low-complexity repeats (e.g., Asn-rich tracts in Dictyostelium). The concept spans signal/transit peptides of secreted/membrane/organellar proteins and intrinsically disordered tails in soluble enzymes and viral accessory proteins. | +6.76 | 10.25 | 3.49 | 13 | 9 |
| #6125 | Disordered PTM/cleavage SLiMs | Short linear motifs in intrinsically disordered/low-complexity regions, often S/T/P/G- and aromatic-rich, that include proteolytic-processing/PTM hotspots and flexible linkers adjacent to transmembrane segments; the feature favors flexible coils or short amphipathic helices and systematically avoids hydrophobic transmembrane cores | +6.17 | 8.54 | 2.37 | 13 | 7 |
| #3503 | Cationic/hydrophobic low-complexity segments | Compositionally biased and low-complexity segments enriched in hydrophobic (L/V/I/A), basic (K/R), and Ser/Thr/Pro residues (with underrepresented Trp). The feature marks both cationic Ser/Thr/Pro-enriched segments and hydrophobic leader-like stretches, capturing N-terminal targeting peptides, micropeptides, and internal polybasic/low-complexity motifs in diverse proteins. | +5.98 | 8.04 | 2.06 | 11 | 2 |
| #678 | Sparse terminal residue peaks | Sparse residue-level activations distributed across short proteins and protein segments, with frequent firing in disordered N-terminal/C-terminal tails as well as in some short transmembrane or compact domains | +5.78 | 8.42 | 2.64 | 20 | 2 |
| #14891 | Disordered termini and cleavage detector | Residue-level detector of intrinsically disordered, flexible termini and proteolytic processing junctions—especially N-termini (including the first residue of the mature chain after propeptide/leader cleavage)—with no strict amino‑acid specificity and a bias toward small/charged residues; common in small, secreted/precursor and viral proteins but also present in disordered regions of diverse proteins. | +5.60 | 8.12 | 2.53 | 17 | 3 |
| #7523 | Diffuse isoleucine composition bias | Weak global preference for isoleucine (and closely related aliphatic hydrophobes), captured primarily as a diffuse composition signal rather than discrete site recognition. | +5.34 | 11.77 | 6.43 | 6 | 4 |
| #13480 | Alanine and small-residue enrichment | A sequence-composition feature that detects small, non-aromatic residues—especially alanine—and their enrichment, lighting up individual Ala residues and stretches enriched in A/G/S/T/P/V (alanine-rich, low-complexity segments) irrespective of protein function or secondary structure. | +5.07 | 9.49 | 4.43 | 23 | 8 |
| #2692 | Compositionally biased disordered LCRs | Feature targets compositionally biased, intrinsically disordered low‑complexity regions with long contiguous runs strongly enriched in small/polar (Gly/Ser/Asn/Thr) or acidic (Asp/Glu) residues; occasional activation on highly basic protamine‑like LCRs. Mere disorder, generic tails, or coiled‑coils are insufficient without such compositional bias. | +4.60 | 7.71 | 3.11 | 13 | 6 |
| #6176 | Sparse activation in small proteins | Sparse, low-amplitude activation distributed across small or short proteins, with no strict residue preference and peaks landing in a variety of structural contexts. | +4.59 | 8.55 | 3.96 | 6 | 5 |
| #9237 | Glycine detector in disordered regions | Residue-identity detector for glycine (G), with a mild bias toward intrinsically disordered/low-complexity segments (often including very N-terminal residues); pan-taxonomic and function-agnostic, but activation is sparse and frequently absent, becoming detectable mainly in glycine-rich low-complexity contexts; occasional weak responses to other small/polar residues. | +4.52 | 10.41 | 5.89 | 15 | 10 |
| #6356 | Ser/Gly/Pro-rich low-complexity IDRs | Often fires on intrinsically disordered, low-complexity regions (IDRs) enriched in Ser/Gly/Pro and basic (Lys/Arg) residues, frequently in secreted/viral proteins, but also activates on short peptides and small folded proteins with no clear sequence motif. Includes Ser-rich PTM-prone segments within unstructured regions. | +4.04 | 7.79 | 3.75 | 14 | 8 |
| #9395 | Short terminal targeting segment | Terminal or short internal targeting/interaction segment: a short, compositionally biased or hydrophobic patch near a protein terminus, or occasionally an internal low‑complexity loop, that functions as a signal peptide or signal‑anchor (Sec/ER/thylakoid/Tat entry) when hydrophobic, or as a basic/Ser/Lys/Arg/Pro‑rich low‑complexity segment that mediates localization or partner binding in soluble proteins; typically the single prominent terminal segment of short secreted, membrane, or regulatory proteins | +3.92 | 6.82 | 2.89 | 10 | 5 |
| #2319 | Threonine residue detector | Residue-identity detector for threonine (Thr): activates on individual Thr residues across diverse proteins, with a mild enrichment in low-complexity/disordered, repeat-rich, and secretory regions; still marks Thr within well-structured domains; occasional weak spillover to serine. | +3.86 | 8.77 | 4.91 | 14 | 5 |
| #6384 | Widespread activation in folded proteins | Proteins showing broad, diffuse activation across large fractions of their sequence, with peaks frequently occurring in folded regions of single-domain or multi-domain enzymes and structural proteins. | +3.78 | 6.35 | 2.57 | 81 | 23 |
| #14000 | IDR and signal peptide detector | Generic detector of low-complexity/intrinsically disordered segments and short hydrophobic N‑terminal stretches (signal peptides/first TM anchors), with preference for S/T/P/G/A/N- and N/Q‑rich tracts; largely avoids well‑ordered helical cores | +3.77 | 6.03 | 2.26 | 10 | 2 |
| #583 | Asp-biased disordered region detector | Detector of Asp residues embedded in low-complexity/disordered segments, with a strong bias for Asp over Glu; fires broadly in polar or acidic-enriched tracts and only weakly on isolated Asp in structured domains. | +3.75 | 9.80 | 6.05 | 22 | 21 |
| #4358 | Low complexity disordered regions | Compositionally biased, low-complexity segments enriched in small/polar residues (serine, glycine, asparagine, proline) and often basic (lysine/arginine) or acidic tracts; the feature preferentially marks IDRs in N/C-terminal tails and repeat-rich regions, though some activations also occur within folded domains. | +3.61 | 6.30 | 2.70 | 14 | 8 |
| #14866 | Intrinsic disorder and low-complexity regions | Intrinsic disorder/low-complexity signal: the feature marks compositionally biased, non-globular regions enriched in polar/charged and small residues (S/T/E/D/R/K/G/P; often A/L), typically flexible N- or C‑terminal tails, linkers, and propeptides that host short linear motifs or processing sites, while avoiding structured domains. | +3.58 | 7.80 | 4.22 | 38 | 35 |
| #1260 | Disordered low-complexity basic segments | Broadly distributed feature with notable but non-exclusive enrichment in compositionally biased, intrinsically disordered low-complexity regions (LCRs), including Gly/Ser/Pro-rich and Arg/Lys-rich tracts or simple repeats. | +3.49 | 6.26 | 2.77 | 14 | 7 |
| #6757 | N-terminal disordered propeptide signature | Generic signature of small proteins, intrinsically disordered and low-complexity segments, and precursor/propeptide regions (S/T/P/G/A-enriched), often at N-termini, rather than structured catalytic or metal-binding motifs | +3.23 | 5.22 | 1.99 | 23 | 2 |
| #11358 | S/G/P/R-rich low-complexity tracts | Compositionally biased, low-complexity tracts enriched in Ser/Gly with frequent Pro/Arg, found in disordered regions of diverse proteins including small secreted peptide precursors, viral accessory proteins, and intracellular disordered segments; the feature marks residues within or just downstream of simple S/G/P/R-rich repeats. | +3.23 | 5.89 | 2.66 | 8 | 2 |
| #12620 | Disordered low-complexity regions | Intrinsically disordered, low‑complexity, compositionally biased regions/tails (IDRs), typically enriched in Ser/Gly/Pro/Ala/Thr and often occurring as acidic (Asp/Glu) or basic (Arg/Lys) tracts or Gln/Asn‑/Gln‑rich repeats; these segments are common in secreted precursors, viral proteins, micropeptides, and testis‑associated proteins, and can also occur as low‑complexity termini or surface loops appended to otherwise folded enzymes; they frequently coincide with low predicted structural confidence. | +3.22 | 6.05 | 2.83 | 17 | 2 |
| #5474 | Disordered Arg–Pro motifs | Detector enriched on basic/polar residues near Proline in flexible or disordered protein segments, with frequent firing on Arg and Pro within Arg-Pro (R-P, RRP) motifs and on residues preceding Pro (e.g., S/T-P) inside low-complexity or disordered stretches. Also fires sporadically in non-disordered regions on similar local sequence contexts. | +3.19 | 5.16 | 1.98 | 9 | 3 |
| #12662 | Sparse short-stretch disordered activations | Sparse, short-stretch activations distributed across diverse protein contexts, with a tendency to fire in non-conserved insert regions, disordered/low-complexity segments, and short N-terminal propeptide regions, but also occurring within structured domains and transmembrane helices. | +3.19 | 5.86 | 2.68 | 8 | 5 |
| #4740 | Disordered loops and linkers | Low-complexity / flexible regions enriched in polar (Ser/Thr/Asn/Gln), basic (Lys/Arg), and Gly/Pro residues, including disordered linkers, propeptides, surface loops, and flexible segments of membrane and globular proteins across taxa (including viral proteins, small secreted peptide precursors, and membrane-protein cytosolic loops) | +3.18 | 6.12 | 2.94 | 13 | 10 |
| #627 | Generic serine detector | Generic serine detector: the feature marks serine residues, with weaker affinity for threonine (and occasionally proline), and is especially prominent in low-complexity S/T-rich stretches; it is context- and structure-agnostic and does not correspond to a specific functional motif. | +3.15 | 8.81 | 5.66 | 17 | 13 |
| #14534 | Unknown generic feature | Unknown generic feature | +12.51 | 27.11 | 14.60 | 236 | 181 |
| #9005 | Unknown generic feature | Unknown generic feature | +12.28 | 22.83 | 10.55 | 217 | 165 |
| #1803 | Unknown generic feature | Unknown generic feature | +7.70 | 25.00 | 17.30 | 237 | 182 |
Part 2 · Differential coordinates
Features firing on just the isoform-unique (extension / alt-frame) residues.
Unique-region features — 322 distinct (prevalence ≥ 2)
| Feature | Label | Description | Activation | Prevalence |
|---|---|---|---|---|
| #220 | Phenylalanine-rich hydrophobic motif detector | Phenylalanine-focused residue identity feature: detects Phe (F) residues, with a preference for F-rich, hydrophobic stretches (e.g., signal peptides and transmembrane helices), but independent of secondary structure or specific function | 15.63 | 2 |
| #13702 | Regulatory low-complexity IDRs | Eukaryotic intrinsically disordered, low‑complexity regulatory segments (terminal tails and inter‑domain linkers) enriched in Pro/Ser/Thr/acidic composition and short linear motifs for modular protein–protein interactions—frequently including WW‑domain–binding PY/PPxY segments and PTM hotspots—with reduced but not absent activation within folded recognition/catalytic domains (WW, PTB/PID, chromo/chromoshadow, SET). | 11.77 | 46 |
| #7523 | Diffuse isoleucine composition bias | Weak global preference for isoleucine (and closely related aliphatic hydrophobes), captured primarily as a diffuse composition signal rather than discrete site recognition. | 11.77 | 2 |
| #8545 | N-terminus accessibility sensor | Detector of accessible peptide chain termini—primarily the extreme N-terminus (initiator methionine and immediate neighbors) in flexible, unstructured tails; position-specific rather than residue-specific—with occasional weak recognition of the C-terminus; common but not universal. | 10.51 | 2 |
| #9852 | N-terminus activation; hydrophobic helices secondary | Extreme N-terminal regions of proteins, starting at the initiator methionine and extending over a short stretch, with occasional secondary activation at internal hydrophobic/amphipathic helices. | 10.44 | 53 |
| #9237 | Glycine detector in disordered regions | Residue-identity detector for glycine (G), with a mild bias toward intrinsically disordered/low-complexity segments (often including very N-terminal residues); pan-taxonomic and function-agnostic, but activation is sparse and frequently absent, becoming detectable mainly in glycine-rich low-complexity contexts; occasional weak responses to other small/polar residues. | 10.41 | 5 |
| #14493 | Short hydrophobic helical runs | Short hydrophobic helical segments rich in I/V/F (sometimes M), recognized in both membrane-targeting/insertion contexts (signal peptides, signal-anchors, transmembrane helices) and in hydrophobic helical patches within otherwise soluble small proteins. The feature flags contiguous aliphatic/hydrophobic runs across diverse proteins, including viral membrane/accessory proteins, secreted peptide/effector precursors, small ORFs/microproteins, and short basic/uncharacterized human proteins. | 10.31 | 4 |
| #7985 | Signal peptides and disordered tails | Compositionally biased, hydrophobic/aromatic-rich segments—often in low-structure regions including N-terminal pre-sequences, flexible linkers/tails, and short exposed stretches within mature chains. These regions are enriched for Leu/Val/Ile/Phe/Tyr with Pro/Gly and often include low-complexity repeats (e.g., Asn-rich tracts in Dictyostelium). The concept spans signal/transit peptides of secreted/membrane/organellar proteins and intrinsically disordered tails in soluble enzymes and viral accessory proteins. | 10.25 | 3 |
| #583 | Asp-biased disordered region detector | Detector of Asp residues embedded in low-complexity/disordered segments, with a strong bias for Asp over Glu; fires broadly in polar or acidic-enriched tracts and only weakly on isolated Asp in structured domains. | 9.80 | 2 |
| #2114 | Chromodomain and integrase-proximal activation | A feature that activates broadly across chromodomain-containing proteins and other chromatin-associated factors, as well as on retrotransposon Gag-Pol polyproteins (in integrase-proximal regions). Activation is distributed widely along the chain with peaks tending to fall within or adjacent to chromodomain folds rather than on catalytic enzyme cores. | 9.73 | 31 |
| #7903 | Arginine-rich polybasic patches | Basic polycationic patches enriched in arginine (often with lysine) within low-complexity/disordered regions, frequently N-terminal. These include classical monopartite NLSs, nucleic acid–binding basic tails, signal peptide n-regions and cytosolic juxtamembrane segments (positive-inside rule), and basic clusters in secreted precursors; typically depleted of acidic residues. | 9.52 | 5 |
| #12385 | Aromatic F/Y cluster motif | Aromatic (phenylalanine/tyrosine) cluster motif: short, hydrophobic low‑complexity segments enriched for F/Y that act as membrane‑interface anchors or aromatic “sticker” patches, most often at N‑termini or adjacent to transmembrane helices; recurrent in small viral/accessory proteins, secreted effectors and toxin/antimicrobial peptides, micropeptides from alt/lncRNA ORFs, and some nucleic‑acid–associated proteins. | 9.51 | 2 |
| #13480 | Alanine and small-residue enrichment | A sequence-composition feature that detects small, non-aromatic residues—especially alanine—and their enrichment, lighting up individual Ala residues and stretches enriched in A/G/S/T/P/V (alanine-rich, low-complexity segments) irrespective of protein function or secondary structure. | 9.49 | 15 |
| #448 | N-terminal leader/targeting segments | N-terminal leader/targeting segments: the feature activates on the extreme N-terminus (first 2–5 residues) of proteins, especially the N-region of signal peptides and organelle transit peptides, and more generally on short, disordered/low-complexity N-terminal tails; strongest at residue ~3 and largely independent of amino-acid identity, occurring across all taxa and functions. | 9.38 | 2 |
| #7384 | Leucine-biased disordered hydrophobic segments | Leucine-biased recognition of intrinsically disordered, low-complexity hydrophobic segments, including disordered insertions within otherwise structured proteins. | 9.24 | 8 |
| #4900 | Arginine-rich segments and residues | Arginine residue identity/basic-tract feature: primarily marks arginine (R) residues—and dense arginine-rich, low-complexity segments—independent of protein family, fold, or function; shows a strong bias for R over other residues (occasional weak lysine signal), appearing both in disordered/repetitive regions and on individual R within structured domains or near transmembrane boundaries. | 8.84 | 8 |
| #627 | Generic serine detector | Generic serine detector: the feature marks serine residues, with weaker affinity for threonine (and occasionally proline), and is especially prominent in low-complexity S/T-rich stretches; it is context- and structure-agnostic and does not correspond to a specific functional motif. | 8.81 | 4 |
| #2319 | Threonine residue detector | Residue-identity detector for threonine (Thr): activates on individual Thr residues across diverse proteins, with a mild enrichment in low-complexity/disordered, repeat-rich, and secretory regions; still marks Thr within well-structured domains; occasional weak spillover to serine. | 8.77 | 9 |
| #6125 | Disordered PTM/cleavage SLiMs | Short linear motifs in intrinsically disordered/low-complexity regions, often S/T/P/G- and aromatic-rich, that include proteolytic-processing/PTM hotspots and flexible linkers adjacent to transmembrane segments; the feature favors flexible coils or short amphipathic helices and systematically avoids hydrophobic transmembrane cores | 8.54 | 3 |
| #11159 | Glycine residue identity detector | Residue-identity detector for glycine (G), largely context-agnostic but with a tendency to fire on glycines in flexible/disordered, often N-terminal or membrane-proximal regions; also active on glycines within transmembrane helices; broadly taxon- and function-independent. | 8.49 | 4 |
| #678 | Sparse terminal residue peaks | Sparse residue-level activations distributed across short proteins and protein segments, with frequent firing in disordered N-terminal/C-terminal tails as well as in some short transmembrane or compact domains | 8.42 | 17 |
| #16243 | Structured-region leucine detector | Generic detector of leucine side chains in structured regions of folded domains, with occasional weaker responses to other hydrophobics; largely indifferent to specific functional motifs or protein class. | 8.35 | 8 |
| #14891 | Disordered termini and cleavage detector | Residue-level detector of intrinsically disordered, flexible termini and proteolytic processing junctions—especially N-termini (including the first residue of the mature chain after propeptide/leader cleavage)—with no strict amino‑acid specificity and a bias toward small/charged residues; common in small, secreted/precursor and viral proteins but also present in disordered regions of diverse proteins. | 8.12 | 14 |
| #5337 | Pro/RGG-rich low-complexity IDRs | Intrinsically disordered, low‑complexity basic segments at termini and long loops, enriched in Pro/Gly and/or Arg/Ser (including RG/RGG or S/T clusters), common across eukaryotic and viral proteins. | 8.06 | 47 |
| #14534 | Unknown generic feature | Unknown generic feature | 27.11 | 54 |
| #1803 | Unknown generic feature | Unknown generic feature | 25.00 | 54 |
| #9005 | Unknown generic feature | Unknown generic feature | 22.83 | 54 |
| #14895 | Unknown generic feature | Unknown generic feature | 19.38 | 54 |
| #9214 | Unknown generic feature | Unknown generic feature | 14.62 | 54 |
| #9194 | Unknown generic feature | Unknown generic feature | 12.60 | 54 |
Top-K sparse-autoencoder activations on the ESM-C residual stream. Activation is a feature's peak value over its region; Prevalence is the number of residues it fires on; Δ activation is isoform − canonical. Only the top 30 features per category are saved; generic / "unknown" features are hidden until you expand a table. Read more about SAE features →
Clinical variants
Differential region — N-terminal extension (isoform-unique)
| Pos (iso) | AA change | Consequence | Source | Clin. sig. | AF (gnomAD) | Impact | AlphaMissense | ESM-C ΔLLR | Link |
|---|---|---|---|---|---|---|---|---|---|
| 1 | R→G | missense_variant | gnomAD | — | 3.60e-06 | — | N/A | -0.21 | chr17-48101389-T-C |
| 1 | R→R | synonymous_variant | gnomAD | — | 1.20e-06 | — | N/A | 0.00 | chr17-48101389-T-G |
| 2 | D→D | synonymous_variant | gnomAD | — | 1.20e-06 | — | N/A | 0.00 | chr17-48101384-G-A |
| 2 | D→V | missense_variant | gnomAD | — | 2.40e-06 | — | N/A | 0.30 | chr17-48101385-T-A |
| 2 | D→G | missense_variant | gnomAD | — | 7.20e-06 | — | N/A | 1.48 | chr17-48101385-T-C |
| 3 | A→V | missense_variant | gnomAD | — | 2.40e-06 | — | N/A | 0.55 | chr17-48101382-G-A |
| 3 | A→S | missense_variant | gnomAD | — | 1.20e-06 | — | N/A | -1.19 | chr17-48101383-C-A |
| 3 | A→T | missense_variant | gnomAD | — | 9.60e-06 | — | N/A | -1.50 | chr17-48101383-C-T |
| 5 | A→A | synonymous_variant | gnomAD | — | 4.80e-06 | — | N/A | 0.00 | chr17-48101375-C-G |
| 5 | A→P | missense_variant | gnomAD | — | 1.20e-06 | — | N/A | -0.52 | chr17-48101377-C-G |
| 5 | A→T | missense_variant | gnomAD | — | 1.20e-06 | — | N/A | -1.33 | chr17-48101377-C-T |
| 6 | — | inframe_insertion | gnomAD | — | 1.32e-05 | — | N/A | — | chr17-48101372-G-GGCC |
| 7 | T→T | synonymous_variant | gnomAD | — | 2.40e-06 | — | N/A | 0.00 | chr17-48101369-G-C |
| 7 | — | frameshift_variant | gnomAD | — | 1.20e-06 | LoF | — | — | chr17-48101371-TG-T |
| 8 | R→R | synonymous_variant | gnomAD | — | 1.20e-06 | — | N/A | 0.00 | chr17-48101366-C-T |
| 10 | A→A | synonymous_variant | gnomAD | — | 4.80e-06 | — | N/A | 0.00 | chr17-48101360-A-T |
| 10 | A→T | missense_variant | gnomAD | — | 3.60e-06 | — | N/A | -1.54 | chr17-48101362-C-T |
| 12 | F→F | synonymous_variant | gnomAD | — | 1.20e-06 | — | N/A | 0.00 | chr17-48101354-G-A |
| 13 | L→L | synonymous_variant | gnomAD | — | 1.20e-05 | — | N/A | 0.00 | chr17-48101351-C-G |
| 14 | G→G | synonymous_variant | gnomAD | — | 2.40e-06 | — | N/A | 0.00 | chr17-48101348-G-A |
| 14 | G→G | synonymous_variant | gnomAD | — | 3.60e-06 | — | N/A | 0.00 | chr17-48101348-G-T |
| 14 | — | frameshift_variant | gnomAD | — | 2.40e-06 | LoF | — | — | chr17-48101348-GC-G |
| 14 | G→R | missense_variant | gnomAD | — | 1.20e-06 | — | N/A | 0.33 | chr17-48101350-C-G |
| 15 | A→D | missense_variant | gnomAD | — | 3.60e-06 | — | N/A | -1.89 | chr17-48101346-G-T |
| 15 | A→S | missense_variant | gnomAD | — | 6.00e-06 | — | N/A | -0.78 | chr17-48101347-C-A |
| 15 | A→T | missense_variant | gnomAD | — | 7.08e-05 | — | N/A | -1.38 | chr17-48101347-C-T |
| 16 | T→N | missense_variant | gnomAD | — | 1.20e-06 | — | N/A | -1.50 | chr17-48101343-G-T |
| 17 | P→P | synonymous_variant | gnomAD | — | 2.40e-06 | — | N/A | 0.00 | chr17-48101339-G-T |
| 17 | P→T | missense_variant | gnomAD | — | 1.20e-06 | — | N/A | -1.14 | chr17-48101341-G-T |
| 18 | P→S | missense_variant | gnomAD | — | 1.20e-06 | — | N/A | -0.40 | chr17-48101338-G-A |
| 19 | G→S | missense_variant | gnomAD | — | 1.20e-06 | — | N/A | -0.02 | chr17-48101335-C-T |
| 21 | P→P | synonymous_variant | gnomAD | — | 2.40e-06 | — | N/A | 0.00 | chr17-48101327-C-G |
| 21 | P→P | synonymous_variant | gnomAD | — | 4.80e-06 | — | N/A | 0.00 | chr17-48101327-C-T |
| 21 | P→Q | missense_variant | gnomAD | — | 2.40e-06 | — | N/A | -1.05 | chr17-48101328-G-T |
| 22 | — | frameshift_variant | gnomAD | — | 3.60e-06 | LoF | — | — | chr17-48101324-C-CGTCGG |
| 22 | T→P | missense_variant | gnomAD | — | 4.80e-06 | — | N/A | 1.09 | chr17-48101326-T-G |
| 23 | R→R | synonymous_variant | gnomAD | — | 1.20e-06 | — | N/A | 0.00 | chr17-48101321-T-A |
| 25 | A→A | synonymous_variant | gnomAD | — | 1.20e-06 | — | N/A | 0.00 | chr17-48101315-G-A |
| 25 | — | frameshift_variant | gnomAD | — | 1.20e-06 | LoF | — | — | chr17-48101316-G-GCGCGT |
| 25 | A→S | missense_variant | gnomAD | — | 7.19e-06 | — | N/A | -0.03 | chr17-48101317-C-A |
| 25 | — | frameshift_variant | gnomAD | — | 4.80e-06 | LoF | — | — | chr17-48101317-C-CA |
| 26 | S→S | synonymous_variant | gnomAD | — | 7.89e-01 | — | N/A | 0.00 | chr17-48101312-G-A |
| 26 | S→R | missense_variant | gnomAD | — | 4.08e-05 | — | N/A | 0.30 | chr17-48101312-G-C |
| 26 | S→T | missense_variant | gnomAD | — | 1.14e-04 | — | N/A | -0.89 | chr17-48101313-C-G |
| 26 | S→N | missense_variant | gnomAD | — | 1.20e-06 | — | N/A | -2.08 | chr17-48101313-C-T |
| 27 | S→S | synonymous_variant | gnomAD | — | 1.20e-06 | — | N/A | 0.00 | chr17-48101309-G-A |
| 27 | S→N | missense_variant | gnomAD | — | 7.19e-06 | — | N/A | -3.05 | chr17-48101310-C-T |
| 27 | S→G | missense_variant | gnomAD | — | 1.20e-06 | — | N/A | -0.14 | chr17-48101311-T-C |
| 28 | A→A | synonymous_variant | gnomAD | — | 1.20e-06 | — | N/A | 0.00 | chr17-48101306-T-C |
| 28 | A→V | missense_variant | gnomAD | — | 8.39e-05 | — | N/A | -1.84 | chr17-48101307-G-A |
| 28 | A→T | missense_variant | gnomAD | — | 2.40e-06 | — | N/A | -1.86 | chr17-48101308-C-T |
| 29 | A→V | missense_variant | gnomAD | — | 1.20e-06 | — | N/A | -1.18 | chr17-48101304-G-A |
| 30 | P→P | synonymous_variant | gnomAD | — | 1.20e-06 | — | N/A | 0.00 | chr17-48101300-C-A |
| 30 | P→L | missense_variant | gnomAD | — | 2.04e-05 | — | N/A | -0.42 | chr17-48101301-G-A |
| 30 | P→R | missense_variant | gnomAD | — | 2.40e-06 | — | N/A | 0.19 | chr17-48101301-G-C |
| 31 | I→V | missense_variant | gnomAD | — | 1.20e-06 | — | N/A | 1.39 | chr17-48101299-T-C |
| 32 | P→L | missense_variant | gnomAD | — | 1.20e-06 | — | N/A | -0.51 | chr17-48101295-G-A |
| 33 | L→L | synonymous_variant | gnomAD | — | 9.59e-06 | — | N/A | 0.00 | chr17-48101291-G-A |
| 33 | L→F | missense_variant | gnomAD | — | 1.20e-06 | — | N/A | -1.58 | chr17-48101293-G-A |
| 34 | G→G | synonymous_variant | gnomAD | — | 2.40e-06 | — | N/A | 0.00 | chr17-48101288-C-A |
| 34 | G→E | missense_variant | gnomAD | — | 1.20e-06 | — | N/A | -1.55 | chr17-48101289-C-T |
| 34 | G→R | missense_variant | gnomAD | — | 1.20e-06 | — | N/A | 0.92 | chr17-48101290-C-G |
| 35 | L→L | synonymous_variant | gnomAD | — | 1.20e-06 | — | N/A | 0.00 | chr17-48101285-G-A |
| 35 | L→L | synonymous_variant | gnomAD | — | 1.20e-06 | — | N/A | 0.00 | chr17-48101285-G-C |
| 35 | L→F | missense_variant | gnomAD | — | 2.40e-06 | — | N/A | -0.91 | chr17-48101287-G-A |
| 36 | L→F | missense_variant | gnomAD | — | 7.19e-06 | — | N/A | -1.34 | chr17-48101282-C-G |
| 37 | G→G | synonymous_variant | gnomAD | — | 1.20e-06 | — | N/A | 0.00 | chr17-48101279-G-A |
| 38 | A→S | missense_variant | gnomAD | — | 1.20e-06 | — | N/A | -0.08 | chr17-48101278-C-A |
| 38 | A→T | missense_variant | gnomAD | — | 1.20e-06 | — | N/A | -1.48 | chr17-48101278-C-T |
| 40 | L→L | synonymous_variant | gnomAD | — | 1.20e-06 | — | N/A | 0.00 | chr17-48101270-C-T |
| 42 | S→S | synonymous_variant | gnomAD | — | 4.57e-04 | — | N/A | 0.00 | chr17-48077038-G-A |
| 42 | S→S | synonymous_variant | COSMIC | — | — | — | N/A | 0.00 | COSV99848140 |
| 43 | V→V | synonymous_variant | gnomAD | — | 6.91e-07 | — | N/A | 0.00 | chr17-48077035-G-C |
| 43 | V→I | missense_variant | gnomAD | — | 1.52e-05 | — | N/A | -3.64 | chr17-48077037-C-T |
| 43 | V→I | missense_variant | COSMIC | — | — | — | N/A | -3.64 | COSV99848059 |
| 44 | T→T | synonymous_variant | gnomAD | — | 2.76e-06 | — | N/A | 0.00 | chr17-48077032-G-A |
| 44 | T→A | missense_variant | gnomAD | — | 6.91e-07 | — | N/A | -3.11 | chr17-48077034-T-C |
| 44 | — | frameshift_variant | gnomAD | — | 6.91e-07 | LoF | — | — | chr17-48077034-T-TCA |
| 45 | L→V | missense_variant | gnomAD | — | 6.91e-07 | — | N/A | -5.76 | chr17-48077031-G-C |
| 46 | — | inframe_insertion | gnomAD | — | 2.06e-06 | — | N/A | — | chr17-48077027-T-TAAA |
| 47 | T→S | missense_variant | gnomAD | — | 6.18e-06 | — | N/A | -1.13 | chr17-48077024-G-C |
| 47 | T→N | missense_variant | gnomAD | — | 6.87e-07 | — | N/A | -4.07 | chr17-48077024-G-T |
| 50 | L→L | synonymous_variant | gnomAD | — | 6.85e-07 | — | N/A | 0.00 | chr17-48077016-G-A |
| 50 | L→V | missense_variant | gnomAD | — | 2.19e-05 | — | N/A | -6.18 | chr17-48077016-G-C |
| 51 | A→A | synonymous_variant | gnomAD | — | 3.70e-05 | — | N/A | 0.00 | chr17-48077011-C-T |
| 51 | A→V | missense_variant | gnomAD | — | 6.85e-07 | — | N/A | -5.70 | chr17-48077012-G-A |
| 51 | A→V | missense_variant | COSMIC | — | — | — | N/A | -5.70 | COSV56682161 |
| 52 | G→G | synonymous_variant | gnomAD | — | 6.85e-07 | — | N/A | 0.00 | chr17-48077008-G-A |
| 52 | G→G | synonymous_variant | COSMIC | — | — | — | N/A | 0.00 | COSV56681939 |
| 53 | T→T | synonymous_variant | gnomAD | — | 6.85e-07 | — | N/A | 0.00 | chr17-48077005-A-G |
90 variants in the differential region.
Shared canonical core — sequence common to canonical and isoform; AlphaMissense applies here
| Pos (iso) | AA change | Consequence | Source | Clin. sig. | AF (gnomAD) | Impact | AlphaMissense | ESM-C ΔLLR | Link |
|---|---|---|---|---|---|---|---|---|---|
| 54 | M→I | missense_variant | gnomAD | — | 1.37e-06 | damaging | — | -10.94 | chr17-48077002-C-G |
| 54 | M→I | missense_variant | gnomAD | — | 6.85e-07 | damaging | — | -10.94 | chr17-48077002-C-T |
| 54 | M→T | missense_variant | gnomAD | — | 6.85e-07 | damaging | — | -11.50 | chr17-48077003-A-G |
| 54 | M→L | missense_variant | gnomAD | — | 6.85e-07 | damaging | — | -11.19 | chr17-48077004-T-G |
| 54 | M→V | missense_variant | COSMIC | — | — | damaging | — | -11.44 | COSV56682591 |
| 55 | G→G | synonymous_variant | gnomAD | — | 6.85e-07 | — | — | 0.00 | chr17-48076999-C-A |
| 55 | G→V | missense_variant | gnomAD | — | 6.85e-07 | damaging | likely_pathogenic (0.88) | -9.25 | chr17-48077000-C-A |
| 55 | G→W | missense_variant | gnomAD | — | 6.85e-07 | damaging | likely_pathogenic (0.97) | -12.31 | chr17-48077001-C-A |
| 56 | — | frameshift_variant | COSMIC | — | — | LoF | — | — | COSV56681712 |
| 56 | — | frameshift_variant | COSMIC | — | — | LoF | — | — | COSV56681499 |
| 57 | K→K | synonymous_variant | gnomAD | — | 2.74e-06 | — | — | 0.00 | chr17-48076993-T-C |
| 58 | Q→Q | synonymous_variant | gnomAD | — | 6.84e-07 | — | — | 0.00 | chr17-48076990-T-C |
| 58 | Q→K | missense_variant | COSMIC | — | — | damaging | ambiguous (0.37) | -10.00 | COSV56682149 |
| 59 | N→K | missense_variant | gnomAD | — | 1.85e-05 | damaging | likely_pathogenic (0.68) | -8.87 | chr17-48076987-G-C |
| 59 | N→K | missense_variant | ClinVar | Uncertain significance | — | damaging | likely_pathogenic (0.68) | -8.87 | ClinVar:3827926 |
| 60 | K→N | missense_variant | gnomAD | — | 6.84e-07 | damaging | likely_pathogenic (0.73) | -9.94 | chr17-48076984-C-G |
| 60 | — | inframe_deletion | gnomAD | — | 6.84e-07 | — | — | — | chr17-48076984-CTTG-C |
| 60 | K→Q | missense_variant | gnomAD | — | 1.37e-06 | damaging | likely_benign (0.24) | -10.25 | chr17-48076986-T-G |
| 61 | K→N | missense_variant | gnomAD | — | 6.84e-07 | damaging | likely_pathogenic (0.84) | -9.87 | chr17-48076981-C-A |
| 61 | K→R | missense_variant | gnomAD | — | 6.84e-07 | damaging | likely_benign (0.12) | -9.06 | chr17-48076982-T-C |
| 62 | K→N | missense_variant | gnomAD | — | 6.84e-07 | damaging | likely_pathogenic (0.58) | -9.00 | chr17-48076978-T-A |
| 62 | — | inframe_deletion | gnomAD | — | 1.37e-06 | — | — | — | chr17-48076978-TTTC-T |
| 62 | K→I | missense_variant | gnomAD | — | 6.84e-07 | damaging | ambiguous (0.56) | -11.69 | chr17-48076979-T-A |
| 62 | K→R | missense_variant | gnomAD | — | 9.58e-06 | damaging | likely_benign (0.12) | -8.94 | chr17-48076979-T-C |
| 62 | K→* | stop_gained | gnomAD | — | 6.84e-07 | LoF | — | — | chr17-48076980-T-A |
| 62 | K→E | missense_variant | gnomAD | — | 6.84e-07 | damaging | ambiguous (0.42) | -9.69 | chr17-48076980-T-C |
| 63 | V→V | synonymous_variant | gnomAD | — | 1.37e-06 | — | — | 0.00 | chr17-48076975-C-T |
| 63 | V→G | missense_variant | gnomAD | — | 6.85e-07 | damaging | likely_benign (0.14) | -8.75 | chr17-48076976-A-C |
| 63 | V→M | missense_variant | gnomAD | — | 7.53e-06 | damaging | likely_benign (0.13) | -8.37 | chr17-48076977-C-T |
| 63 | V→L | missense_variant | COSMIC | — | — | damaging | likely_benign (0.20) | -8.06 | COSV56682280 |
| 66 | V→V | synonymous_variant | gnomAD | — | 1.37e-06 | — | — | 0.00 | chr17-48076966-C-T |
| 66 | V→L | missense_variant | gnomAD | — | 2.05e-06 | damaging | likely_benign (0.30) | -7.78 | chr17-48076968-C-A |
| 66 | V→M | missense_variant | gnomAD | — | 2.74e-06 | damaging | likely_benign (0.24) | -8.68 | chr17-48076968-C-T |
| 67 | L→L | synonymous_variant | gnomAD | — | 2.05e-06 | — | — | 0.00 | chr17-48076963-T-C |
| 68 | E→K | missense_variant | gnomAD | — | 6.84e-07 | damaging | likely_pathogenic (0.83) | -11.12 | chr17-48076962-C-T |
| 69 | E→Q | missense_variant | gnomAD | — | 6.84e-07 | damaging | likely_pathogenic (0.62) | -9.69 | chr17-48076959-C-G |
| 69 | E→K | missense_variant | COSMIC | — | — | damaging | likely_pathogenic (0.83) | -9.56 | COSV56681509 |
| 70 | E→Q | missense_variant | gnomAD | — | 1.37e-06 | damaging | likely_pathogenic (0.61) | -9.81 | chr17-48076956-C-G |
| 71 | E→E | synonymous_variant | gnomAD | — | 1.37e-06 | — | — | 0.00 | chr17-48076951-T-C |
| 71 | E→G | missense_variant | gnomAD | — | 6.84e-07 | damaging | likely_pathogenic (0.77) | -9.56 | chr17-48076952-T-C |
| 72 | E→Q | missense_variant | COSMIC | — | — | damaging | likely_pathogenic (0.95) | -11.25 | COSV99848313 |
| 73 | E→E | synonymous_variant | gnomAD | — | 2.05e-06 | — | — | 0.00 | chr17-48076945-T-C |
| 75 | V→V | synonymous_variant | gnomAD | — | 2.74e-06 | — | — | 0.00 | chr17-48076939-C-T |
| 75 | V→V | synonymous_variant | COSMIC | — | — | — | — | 0.00 | COSV99848141 |
| 76 | V→V | synonymous_variant | gnomAD | — | 6.84e-07 | — | — | 0.00 | chr17-48076936-C-A |
| 76 | V→V | synonymous_variant | gnomAD | — | 6.84e-07 | — | — | 0.00 | chr17-48076936-C-G |
| 77 | E→V | missense_variant | ClinVar | — | — | damaging | likely_pathogenic (1.00) | -11.31 | ClinVar:4423332 |
| 78 | K→N | missense_variant | gnomAD | — | 6.84e-07 | damaging | likely_pathogenic (0.99) | -9.37 | chr17-48076930-T-A |
| 80 | L→L | synonymous_variant | gnomAD | — | 6.84e-06 | — | — | 0.00 | chr17-48076924-G-A |
| 80 | L→I | missense_variant | COSMIC | — | — | damaging | likely_benign (0.28) | -9.56 | COSV56682110 |
| 81 | D→H | missense_variant | gnomAD | — | 1.37e-06 | damaging | likely_pathogenic (0.94) | -7.89 | chr17-48076923-C-G |
| 81 | D→N | missense_variant | gnomAD | — | 1.85e-05 | damaging | likely_pathogenic (0.56) | -4.01 | chr17-48076923-C-T |
| 82 | R→H | missense_variant | gnomAD | — | 2.74e-06 | damaging | likely_pathogenic (0.79) | -7.03 | chr17-48076919-C-T |
| 82 | R→C | missense_variant | gnomAD | — | 1.23e-05 | damaging | likely_pathogenic (0.87) | -7.12 | chr17-48076920-G-A |
| 83 | R→Q | missense_variant | gnomAD | — | 2.05e-06 | damaging | likely_pathogenic (0.99) | -9.25 | chr17-48076916-C-T |
| 83 | R→R | synonymous_variant | gnomAD | — | 6.84e-07 | — | — | 0.00 | chr17-48076917-G-T |
| 83 | R→R | synonymous_variant | COSMIC | — | — | — | — | 0.00 | COSV56682056 |
| 83 | R→* | stop_gained | COSMIC | — | — | LoF | — | — | COSV56681749 |
| 84 | V→V | synonymous_variant | gnomAD | — | 2.05e-06 | — | — | 0.00 | chr17-48076912-C-T |
| 84 | V→A | missense_variant | gnomAD | — | 1.37e-06 | damaging | likely_pathogenic (0.92) | -7.43 | chr17-48076913-A-G |
| 85 | V→V | synonymous_variant | gnomAD | — | 1.37e-06 | — | — | 0.00 | chr17-48076909-T-C |
| 85 | V→L | missense_variant | gnomAD | — | 6.84e-07 | damaging | likely_pathogenic (0.76) | -6.96 | chr17-48076911-C-A |
| 86 | K→K | synonymous_variant | gnomAD | — | 6.84e-07 | — | — | 0.00 | chr17-48076906-C-T |
| 87 | G→G | synonymous_variant | gnomAD | — | 3.42e-06 | — | — | 0.00 | chr17-48076903-G-A |
| 87 | G→S | missense_variant | gnomAD | — | 6.84e-07 | damaging | likely_pathogenic (0.96) | -7.97 | chr17-48076905-C-T |
| 87 | G→S | missense_variant | COSMIC | — | — | damaging | likely_pathogenic (0.96) | -7.97 | COSV99847936 |
| 88 | K→R | missense_variant | gnomAD | — | 2.05e-06 | — | likely_benign (0.09) | -5.18 | chr17-48076901-T-C |
| 90 | E→D | missense_variant | gnomAD | — | 6.85e-07 | damaging | likely_pathogenic (1.00) | -7.72 | chr17-48076894-C-G |
| 90 | E→K | missense_variant | COSMIC | — | — | damaging | likely_pathogenic (1.00) | -10.19 | COSV56682856 |
| 91 | Y→Y | synonymous_variant | gnomAD | — | 1.71e-05 | — | — | 0.00 | chr17-48076891-G-A |
| 92 | L→L | synonymous_variant | gnomAD | — | 3.43e-06 | — | — | 0.00 | chr17-48076888-G-A |
| 92 | L→R | missense_variant | gnomAD | — | 6.85e-07 | damaging | likely_pathogenic (0.84) | -9.56 | chr17-48076889-A-C |
| 92 | L→F | missense_variant | COSMIC | — | — | — | ambiguous (0.35) | -6.53 | COSV56681814 |
| 93 | L→L | synonymous_variant | gnomAD | — | 6.85e-07 | — | — | 0.00 | chr17-48076885-T-G |
| 94 | K→K | synonymous_variant | gnomAD | — | 6.85e-07 | — | — | 0.00 | chr17-48076882-C-T |
| 94 | K→E | missense_variant | gnomAD | — | 6.85e-07 | damaging | likely_pathogenic (1.00) | -11.19 | chr17-48076884-T-C |
| 97 | G→G | synonymous_variant | gnomAD | — | 4.39e-05 | — | — | 0.00 | chr17-48076873-T-C |
| 98 | F→F | synonymous_variant | gnomAD | — | 6.86e-07 | — | — | 0.00 | chr17-48076870-G-A |
| 98 | F→L | missense_variant | COSMIC | — | — | damaging | likely_pathogenic (1.00) | -10.37 | COSV56682209 |
| 101 | E→K | missense_variant | gnomAD | — | 7.10e-07 | damaging | likely_pathogenic (0.84) | -9.50 | chr17-48076177-C-T |
| 101 | E→Q | missense_variant | ClinVar | — | — | damaging | likely_pathogenic (0.60) | -11.12 | ClinVar:4423331 |
| 101 | E→Q | missense_variant | COSMIC | — | — | damaging | likely_pathogenic (0.60) | -11.12 | COSV56681530 |
| 102 | D→H | missense_variant | gnomAD | — | 7.06e-07 | damaging | likely_pathogenic (0.97) | -12.12 | chr17-48076174-C-G |
| 103 | N→D | missense_variant | gnomAD | — | 1.41e-06 | damaging | likely_pathogenic (0.97) | -10.37 | chr17-48076171-T-C |
| 103 | N→S | missense_variant | COSMIC | — | — | damaging | likely_pathogenic (0.59) | -9.00 | COSV56681968 |
| 106 | E→D | missense_variant | COSMIC | — | — | damaging | likely_pathogenic (1.00) | -10.56 | COSV99848274 |
| 108 | E→D | missense_variant | COSMIC | — | — | damaging | likely_pathogenic (0.96) | -8.69 | COSV99848151 |
| 110 | N→N | synonymous_variant | gnomAD | — | 1.38e-06 | — | — | 0.00 | chr17-48076148-G-A |
| 111 | L→L | synonymous_variant | gnomAD | — | 1.38e-06 | — | — | 0.00 | chr17-48076147-G-A |
| 114 | P→P | synonymous_variant | gnomAD | — | 1.51e-05 | — | — | 0.00 | chr17-48076136-G-A |
| 114 | P→P | synonymous_variant | gnomAD | — | 6.88e-07 | — | — | 0.00 | chr17-48076136-G-C |
| 114 | P→P | synonymous_variant | COSMIC | — | — | — | — | 0.00 | COSV99847948 |
| 115 | D→N | missense_variant | COSMIC | — | — | damaging | likely_pathogenic (0.96) | -8.31 | COSV99847954 |
| 116 | L→L | synonymous_variant | gnomAD | — | 1.23e-05 | — | — | 0.00 | chr17-48076130-G-A |
| 118 | A→A | synonymous_variant | gnomAD | — | 5.48e-06 | — | — | 0.00 | chr17-48076124-A-G |
| 118 | A→A | synonymous_variant | gnomAD | — | 6.85e-07 | — | — | 0.00 | chr17-48076124-A-T |
| 118 | A→G | missense_variant | gnomAD | — | 1.37e-06 | damaging | ambiguous (0.38) | -9.94 | chr17-48076125-G-C |
| 121 | L→L | synonymous_variant | gnomAD | — | 1.37e-06 | — | — | 0.00 | chr17-48076115-C-T |
| 121 | L→L | synonymous_variant | gnomAD | — | 6.85e-07 | — | — | 0.00 | chr17-48076117-G-A |
| 121 | L→L | synonymous_variant | COSMIC | — | — | — | — | 0.00 | COSV105045650 |
| 122 | Q→* | stop_gained | COSMIC | — | — | LoF | — | — | COSV99848158 |
| 123 | S→L | missense_variant | gnomAD | — | 2.74e-06 | damaging | likely_benign (0.27) | -9.37 | chr17-48076110-G-A |
| 123 | S→L | missense_variant | ClinVar | Uncertain significance | — | damaging | likely_benign (0.27) | -9.37 | ClinVar:4648788 |
| 124 | Q→R | missense_variant | gnomAD | — | 6.85e-07 | — | likely_benign (0.33) | -7.47 | chr17-48076107-T-C |
| 124 | Q→* | stop_gained | COSMIC | — | — | LoF | — | — | COSV99848270 |
| 125 | K→R | missense_variant | gnomAD | — | 6.85e-07 | — | likely_benign (0.12) | -6.44 | chr17-48076104-T-C |
| 125 | K→Q | missense_variant | ClinVar | Uncertain significance | — | damaging | likely_benign (0.33) | -8.75 | ClinVar:4531663 |
| 126 | T→A | missense_variant | gnomAD | — | 6.85e-07 | — | likely_benign (0.06) | -6.06 | chr17-48076102-T-C |
| 128 | H→P | missense_variant | gnomAD | — | 4.79e-06 | — | likely_benign (0.11) | -7.21 | chr17-48076095-T-G |
| 130 | T→R | missense_variant | gnomAD | — | 2.74e-06 | damaging | likely_benign (0.23) | -7.90 | chr17-48076089-G-C |
| 131 | D→G | missense_variant | gnomAD | — | 6.84e-07 | — | likely_benign (0.10) | -6.27 | chr17-48076086-T-C |
| 131 | D→N | missense_variant | gnomAD | — | 6.84e-07 | — | likely_benign (0.10) | -6.15 | chr17-48076087-C-T |
| 131 | D→G | missense_variant | ClinVar | Uncertain significance | — | — | likely_benign (0.10) | -6.27 | ClinVar:4220041 |
| 131 | D→N | missense_variant | ClinVar | Uncertain significance | — | — | likely_benign (0.10) | -6.15 | ClinVar:4220044 |
| 132 | K→K | synonymous_variant | gnomAD | — | 4.04e-05 | — | — | 0.00 | chr17-48076082-T-C |
| 132 | K→I | missense_variant | gnomAD | — | 6.84e-07 | damaging | ambiguous (0.43) | -10.75 | chr17-48076083-T-A |
| 132 | K→I | missense_variant | ClinVar | Uncertain significance | — | damaging | ambiguous (0.43) | -10.75 | ClinVar:4220042 |
| 133 | S→S | synonymous_variant | gnomAD | — | 4.79e-06 | — | — | 0.00 | chr17-48076079-T-G |
| 134 | E→K | missense_variant | COSMIC | — | — | damaging | likely_benign (0.26) | -8.68 | COSV99848247 |
| 135 | G→R | missense_variant | gnomAD | — | 6.84e-07 | damaging | likely_pathogenic (0.58) | -8.76 | chr17-48076075-C-G |
| 135 | G→R | missense_variant | COSMIC | — | — | damaging | likely_pathogenic (0.58) | -8.76 | COSV56682528 |
| 138 | R→H | missense_variant | gnomAD | — | 7.53e-06 | damaging | likely_pathogenic (0.89) | -9.31 | chr17-48076065-C-T |
| 138 | R→C | missense_variant | gnomAD | — | 6.16e-06 | damaging | likely_pathogenic (0.96) | -9.31 | chr17-48076066-G-A |
| 138 | R→H | missense_variant | ClinVar | Uncertain significance | — | damaging | likely_pathogenic (0.89) | -9.31 | ClinVar:2290145 |
| 138 | R→H | missense_variant | COSMIC | — | — | damaging | likely_pathogenic (0.89) | -9.31 | COSV56682901 |
| 138 | R→C | missense_variant | COSMIC | — | — | damaging | likely_pathogenic (0.96) | -9.31 | COSV56682920 |
| 139 | K→R | missense_variant | gnomAD | — | 8.90e-06 | damaging | likely_benign (0.11) | -7.87 | chr17-48076062-T-C |
| 140 | A→T | missense_variant | gnomAD | — | 1.37e-06 | — | likely_benign (0.08) | -6.34 | chr17-48076060-C-T |
| 140 | A→T | missense_variant | ClinVar | Uncertain significance | — | — | likely_benign (0.08) | -6.34 | ClinVar:4220043 |
| 141 | D→G | missense_variant | gnomAD | — | 6.85e-07 | damaging | likely_benign (0.15) | -8.36 | chr17-48076056-T-C |
| 142 | S→A | missense_variant | gnomAD | — | 6.85e-07 | damaging | likely_benign (0.06) | -9.87 | chr17-48076054-A-C |
| 142 | S→T | missense_variant | gnomAD | — | 1.37e-06 | — | likely_benign (0.05) | -7.31 | chr17-48076054-A-T |
| 146 | D→V | missense_variant | gnomAD | — | 1.37e-06 | damaging | likely_benign (0.13) | -9.56 | chr17-48076041-T-A |
| 146 | D→V | missense_variant | ClinVar | Uncertain significance | — | damaging | likely_benign (0.13) | -9.56 | ClinVar:2307098 |
| 146 | D→N | missense_variant | COSMIC | — | — | damaging | likely_benign (0.11) | -7.93 | COSV56682563 |
| 147 | K→K | synonymous_variant | gnomAD | — | 1.37e-06 | — | — | 0.00 | chr17-48076037-C-T |
| 147 | K→R | missense_variant | gnomAD | — | 2.05e-06 | — | likely_benign (0.09) | -4.90 | chr17-48076038-T-C |
| 147 | K→E | missense_variant | gnomAD | — | 6.85e-07 | damaging | likely_benign (0.18) | -11.05 | chr17-48076039-T-C |
| 147 | K→K | synonymous_variant | COSMIC | — | — | — | — | 0.00 | COSV56682520 |
| 147 | K→N | missense_variant | COSMIC | — | — | damaging | likely_benign (0.29) | -9.30 | COSV99848170 |
| 148 | G→G | synonymous_variant | gnomAD | — | 2.06e-06 | — | — | 0.00 | chr17-48076034-T-C |
| 148 | — | frameshift_variant | gnomAD | — | 1.37e-06 | LoF | — | — | chr17-48076035-CCCTT-C |
| 148 | G→R | missense_variant | gnomAD | — | 2.74e-06 | damaging | ambiguous (0.54) | -7.55 | chr17-48076036-C-T |
| 149 | E→E | synonymous_variant | gnomAD | — | 6.86e-07 | — | — | 0.00 | chr17-48076031-C-T |
| 149 | E→Q | missense_variant | COSMIC | — | — | damaging | likely_benign (0.23) | -8.62 | COSV99848203 |
| 150 | E→D | missense_variant | gnomAD | — | 6.86e-07 | — | likely_benign (0.05) | -5.84 | chr17-48076028-C-G |
| 151 | S→S | synonymous_variant | gnomAD | — | 4.81e-06 | — | — | 0.00 | chr17-48076025-G-A |
| 151 | S→G | missense_variant | gnomAD | — | 6.86e-07 | — | likely_benign (0.06) | -5.94 | chr17-48076027-T-C |
| 151 | S→S | synonymous_variant | COSMIC | — | — | — | — | 0.00 | COSV99848275 |
| 153 | P→P | synonymous_variant | gnomAD | — | 2.07e-06 | — | — | 0.00 | chr17-48076019-T-C |
| 153 | P→L | missense_variant | gnomAD | — | 1.38e-06 | — | likely_benign (0.09) | -6.30 | chr17-48076020-G-A |
| 155 | K→N | missense_variant | gnomAD | — | 6.90e-07 | damaging | likely_pathogenic (0.93) | -10.37 | chr17-48076013-C-A |
| 155 | K→K | synonymous_variant | gnomAD | — | 1.38e-06 | — | — | 0.00 | chr17-48076013-C-T |
| 156 | K→N | missense_variant | gnomAD | — | 6.91e-07 | damaging | likely_pathogenic (0.91) | -10.25 | chr17-48076010-C-G |
| 156 | K→R | missense_variant | gnomAD | — | 1.38e-06 | — | likely_benign (0.09) | -5.37 | chr17-48076011-T-C |
| 156 | K→N | missense_variant | COSMIC | — | — | damaging | likely_pathogenic (0.91) | -10.25 | COSV56681336 |
| 156 | K→N | missense_variant | COSMIC | — | — | damaging | likely_pathogenic (0.91) | -10.25 | COSV99848297 |
| 157 | — | inframe_deletion | gnomAD | — | 2.77e-06 | — | — | — | chr17-48076007-TTTC-T |
| 157 | K→R | missense_variant | gnomAD | — | 6.92e-07 | — | likely_benign (0.09) | -6.75 | chr17-48076008-T-C |
| 158 | E→G | missense_variant | gnomAD | — | 6.92e-07 | damaging | likely_benign (0.33) | -8.81 | chr17-48076005-T-C |
| 159 | E→V | missense_variant | gnomAD | — | 1.39e-06 | damaging | likely_benign (0.30) | -8.24 | chr17-48076002-T-A |
| 159 | E→Q | missense_variant | gnomAD | — | 6.94e-07 | damaging | ambiguous (0.43) | -8.99 | chr17-48076003-C-G |
| 160 | S→S | synonymous_variant | gnomAD | — | 6.85e-07 | — | — | 0.00 | chr17-48075098-T-A |
| 160 | S→L | missense_variant | gnomAD | — | 6.85e-07 | — | likely_benign (0.08) | -6.09 | chr17-48075099-G-A |
| 161 | E→A | missense_variant | gnomAD | — | 6.85e-07 | damaging | ambiguous (0.50) | -9.43 | chr17-48075096-T-G |
| 163 | P→S | missense_variant | gnomAD | — | 6.84e-07 | — | ambiguous (0.53) | -5.76 | chr17-48075091-G-A |
| 163 | P→Q | missense_variant | COSMIC | — | — | damaging | likely_pathogenic (0.61) | -7.20 | COSV106082017 |
| 163 | P→S | missense_variant | COSMIC | — | — | — | ambiguous (0.53) | -5.76 | COSV56682997 |
| 164 | R→Q | missense_variant | gnomAD | — | 6.84e-07 | damaging | likely_pathogenic (0.92) | -8.69 | chr17-48075087-C-T |
| 164 | R→* | stop_gained | gnomAD | — | 2.05e-06 | LoF | — | — | chr17-48075088-G-A |
| 164 | R→* | stop_gained | COSMIC | — | — | LoF | — | — | COSV56682269 |
| 166 | F→S | missense_variant | gnomAD | — | 6.84e-07 | damaging | likely_pathogenic (1.00) | -10.44 | chr17-48075081-A-G |
| 167 | A→A | synonymous_variant | gnomAD | — | 1.37e-06 | — | — | 0.00 | chr17-48075077-A-C |
| 167 | A→A | synonymous_variant | gnomAD | — | 6.84e-07 | — | — | 0.00 | chr17-48075077-A-G |
| 168 | R→Q | missense_variant | gnomAD | — | 6.84e-07 | damaging | likely_pathogenic (0.99) | -8.12 | chr17-48075075-C-T |
| 168 | R→* | stop_gained | gnomAD | — | 6.84e-07 | LoF | — | — | chr17-48075076-G-A |
| 168 | R→R | synonymous_variant | COSMIC | — | — | — | — | 0.00 | COSV107309107 |
| 168 | R→L | missense_variant | COSMIC | — | — | damaging | likely_pathogenic (1.00) | -10.37 | COSV99847942 |
| 168 | R→Q | missense_variant | COSMIC | — | — | damaging | likely_pathogenic (0.99) | -8.12 | COSV56683114 |
| 168 | R→* | stop_gained | COSMIC | — | — | LoF | — | — | COSV104387334 |
| 169 | G→D | missense_variant | COSMIC | — | — | damaging | likely_pathogenic (0.89) | -9.62 | COSV105045685 |
| 171 | E→K | missense_variant | gnomAD | — | 6.84e-07 | damaging | likely_pathogenic (0.87) | -7.54 | chr17-48075067-C-T |
| 171 | E→E | synonymous_variant | COSMIC | — | — | — | — | 0.00 | COSV56681442 |
| 172 | P→P | synonymous_variant | gnomAD | — | 1.97e-04 | — | — | 0.00 | chr17-48075062-C-T |
| 172 | P→L | missense_variant | gnomAD | — | 7.52e-06 | damaging | likely_pathogenic (1.00) | -10.81 | chr17-48075063-G-A |
| 172 | P→P | synonymous_variant | COSMIC | — | — | — | — | 0.00 | COSV99848187 |
| 172 | P→Q | missense_variant | COSMIC | — | — | damaging | likely_pathogenic (1.00) | -12.12 | COSV56682499 |
| 172 | P→S | missense_variant | COSMIC | — | — | damaging | likely_pathogenic (1.00) | -9.69 | COSV99848079 |
| 173 | E→D | missense_variant | gnomAD | — | 1.37e-06 | damaging | likely_pathogenic (0.92) | -9.06 | chr17-48075059-C-A |
| 173 | E→E | synonymous_variant | gnomAD | — | 1.37e-06 | — | — | 0.00 | chr17-48075059-C-T |
| 173 | E→D | missense_variant | ClinVar | Uncertain significance | — | damaging | likely_pathogenic (0.92) | -9.06 | ClinVar:4648787 |
| 173 | E→K | missense_variant | COSMIC | — | — | damaging | likely_pathogenic (1.00) | -12.12 | COSV56681862 |
| 174 | R→R | synonymous_variant | gnomAD | — | 6.84e-07 | — | — | 0.00 | chr17-48075056-C-T |
| 174 | R→Q | missense_variant | COSMIC | — | — | damaging | likely_pathogenic (0.98) | -7.87 | COSV56682101 |
| 174 | R→L | missense_variant | COSMIC | — | — | damaging | likely_pathogenic (0.99) | -10.12 | COSV99848174 |
| 175 | I→V | missense_variant | gnomAD | — | 6.84e-07 | damaging | likely_pathogenic (0.96) | -7.69 | chr17-48075055-T-C |
| 177 | G→A | missense_variant | COSMIC | — | — | damaging | likely_pathogenic (1.00) | -12.44 | COSV56682437 |
| 179 | T→T | synonymous_variant | gnomAD | — | 2.74e-06 | — | — | 0.00 | chr17-48075041-T-C |
| 181 | S→S | synonymous_variant | gnomAD | — | 6.84e-07 | — | — | 0.00 | chr17-48075035-G-A |
| 182 | S→S | synonymous_variant | COSMIC | — | — | — | — | 0.00 | COSV56681832 |
| 184 | E→Q | missense_variant | gnomAD | — | 6.84e-07 | damaging | likely_pathogenic (0.92) | -10.00 | chr17-48075028-C-G |
| 185 | L→L | synonymous_variant | gnomAD | — | 9.59e-06 | — | — | 0.00 | chr17-48075023-G-A |
| 186 | M→L | missense_variant | gnomAD | — | 6.85e-07 | damaging | likely_pathogenic (0.91) | -10.25 | chr17-48075022-T-A |
| 186 | M→V | missense_variant | gnomAD | — | 6.85e-07 | damaging | likely_pathogenic (0.99) | -11.50 | chr17-48075022-T-C |
| 186 | — | inframe_deletion | COSMIC | — | — | — | — | — | COSV56681204 |
| 187 | F→F | synonymous_variant | gnomAD | — | 6.85e-07 | — | — | 0.00 | chr17-48075017-G-A |
| 188 | L→P | missense_variant | gnomAD | — | 6.85e-07 | damaging | likely_pathogenic (1.00) | -13.06 | chr17-48075015-A-G |
| 188 | L→L | synonymous_variant | COSMIC | — | — | — | — | 0.00 | COSV99848148 |
| 188 | L→L | synonymous_variant | COSMIC | — | — | — | — | 0.00 | COSV56681918 |
| 189 | M→I | missense_variant | gnomAD | — | 6.86e-07 | damaging | likely_pathogenic (0.97) | -8.12 | chr17-48075011-C-A |
| 189 | M→T | missense_variant | gnomAD | — | 6.86e-07 | damaging | likely_pathogenic (1.00) | -11.81 | chr17-48075012-A-G |
| 190 | K→K | synonymous_variant | gnomAD | — | 6.86e-07 | — | — | 0.00 | chr17-48075008-T-C |
| 190 | K→T | missense_variant | gnomAD | — | 6.86e-07 | damaging | likely_pathogenic (1.00) | -11.87 | chr17-48075009-T-G |
| 192 | K→R | missense_variant | gnomAD | — | 8.28e-06 | — | likely_benign (0.14) | -6.94 | chr17-48071577-T-C |
| 193 | N→K | missense_variant | COSMIC | — | — | damaging | likely_pathogenic (0.89) | -8.94 | COSV99848229 |
| 194 | S→F | missense_variant | COSMIC | — | — | damaging | likely_pathogenic (0.99) | -10.37 | COSV56681364 |
| 197 | A→A | synonymous_variant | gnomAD | — | 6.87e-07 | — | — | 0.00 | chr17-48071561-A-G |
| 198 | D→D | synonymous_variant | gnomAD | — | 6.87e-07 | — | — | 0.00 | chr17-48071558-G-A |
| 200 | V→L | missense_variant | gnomAD | — | 6.86e-07 | damaging | likely_pathogenic (1.00) | -11.25 | chr17-48071554-C-G |
| 202 | A→A | synonymous_variant | gnomAD | — | 6.85e-07 | — | — | 0.00 | chr17-48071546-G-A |
| 202 | A→D | missense_variant | gnomAD | — | 6.85e-07 | damaging | likely_pathogenic (1.00) | -12.00 | chr17-48071547-G-T |
| 202 | A→V | missense_variant | COSMIC | — | — | damaging | likely_pathogenic (1.00) | -10.62 | COSV99848117 |
| 204 | E→* | stop_gained | gnomAD | — | 6.85e-07 | LoF | — | — | chr17-48071542-C-A |
| 204 | E→K | missense_variant | COSMIC | — | — | damaging | likely_pathogenic (0.99) | -10.94 | COSV56681808 |
| 206 | N→S | missense_variant | gnomAD | — | 6.84e-07 | damaging | likely_pathogenic (0.91) | -12.44 | chr17-48071535-T-C |
| 207 | V→F | missense_variant | gnomAD | — | 2.74e-06 | damaging | likely_pathogenic (0.57) | -8.73 | chr17-48071533-C-A |
| 207 | V→V | synonymous_variant | COSMIC | — | — | — | — | 0.00 | COSV56681933 |
| 208 | K→K | synonymous_variant | COSMIC | — | — | — | — | 0.00 | COSV56682664 |
| 209 | C→C | synonymous_variant | gnomAD | — | 1.37e-06 | — | — | 0.00 | chr17-48071525-G-A |
| 209 | C→S | missense_variant | gnomAD | — | 6.84e-07 | damaging | likely_pathogenic (1.00) | -5.97 | chr17-48071526-C-G |
| 210 | P→L | missense_variant | gnomAD | — | 6.84e-07 | damaging | likely_pathogenic (1.00) | -10.62 | chr17-48071523-G-A |
| 211 | Q→Q | synonymous_variant | gnomAD | — | 6.84e-07 | — | — | 0.00 | chr17-48071519-C-T |
| 211 | Q→* | stop_gained | gnomAD | — | 6.84e-07 | LoF | — | — | chr17-48071521-G-A |
| 212 | V→F | missense_variant | gnomAD | — | 2.05e-06 | damaging | likely_pathogenic (0.99) | -12.05 | chr17-48071518-C-A |
| 214 | I→I | synonymous_variant | COSMIC | — | — | — | — | 0.00 | COSV56681722 |
| 214 | I→T | missense_variant | COSMIC | — | — | damaging | likely_pathogenic (1.00) | -11.06 | COSV56682067 |
| 215 | S→S | synonymous_variant | gnomAD | — | 2.05e-06 | — | — | 0.00 | chr17-48071507-G-T |
| 216 | F→F | synonymous_variant | gnomAD | — | 6.84e-07 | — | — | 0.00 | chr17-48071504-G-A |
| 217 | Y→Y | synonymous_variant | gnomAD | — | 1.37e-06 | — | — | 0.00 | chr17-48071501-A-G |
| 217 | Y→C | missense_variant | gnomAD | — | 6.84e-07 | damaging | likely_pathogenic (1.00) | -11.94 | chr17-48071502-T-C |
| 220 | R→K | missense_variant | COSMIC | — | — | damaging | likely_pathogenic (0.99) | -11.06 | COSV56682552 |
| 222 | T→T | synonymous_variant | gnomAD | — | 4.11e-06 | — | — | 0.00 | chr17-48071486-C-T |
| 222 | T→K | missense_variant | COSMIC | — | — | damaging | likely_pathogenic (0.99) | -13.94 | COSV56681836 |
| 224 | H→H | synonymous_variant | gnomAD | — | 6.84e-07 | — | — | 0.00 | chr17-48071480-A-G |
| 225 | S→S | synonymous_variant | gnomAD | — | 6.84e-07 | — | — | 0.00 | chr17-48071477-G-A |
| 225 | S→F | missense_variant | gnomAD | — | 6.84e-07 | damaging | likely_pathogenic (0.97) | -11.12 | chr17-48071478-G-A |
| 225 | S→A | missense_variant | gnomAD | — | 6.84e-07 | damaging | likely_benign (0.27) | -9.06 | chr17-48071479-A-C |
| 225 | S→P | missense_variant | gnomAD | — | 6.84e-07 | damaging | likely_pathogenic (0.98) | -11.81 | chr17-48071479-A-G |
| 225 | S→T | missense_variant | gnomAD | — | 6.84e-07 | damaging | ambiguous (0.44) | -10.19 | chr17-48071479-A-T |
| 226 | Y→Y | synonymous_variant | gnomAD | — | 2.94e-05 | — | — | 0.00 | chr17-48071474-G-A |
| 226 | Y→* | stop_gained | gnomAD | — | 6.85e-07 | LoF | — | — | chr17-48071474-G-T |
| 227 | P→S | missense_variant | gnomAD | — | 6.85e-07 | damaging | likely_pathogenic (0.58) | -8.75 | chr17-48071473-G-A |
| 227 | P→S | missense_variant | COSMIC | — | — | damaging | likely_pathogenic (0.58) | -8.75 | COSV108798400 |
| 228 | S→S | synonymous_variant | gnomAD | — | 4.11e-06 | — | — | 0.00 | chr17-48071468-C-T |
| 228 | S→L | missense_variant | gnomAD | — | 2.05e-06 | damaging | likely_benign (0.21) | -8.42 | chr17-48071469-G-A |
| 228 | S→L | missense_variant | ClinVar | Uncertain significance | — | damaging | likely_benign (0.21) | -8.42 | ClinVar:3138046 |
| 228 | S→S | synonymous_variant | COSMIC | — | — | — | — | 0.00 | COSV56681272 |
| 228 | S→* | stop_gained | COSMIC | — | — | LoF | — | — | COSV99848014 |
| 229 | E→K | missense_variant | COSMIC | — | — | damaging | likely_pathogenic (0.75) | -12.00 | COSV104555665 |
| 230 | D→E | missense_variant | gnomAD | — | 4.80e-06 | — | likely_benign (0.09) | -5.24 | chr17-48071462-A-C |
| 230 | D→D | synonymous_variant | gnomAD | — | 6.85e-07 | — | — | 0.00 | chr17-48071462-A-G |
| 230 | D→Y | missense_variant | gnomAD | — | 6.85e-07 | damaging | likely_pathogenic (0.71) | -11.56 | chr17-48071464-C-A |
| 230 | D→N | missense_variant | gnomAD | — | 1.37e-06 | damaging | likely_benign (0.31) | -10.18 | chr17-48071464-C-T |
| 230 | D→E | missense_variant | ClinVar | Uncertain significance | — | — | likely_benign (0.09) | -5.24 | ClinVar:2517719 |
| 232 | — | inframe_deletion | gnomAD | — | 2.06e-06 | — | — | — | chr17-48071456-GTCA-G |
| 232 | D→N | missense_variant | COSMIC | — | — | damaging | likely_benign (0.19) | -8.62 | COSV99848218 |
| 233 | K→N | missense_variant | COSMIC | — | — | damaging | likely_pathogenic (0.72) | -9.62 | COSV99848111 |
| 233 | — | frameshift_variant | COSMIC | — | — | LoF | — | — | COSV56682928 |
| 234 | K→* | stop_gained | gnomAD | — | 6.85e-07 | LoF | — | — | chr17-48071452-T-A |
| 235 | D→V | missense_variant | gnomAD | — | 6.86e-07 | damaging | ambiguous (0.45) | -9.47 | chr17-48071448-T-A |
| 235 | — | frameshift_variant | gnomAD | — | 6.86e-07 | LoF | — | — | chr17-48071449-C-CT |
| 235 | — | frameshift_variant | gnomAD | — | 6.86e-07 | LoF | — | — | chr17-48071449-CTTTT-C |
| 235 | D→H | missense_variant | COSMIC | — | — | damaging | likely_pathogenic (0.73) | -10.72 | COSV56682228 |
| 235 | D→Y | missense_variant | COSMIC | — | — | damaging | likely_pathogenic (0.65) | -9.97 | COSV56681987 |
| 236 | — | frameshift_variant | gnomAD | — | 6.86e-07 | LoF | — | — | chr17-48071445-TC-T |
| 236 | D→Y | missense_variant | gnomAD | — | 1.37e-06 | damaging | likely_pathogenic (0.68) | -10.43 | chr17-48071446-C-A |
| 236 | D→N | missense_variant | gnomAD | — | 6.86e-07 | — | likely_benign (0.27) | -7.43 | chr17-48071446-C-T |
| 236 | — | frameshift_variant | gnomAD | — | 6.86e-07 | LoF | — | — | chr17-48071446-CA-C |
| 238 | N→K | missense_variant | gnomAD | — | 6.87e-07 | damaging | likely_pathogenic (0.68) | -7.21 | chr17-48071438-G-C |
| 238 | — | frameshift_variant | gnomAD | — | 1.37e-06 | LoF | — | — | chr17-48071439-TTCTTG-T |
| 238 | N→K | missense_variant | COSMIC | — | — | damaging | likely_pathogenic (0.68) | -7.21 | COSV56682095 |
281 variants in the shared canonical core.
LoF = frameshift / stop-gain / splice-disrupting — inherently loss-of-function, flagged by consequence (AlphaMissense and ESM-C score only missense/substitutions, so they are blank here by design, not by absence of impact). AlphaMissense is computed in the canonical reading frame, so it scores the shared core but reads N/A across an isoform-unique extension — use ESM-C ΔLLR there.