CBX1
EXTENDED 226 aa (canonical 185 aa) · UniProt P83916 · CDLMPS
chr17:48101353:-:CTG:ENST00000225603.9
AI summary A confidently-folded but structurally unmoored 42-aa N-terminal extension is added ahead of the HP1-beta chromoshadow/chromo core, with no domain or localization change.
The extension folds into a discrete strand element (pLDDT ~0.80) yet forms almost no contacts with the canonical CBX1 body and sits at very high inter-region PAE, meaning it is a locally structured appendage rather than an integrated addition to HP1-beta's chromodomain/chromoshadow architecture. The large shared-region RMSD is not trustworthy as a core refold given the low global pTM, so it should not be read as CBX1's H3K9me-reading fold being reorganized. Separately, a modest whole-protein biophysical shift (charge/hydropathy) and an SAE feature-magnitude shift indicate the extension has a detectable but non-domain-forming compositional signature.
CBX1/HP1-beta's core function depends on its chromodomain reading H3K9me3 and its chromoshadow domain mediating dimerization/partner recruitment (SUV39H1, PurB/Sp3, KAP-1); since S1 finds no real InterPro domain gained or lost in the added region, this extension does not appear to add or remove a functional module, and DeepLoc still calls both isoform and canonical nuclear at a calibrated, confident baseline, so nuclear/chromatin targeting is not disrupted. The unmoored, low-confidence nature of the extension's fold placement makes it unlikely to meaningfully engage or block the reader/adaptor functions central to CBX1's known heterochromatin and DNA-damage-response roles, though a subtle biophysical/representational change at the N-terminus cannot be ruled out.
The P2 shared-region RMSD (14.72 Å) is not trustworthy given low pTM (~0.35-0.39) and high inter-region PAE, so no core-refold claim is made; the extension's fold-confidence (P1) is real but its structural integration with the canonical body is unresolved, and mass-spec could not validate any unique peptide from this extension.
Folding
Evidence — click any tile for the differential-region detail
LLM reasoning
How well is the isoform-unique region conserved at the amino-acid level across primates (mean percent identity)? High AA identity in close relatives argues the alternative protein is real and translated, not a sequencing or annotation artifact. (Reading-frame intactness is shown as context.)
Comparison
Sequence conservation · primate
mean AA identity over the differential ORF is the score basis; frame intactness across the clade is context
| Conservation metric | Canonical ORF | Differential ORF |
|---|---|---|
| Mean AA identity | 100% | 95% |
| Frame intact (fraction of species) | 84% | 88% |
| Species aligned | 25 | 25 |
| Species frame-intact | 21 | 22 |
| Start codon conserved | 100% | 100% |
| Deepest intact species | Propithecus_coquereli | Propithecus_coquereli |
| Phylo depth (MRCA) | 7 | 7 |
The same amino-acid identity test across mammals. Conservation over deeper evolutionary distance is stronger evidence the ORF is under selection to be translated. (Reading-frame intactness is shown as context.)
Comparison
Sequence conservation · mammalian
mean AA identity over the differential ORF is the score basis; frame intactness across the clade is context
| Conservation metric | Canonical ORF | Differential ORF |
|---|---|---|
| Mean AA identity | 98% | 88% |
| Frame intact (fraction of species) | 15% | 45% |
| Species aligned | 20 | 20 |
| Species frame-intact | 3 | 9 |
| Start codon conserved | 100% | 83% |
| Deepest intact species | Callithrix_jacchus | Loxodonta_africana |
| Phylo depth (MRCA) | 6 | 12 |
phyloP / phastCons measure per-base evolutionary constraint. A high absolute phyloP over the differential region means the sequence itself is under strong purifying selection — evidence it is coding and does something. (Enrichment over the shared core is shown as context only.)
Comparison
Per-base conservation · differential region
absolute mean phyloP over the differential region is the score basis (high = strong purifying selection); the shared column and enrichment are context, not the claim
| Track | Differential | Shared | vs shared |
|---|---|---|---|
| phyloP mean | 3.29 | 4.95 | 0.663 |
| phastCons mean | 0.688 | 0.92 | — |
Start site · canonical vs isoform
initiation context + per-base conservation at each start codon
| Property | Canonical | Isoform |
|---|---|---|
| Start codon | ATG | CTG |
| Kozak context (−9..+4) | GCGGGCACTATGG | GCTGCCTTCCTGG |
| phyloP at start codon | 6.43 | 2.31 |
| phastCons at start codon | 1 | 0.916 |
| phyloP over Kozak window | 6.31 | 1.88 |
| phastCons over Kozak window | 1 | 0.777 |
| Kozak mismatch — full consensus | 3 | 4 |
| Kozak window GC content | 0.692 | 0.692 |
LLM reasoning
Is the alternative start used in more than one cell line? Reproducible initiation across independent samples argues against a one-off ribosome artifact.
Comparison
Per-cell-line usage · canonical vs isoform
Fisher q = 2.46e-12
| Cell line | Canonical | Isoform | p-value |
|---|---|---|---|
| HeLa | — | 3.46 | 0.000479 |
| K562 | — | 1.11 | 2.46e-12 |
| U2OS | — | 1.25 | 2.09e-07 |
| RPE1 Async | — | 1.13 | 5.13e-05 |
| RPE1 Que | — | 0.225 | 1.1e-07 |
How efficiently ribosomes initiate at this start (TIS) relative to background, per cell line. Strong, reproducible initiation supports a genuine translation event.
Comparison
Start-Site Usage · canonical vs isoform
ribosome initiation efficiency at the canonical start vs this alternative start, per cell line
| Cell line | Canonical | Isoform |
|---|---|---|
| HeLa | — | 0.0945 |
| K562 | — | 0.0511 |
| U2OS | — | 0.0193 |
| RPE1 Async | — | 0.0268 |
| RPE1 Que | — | 0.0111 |
Were tryptic peptides unique to the differential region detected by mass spec? Direct peptide evidence is the strongest proof the isoform protein exists.
Comparison
Peptide Evidence (canonical vs isoform)
isoform-unique peptides are direct evidence the alternative protein exists
| Feature | Canonical | Isoform |
|---|---|---|
| Tryptic peptides (in-silico) | 28 | 10 |
| Validated by mass-spec | 0 | 0 |
| Isoform-unique peptides | — | 10 |
Details
Peptide Evidence (canonical vs isoform)
- peptide MGATPPGDPTR 0–11
- peptide GATPPGDPTR 1–11
- peptide ASSAAPIPLGLLGAALSSVTLYTR 12–36
- peptide LAGTMGK 37–44
- peptide MGATPPGDPTRR 0–12
- peptide GATPPGDPTRR 1–12
- peptide RASSAAPIPLGLLGAALSSVTLYTR 11–36
- peptide ASSAAPIPLGLLGAALSSVTLYTRK 12–37
- peptide KLAGTMGK 36–44
- peptide LAGTMGKK 37–45
LLM reasoning
Do the predicted localization features (DeepLoc prediction, sorting signals, or membrane association) differ between the canonical and the isoform? A protein with changed localization features acts in a different cellular context.
Comparison
Subcellular localization (DeepLoc)
no localization-feature change
| Property | Canonical | Isoform |
|---|---|---|
| Predicted location | Nucleus | Nucleus |
| Sorting signals | Nuclear localization signal | Nuclear localization signal |
| Membrane | Soluble | Soluble |
Do N-terminal targeting signals (SignalP secretion, TargetP mitochondrial/chloroplast) differ between canonical and isoform? N-terminal changes most directly add or remove targeting peptides.
LLM reasoning
Does healthy human germline variation (gnomAD) avoid this region (depletion ratio < 1×), and is it intrinsically constrained (ESM-C constraint delta > 0)? Depletion of population variation plus high sequence constraint mean the region resists change — it is functionally important. Scored only where the unique region is canonical coding sequence, i.e. on truncations. gnomAD is a tolerance catalogue, not a disease one; disease/cancer variants (ClinVar / COSMIC) live in M2.
Comparison
Germline Variants · differential vs shared region
gnomAD (population) variant density per nucleotide — depletion ratio < 1× means healthy human variation avoids the region (constrained), the M1 basis alongside ESM-C constraint
| Variant set | Differential | Shared | Depletion ratio |
|---|---|---|---|
| gnomAD variants | 68 | 192 | 1.6× |
Sequence constraint (ESM-C) · differential vs shared region
lower LLR = more constrained; enrichment = differential / conserved
| Property | Differential | Shared | Enrichment |
|---|---|---|---|
| ESM-C mean LLR | -1.79 | -0.00445 | — |
| Constrained positions | 0 | 0 | — |
Predicted-damaging variants · germline (gnomAD)
AlphaMissense / ESM-C / LoF calls per region (length-normalized)
| Variant set | Differential | Shared | Enrichment |
|---|---|---|---|
| Scorable variants | 66 | 188 | 1.6× |
| Damaging variants | 5 | 103 | 0.22× |
| — of which loss-of-function | 5 | 13 | 1.7× |
| AlphaMissense-pathogenic | 0 | 58 | 0× |
Predictor scores · germline (gnomAD)
scored: 319 ESM-C · 168 AlphaMissense
| Property | Differential | Shared |
|---|---|---|
| Mean ΔLLR (ESM-C) | -0.853 | -5.64 |
| Min ΔLLR (ESM-C) | -6.18 | -13.1 |
| Mean AlphaMissense | — | 0.578 |
Are disease (ClinVar / COSMIC) variants enriched per nucleotide in the differential region versus the shared core (disease enrichment ratio ≥ 1×)? Disease variants concentrating in the unique region tie it to phenotype.
Comparison
Clinical-variant burden · differential vs shared region
counts per region; ratio is density-normalized (variants per nucleotide, differential ÷ shared — ≥1× = disease variants concentrate in the differential region, the M2 basis)
| Variant set | Differential | Shared | Enrichment |
|---|---|---|---|
| Disease variants | 4 | 89 | 0.2× |
| Pathogenic | 0 | 0 | — |
Predicted-damaging variants · disease (ClinVar/COSMIC)
AlphaMissense / ESM-C / LoF calls per region (length-normalized)
| Variant set | Differential | Shared | Enrichment |
|---|---|---|---|
| Scorable variants | 4 | 88 | 0.21× |
| Damaging variants | 0 | 66 | 0× |
| — of which loss-of-function | 0 | 9 | 0× |
| AlphaMissense-pathogenic | 0 | 43 | 0× |
Predictor scores · disease (ClinVar/COSMIC)
scored: 319 ESM-C · 168 AlphaMissense
| Property | Differential | Shared |
|---|---|---|
| Mean ΔLLR (ESM-C) | -2.42 | -7.53 |
| Min ΔLLR (ESM-C) | -5.89 | -13.9 |
| Mean AlphaMissense | — | 0.68 |
LLM reasoning
Does the gained/lost region fold into ordered structure (pLDDT) and shift the protein's biophysical character (charge, hydropathy, disorder)? Structured, biophysically distinct regions are more likely to be functional.
Comparison
Structure (ESMFold2) · canonical vs isoform
TM-score 0.454 · RMSD 4.65 Å · 3 interface contacts
| Metric | Canonical | Isoform |
|---|---|---|
| pLDDT (whole protein) | 0.822 | 0.833 |
Fold confidence · differential vs shared region
mean ESMFold2 pLDDT in each region
| Metric | Differential | Shared | Enrichment |
|---|---|---|---|
| pLDDT | 0.76 | 0.733 | 1 |
The shared region is the stretch of protein identical in both the isoform and the canonical (the canonical body for an extension; the post-truncation body for a truncation). Folded in both contexts it normally comes out nearly identical, so its Cα backbone RMSD ≈ 0. A high shared-region RMSD (superposed on the shared residues only, from the ESMFold2 structures) means the extension or truncation reorganizes how the retained region folds — a rare, high-interest functional signal. TM-score is a length-normalized companion; the check only fires when both structures are confidently folded, and uORF/altORF isoforms (no shared region) are not evaluable.
Comparison
Shared-region structural change
Cα RMSD superposed on the shared residues only; TM-score is length-normalized · shared-region Cα RMSD 14.7 Å · shared TM-score 0.48 · shared region 185 aa · min shared pLDDT 0.822 · global TM-score 0.454 · global RMSD 4.65 Å
| Property | Canonical | Isoform |
|---|---|---|
| Shared-region pLDDT | 0.822 | 0.849 |
Does the differential region contain actual secondary structure — a helix or strand — rather than coil? ESMFold2 predicts coordinates but no secondary structure, so elements are assigned from those coordinates (P-SEA, Cα geometry). Because they are derived from a prediction, each element carries its own mean pLDDT: a geometrically clean helix running through a disordered stretch is geometry fitted to a guess, so BOTH length and confidence are required to score. The direction differs by ORF type — an extension GAINS the element (a candidate functional addition), a truncation LOSES one, and there the element is read off the canonical structure because the removed segment exists only in it. This says nothing about whether the element is integrated with the rest of the fold; the contact and PAE evidence in this same category answers that.
Comparison
Secondary structure · differential vs shared region
helices and strands assigned from the predicted coordinates (P-SEA)
| Metric | Differential | Shared |
|---|---|---|
| Alpha helices | 0 | 3 |
| Beta strands | 1 | 2 |
| Longest element (aa) | 22 | 18 |
| Mean pLDDT | 0.80 | 0.90 |
Elements and coordinates
1 in the differential region, 5 in the shared core — residue numbering is 1-based on the protein holding the region
| Isoform-unique | Shared core |
|---|---|
| beta strand 7–28 22 aa · pLDDT 0.80 | alpha helix 102–116 15 aa · pLDDT 0.88 |
| — | beta strand 123–140 18 aa · pLDDT 0.80 |
| — | beta strand 184–189 6 aa · pLDDT 0.93 |
| — | alpha helix 190–195 6 aa · pLDDT 0.96 |
| — | alpha helix 198–208 11 aa · pLDDT 0.96 |
Below threshold
1 in the differential region, 4 in the shared core — shorter than 6 aa or below pLDDT 0.70, so not counted above
| Isoform-unique | Shared core |
|---|---|
| beta strand 36–45 10 aa · pLDDT 0.68 | beta strand 76–80 5 aa · pLDDT 0.95 |
| — | beta strand 91–94 4 aa · pLDDT 0.95 |
| — | beta strand 172–176 5 aa · pLDDT 0.92 |
| — | beta strand 209–213 5 aa · pLDDT 0.88 |
LLM reasoning
Does the isoform gain or lose a real InterPro functional domain in the differential region (disorder/structural-only signatures excluded)? Gaining or losing a domain changes function directly.
Comparison
Domains & motifs (canonical vs isoform)
gained = only in the isoform; lost = only in the canonical
| Feature | Canonical | Isoform |
|---|---|---|
| InterPro domains | 15 | 16 |
| Short linear motifs | 1 | 3 |
Details
Domains & motifs (canonical vs isoform)
- gained Coil domain
- gained CDK_Sites_SP_TP motif
- gained SH3_ClassII motif
Biophysical character of the isoform-differential region versus the shared canonical core — pI, hydropathy, charge, disorder and related properties. Descriptive; the folding (P1) score keys off the GRAVY / charge / disorder deltas.
Comparison
Biophysics · differential vs shared region
highlighted rows are enriched in the differential region
| Property | Differential | Shared | Enrichment |
|---|---|---|---|
| Isoelectric point (pI) | 11.4 | 4.54 | 2.52 |
| Hydropathy (GRAVY) | 0.173 | -1.28 | -0.135 |
| Fraction charged | 0.122 | 0.46 | 0.266 |
| Disorder fraction | 0.139 | 0.232 | 0.6 |
| Disorder-promoting | 0.634 | 0.681 | 0.931 |
| Low-complexity fraction | 0.317 | 0.222 | 1.43 |
| Prion-like fraction | 0.244 | 0.2 | 1.22 |
| LLPS score | 0.163 | 0.192 | 0.848 |
| π–π propensity | 0.0976 | 0.173 | 0.564 |
| Aromaticity | 0.0244 | 0.0703 | 0.347 |
| Instability index | 29.4 | 53 | 0.554 |
| Shannon entropy | 3.34 | 3.9 | 0.857 |
| Normalized complexity | 0.773 | 0.902 | 0.857 |
Sparse-autoencoder (SAE) interpretability features on the ESM-C residual stream, comparing the isoform against the canonical protein. This is the S3 criterion — a presence check that fires when any interpretable feature is gained or lost. Only the top 30 features per category (ranked by activation, prevalence ≥ 2) are saved; each table shows the top 10 interpretable features, and generic / "unknown" features are hidden unless you expand it. What are SAE features? →
Part 1 · Global feature comparison
Which SAE features fire in the isoform vs the canonical protein, across the whole sequence.
Isoform-only features — 135 total (prevalence ≥ 2)
| Feature | Label | Description | Activation | Prevalence |
|---|---|---|---|---|
| #4992 | Small-residue disordered N-terminus | Context-dependent free N-terminus signature: short, disordered amino-terminal stretches beginning with the initiator Met followed by residues compatible with MAP cleavage or N-acetylation (e.g., M–G/A/S/T/V, and other contexts including M–D/E/L); strongest emphasis on small residues at positions 2–3; activation is absent when the native N-terminus is structured or starts with a signal peptide/transmembrane segment. | 12.65 | 2 |
| #448 | N-terminal leader/targeting segments | N-terminal leader/targeting segments: the feature activates on the extreme N-terminus (first 2–5 residues) of proteins, especially the N-region of signal peptides and organelle transit peptides, and more generally on short, disordered/low-complexity N-terminal tails; strongest at residue ~3 and largely independent of amino-acid identity, occurring across all taxa and functions. | 12.09 | 2 |
| #9852 | N-terminus activation; hydrophobic helices secondary | Extreme N-terminal regions of proteins, starting at the initiator methionine and extending over a short stretch, with occasional secondary activation at internal hydrophobic/amphipathic helices. | 11.45 | 40 |
| #11504 | Initiator methionine at M1 | Initiator methionine at the very start of the polypeptide chain (M1), i.e., the translation start residue, independent of protein family, taxonomy, membrane association, or presence of signal/leader/propeptide regions; often annotated as post-translationally removed and typically situated in a flexible, coil-like N-terminus. | 10.97 | 2 |
| #8545 | N-terminus accessibility sensor | Detector of accessible peptide chain termini—primarily the extreme N-terminus (initiator methionine and immediate neighbors) in flexible, unstructured tails; position-specific rather than residue-specific—with occasional weak recognition of the C-terminus; common but not universal. | 10.86 | 2 |
| #3609 | Position-2 N-terminus residue detector | Detector of the identity of the second residue at the extreme N-terminus of proteins, with strongest preference for small residues (especially Ser, Ala, and Gly; also Thr/Pro), i.e., an N-terminal position-2 signal commonly associated with generic, disordered starts and co-translational N-terminus processing | 10.31 | 3 |
| #9998 | Weak N+3 positional cue | Absolute N-terminal positional cue centered near the fourth residue (N+3), independent of amino-acid identity; most often within signal/transit peptides of precursors but also present in non-secreted proteins; the effect is weak and often does not produce a discrete detected site. | 9.57 | 2 |
| #14493 | Short hydrophobic helical runs | Short hydrophobic helical segments rich in I/V/F (sometimes M), recognized in both membrane-targeting/insertion contexts (signal peptides, signal-anchors, transmembrane helices) and in hydrophobic helical patches within otherwise soluble small proteins. The feature flags contiguous aliphatic/hydrophobic runs across diverse proteins, including viral membrane/accessory proteins, secreted peptide/effector precursors, small ORFs/microproteins, and short basic/uncharacterized human proteins. | 9.28 | 3 |
| #11159 | Glycine residue identity detector | Residue-identity detector for glycine (G), largely context-agnostic but with a tendency to fire on glycines in flexible/disordered, often N-terminal or membrane-proximal regions; also active on glycines within transmembrane helices; broadly taxon- and function-independent. | 8.19 | 4 |
| #7903 | Arginine-rich polybasic patches | Basic polycationic patches enriched in arginine (often with lysine) within low-complexity/disordered regions, frequently N-terminal. These include classical monopartite NLSs, nucleic acid–binding basic tails, signal peptide n-regions and cytosolic juxtamembrane segments (positive-inside rule), and basic clusters in secreted precursors; typically depleted of acidic residues. | 8.12 | 5 |
| #5337 | Pro/RGG-rich low-complexity IDRs | Intrinsically disordered, low‑complexity basic segments at termini and long loops, enriched in Pro/Gly and/or Arg/Ser (including RG/RGG or S/T clusters), common across eukaryotic and viral proteins. | 7.68 | 33 |
| #5241 | Early N-terminal positional signal | Generic early N-terminus positional signal peaking at residue ~5–7, found broadly within the disordered N-tail of proteins, including both cleavable leader/targeting segments (signal peptides, propeptides) of precursor proteins and disordered N-terminal regions of mature proteins; residue identity is not specific | 7.34 | 2 |
| #14646 | Short N-terminal leader segment | Short N-terminal segments within the first ~10–40 amino acids, often Lys/Arg-enriched but sometimes purely hydrophobic. These commonly correspond to Sec-type signal peptides, signal anchors, or other short N-terminal segments preceding/overlapping the first hydrophobic helix, and also include disordered N-terminal tails in soluble proteins. | 7.32 | 39 |
| #3670 | Terminal disordered peptide detector | Short unstructured peptides and N-terminal/C-terminal segments of larger proteins; occasional firing near N-terminal leader/signal sequences and basic targeting motifs (e.g., NLS); broadly residue-tolerant with bias toward hydrophobic, Gly/Pro, and basic residues; appears in viral accessory proteins, microproteins, and secreted precursors but also in diverse bacterial/archaeal enzymes, with sparse activation overall. | 7.26 | 7 |
| #6826 | Sparse internal aliphatic peaks | Sparse residue-level feature firing at scattered internal positions across diverse proteins, with peaks often landing on aliphatic/small residues (A, P, V, I, L) flanked by polar or charged context, frequently in loop-like or low-complexity stretches. | 6.56 | 2 |
| #7384 | Leucine-biased disordered hydrophobic segments | Leucine-biased recognition of intrinsically disordered, low-complexity hydrophobic segments, including disordered insertions within otherwise structured proteins. | 6.43 | 6 |
| #15324 | Low-complexity disordered segments | A detector of low‑complexity, intrinsically disordered, Ser/Thr/Gly/Pro/Ala‑rich segments (often N‑terminal), including proline‑rich stretches and occasional closely spaced cysteines; these regions occur in viral microproteins, secretory precursors/extracellular repeats, membrane‑proximal tails, and organelle transit peptides rather than in a specific folded domain. | 5.94 | 7 |
| #8488 | Ala/Thr-rich N-terminal disorder | Ala/Thr-enriched composition feature with a bias for low-complexity intrinsically disordered regions and N-terminal prepro/signal-peptide segments; the feature reflects small-residue (A/T, secondarily S/P) composition often in flexible regions but also appears at A/T residues in some structured contexts. | 5.69 | 10 |
| #787 | Basic disordered microprotein regions | The feature activates within small proteins, viral accessory/regulatory proteins, and short ORFs, often on disordered or low-complexity segments. Activation is frequently associated with basic-residue–containing patches (Lys/Arg) but is not restricted to a particular position in the protein. | 5.64 | 10 |
| #13063 | Serine-rich region detector | Serine residues across a wide range of structural contexts, with a bias toward serine-rich segments in intrinsically disordered or low-complexity regions (often N-terminal tails), but also firing on serines within transmembrane helices and other structured contexts; the feature behaves primarily as a serine detector with some preference for S-rich patches. | 5.57 | 4 |
| #5021 | Polybasic intrinsically disordered regions | Low-complexity, often intrinsically disordered regions in short proteins, precursors, and microproteins across taxa, including viral accessory proteins, neuropeptide/hormone precursors, sperm nuclear proteins, flexible enzyme tails/inserts, and antisense-derived/uncharacterized microproteins. Activation is biased toward, but not restricted to, basic (Lys/Arg) clusters, with frequent peaks also at Pro/Ser/Gly residues in disordered context. | 5.31 | 4 |
| #4579 | Acidic polar disordered tails | Low-complexity, disorder-associated segments enriched in Asp/Asn/His/Tyr and depleted of Lys/Arg—frequently N- or C-terminal propeptides and regulatory tails, and occasionally short polar/acidic stretches within other contexts including amphipathic helices in small viral proteins. | 5.26 | 2 |
| #9234 | Charged low-complexity disordered regions | Intrinsically disordered, low-complexity segments enriched for Lys and acidic residues (E/D), plus N/Q, that include both polyampholyte and basic-residue tracts; typically flexible terminal tails and linker regions in diverse eukaryotic and viral proteins, including many small/secreted proteins and membrane-protein loops | 5.20 | 2 |
| #3123 | Low-complexity IDRs and signal peptides | Low-complexity, intrinsically disordered segments enriched in Pro, Ser/Thr and basic residues: the feature prefers compositionally biased coils and polar/low-complexity stretches, and also fires on hydrophobic LVIA-rich segments such as the h-region of signal peptides; it activates broadly across taxa in secreted/membrane proteins, small regulatory peptides, and flexible disordered regions of soluble enzymes. | 5.18 | 5 |
| #3278 | PRQH-rich disordered tails | Compositionally biased, intrinsically disordered low‑complexity segments enriched in Pro/Arg/Gln/His (frequent PR/PQ tracts, Arg- and His-clusters), with occasional sensitivity to Leu‑rich helical stretches (signal peptides or leucine zippers); typically terminal, widespread across taxa, and common in nucleic‑acid–binding and secreted proteins. | 5.04 | 3 |
| #4913 | Low-complexity disordered segments | Low-complexity and compositionally biased sequence stretches, frequently in disordered or flexible regions including N/C-terminal tails, but also extending to short low-complexity tracts within otherwise ordered proteins; peaks favor Ser and other polar residues, with additional hits on aromatic (Trp/Phe/Tyr) and basic-rich contexts, suggesting a composition/local-context signal rather than a specific fold or sharply defined motif | 4.95 | 2 |
| #3660 | Low-complexity basic IDRs and micro-TMs | Generic low-complexity segments that are intrinsically disordered, proline‑rich and/or Lys/Arg‑biased (often also enriched in Ser/Thr) with intermittent hydrophobic residues, together with short hydrophobic helices in small membrane proteins; these include propeptide/processing regions of secreted precursors, basic disordered segments in viral proteins, mucin‑like S/T‑rich patches in glycoproteins, surface loops in enzymes, and single‑pass transmembrane microproteins (e.g., organellar gene products) rather than a specific folded domain. | 4.80 | 7 |
| #9096 | Intrinsically disordered low-complexity regions | Compositionally biased, intrinsically disordered low‑complexity segments (often short tandem‑repeat regions) enriched in small/flexible residues and sometimes cysteine/glutamine, commonly occurring in N‑terminal/propeptide stretches of secreted precursors and small unstructured proteins; peaks can also occur in flexible terminal/loop segments within otherwise structured proteins (including enzymes); peaks can fall at residues within simple repeats or immediately flanking processed peptide segments, but the unifying concept is IDR/low‑complexity or flexibly disordered segments rather than a specific domain. | 4.73 | 12 |
| #1907 | IDRs and signal peptides | Compositionally biased non-globular segments — intrinsically disordered low-complexity regions enriched in small residues (Ala/Pro/Gly/Ser) and short basic patches (e.g., NLS-like) — with additional firing on some hydrophobic helical segments such as signal peptides. | 4.71 | 9 |
| #12009 | Hydrophobic and leader-like segments | Hydrophobic and/or small/turn-forming sequence segments, including but not limited to classical targeting leaders (signal peptides, signal-anchor TMs, mitochondrial/chloroplast transit peptides). The feature is compositional and responds to short hydrophobic patches, propeptides, low-complexity/Pro/Ser/Thr-rich segments, and amphipathic/disordered stretches that may occur anywhere in the sequence, though early/leader regions are common. | 4.63 | 7 |
Canonical-only features — 14 total (prevalence ≥ 2)
| Feature | Label | Description | Activation | Prevalence |
|---|---|---|---|---|
| #3476 | Broad-spectrum PPIase detector | Parvulin/PpiC-type peptidyl-prolyl cis-trans isomerase domain detector that recognizes bacterial PrsA-family foldase chaperones and related PPIase domains, including both catalytically active enzymes and inactive PPIase-like homologs (e.g., virulence-associated Mip, lumenal/periplasmic FKBPs) | 4.06 | 2 |
| #16313 | Catalytic-domain N-terminus motif | Short N-terminal segments at the start of the first catalytic domain, often coinciding with the first β-strand and active-site/substrate-binding residues, marking the beginning of the mature/folded core rather than a generic basic patch. | 2.31 | 3 |
| #1949 | PMT FxRFLDxxQY helix | A conserved internal alpha-helical motif in plant/protist phosphoethanolamine N-methyltransferases (PMTs), centered on the sequence "FQRFLDNVQY" (and close variants). | 2.01 | 2 |
| #9952 | Ordered helical assembly scaffold | A broad “ordered helical/assembly scaffold” signature: long, structured segments—often alpha-helical (TM helices, coiled-coils, helical hairpins) and stable cores of ligand/effector-binding domains—that mark architectural elements used for signal transfer and assembly across diverse molecular machines (signaling, secretion, transport, nucleotide/quinone metabolism, and viral replication), rather than specific catalytic micro-motifs. | 1.95 | 2 |
| #4487 | Sparse internal residue peaks | A diffuse, low-specificity feature with sparse residue-level activations distributed across a heterogeneous set of proteins, including bacterial hydrolases (amidases/Ntn-fold), acyl/D-alanyl carrier proteins, Sec translocon γ-subunits (bacterial, archaeal, and eukaryotic SSS1/SecE), and eukaryotic C2H2 zinc-finger transcription factors (ZIC family). | 1.93 | 2 |
| #8749 | RPA/SSB DBD-C zinc finger | Primarily single-stranded nucleic-acid–binding proteins of the RPA/SSB family and related DNA replication/repair auxiliaries, with the strongest residue-level signal concentrated in their C-terminal DNA-binding/zinc-finger module rather than in the N-terminal/central OB barrels; telomere-binding factors (POT1/TPP1/CST), exosome-cap OB subunits, plant organellar SSBs, and other OB-fold ssDNA/RNA-binding proteins also fall within the feature's scope. | 1.93 | 2 |
| #9937 | C-terminal accessory glycan region | C-terminal accessory region downstream of the catalytic core in glycan‑modifying enzymes—most prominently FGly‑dependent sulfatases but also other carbohydrate‑active hydrolases and lumenal/membrane glycan transferases. The feature targets the distal C‑terminal subdomain or tail (often a coil→alpha‑helix segment, or a flexible basic tail in SULFs) that contributes to substrate engagement/localization rather than the catalytic nucleophile/metal sites. | 1.82 | 2 |
| #11298 | Noncatalytic nucleotide-binding helical scaffold | An alpha-helical scaffold signature found inside nucleotide-binding enzyme cores—prominently in kinesin motor domains and in analogous helical subdomains of other NTPases, kinases, and TIR domains; the signal concentrates on non-catalytic helices adjacent to, but not coincident with, ATP/catalytic motifs, consistent with structural elements used for conformational coupling or interface formation. | 1.81 | 2 |
| #8679 | C-terminal domain boundary marker | Residue-level marker of structural boundaries—especially C‑terminal ends of compact domains or segments—often at helix/turn caps or immediately into linkers/disordered tails; occasionally a single interior residue within a long helix | 1.73 | 2 |
| #5824 | Pro/Gly-rich beta-strand caps | Short Pro/Gly-enriched β-strand edge/turn motifs at strand–loop (and loop–helix) junctions—often at the N-terminus of a domain—i.e., β-hairpin turns and strand N-caps used as structural connectors rather than catalytic sites | 1.71 | 2 |
| #5462 | RBP IDRs and coiled-coils | Extended regions outside the canonical RRM RNA-binding core in eukaryotic (often nuclear) RNA/RNP proteins. The feature lights up both long compositionally biased intrinsically disordered segments (Arg/Gly/Ser/Pro–rich low-complexity stretches, RG/RGG boxes, poly-G/P, and mixed basic/acidic runs) and conserved coiled-coil/helical dimerization regions adjacent to the RNA-binding modules (notably the DBHS-family coiled coil). | 1.70 | 2 |
| #7907 | Domain-edge mixed secondary surface segment | A broad surface segment at domain boundaries—often a region just after a signal peptide or N-terminal cap in secreted/periplasmic carbohydrate-active enzymes, and analogous regions in other proteins—that mixes disordered coils with short beta-strands/loops and lies outside catalytic residues but can host substrate-contact or modification sites. | 1.67 | 2 |
| #7891 | Short localized secondary-structure segments | Short, contiguous local secondary-structure segments (~15-40 aa) within specific small RNA/protein-interaction domains, most prominently CpcD-like domains (in phycobilisome linker polypeptides and ferredoxin-NADP reductases) and R3H domains (in RNA-binding regulatory proteins). Peaks center on a β-strand/turn or short helix within the domain fold. | 1.65 | 2 |
| #3180 | Noncatalytic C-terminal interaction regions | Non-catalytic C-terminal interaction regions—long, charged/polar terminal domains or tails (often low-complexity or repeat-based) that mediate assembly, trafficking, and RNA/protein binding, rather than enzymatic catalysis; includes structured terminal modules (e.g., NTF2-like) and flexible extensions, with a strong bias toward the final domain/segment of the protein. | 1.48 | 2 |
Shared features by |Δ| activation — 941 total (prevalence ≥ 2)
| Feature | Label | Description | Δ activation | Iso act. | Canon act. | Iso prev. | Canon prev. |
|---|---|---|---|---|---|---|---|
| #6125 | Disordered PTM/cleavage SLiMs | Short linear motifs in intrinsically disordered/low-complexity regions, often S/T/P/G- and aromatic-rich, that include proteolytic-processing/PTM hotspots and flexible linkers adjacent to transmembrane segments; the feature favors flexible coils or short amphipathic helices and systematically avoids hydrophobic transmembrane cores | +6.08 | 8.45 | 2.37 | 10 | 7 |
| #14891 | Disordered termini and cleavage detector | Residue-level detector of intrinsically disordered, flexible termini and proteolytic processing junctions—especially N-termini (including the first residue of the mature chain after propeptide/leader cleavage)—with no strict amino‑acid specificity and a bias toward small/charged residues; common in small, secreted/precursor and viral proteins but also present in disordered regions of diverse proteins. | +5.26 | 7.79 | 2.53 | 11 | 3 |
| #7523 | Diffuse isoleucine composition bias | Weak global preference for isoleucine (and closely related aliphatic hydrophobes), captured primarily as a diffuse composition signal rather than discrete site recognition. | +5.13 | 11.55 | 6.43 | 5 | 4 |
| #678 | Sparse terminal residue peaks | Sparse residue-level activations distributed across short proteins and protein segments, with frequent firing in disordered N-terminal/C-terminal tails as well as in some short transmembrane or compact domains | +4.75 | 7.39 | 2.64 | 13 | 2 |
| #6356 | Ser/Gly/Pro-rich low-complexity IDRs | Often fires on intrinsically disordered, low-complexity regions (IDRs) enriched in Ser/Gly/Pro and basic (Lys/Arg) residues, frequently in secreted/viral proteins, but also activates on short peptides and small folded proteins with no clear sequence motif. Includes Ser-rich PTM-prone segments within unstructured regions. | +4.49 | 8.24 | 3.75 | 13 | 8 |
| #2692 | Compositionally biased disordered LCRs | Feature targets compositionally biased, intrinsically disordered low‑complexity regions with long contiguous runs strongly enriched in small/polar (Gly/Ser/Asn/Thr) or acidic (Asp/Glu) residues; occasional activation on highly basic protamine‑like LCRs. Mere disorder, generic tails, or coiled‑coils are insufficient without such compositional bias. | +4.47 | 7.58 | 3.11 | 11 | 6 |
| #3503 | Cationic/hydrophobic low-complexity segments | Compositionally biased and low-complexity segments enriched in hydrophobic (L/V/I/A), basic (K/R), and Ser/Thr/Pro residues (with underrepresented Trp). The feature marks both cationic Ser/Thr/Pro-enriched segments and hydrophobic leader-like stretches, capturing N-terminal targeting peptides, micropeptides, and internal polybasic/low-complexity motifs in diverse proteins. | +4.42 | 6.48 | 2.06 | 9 | 2 |
| #13480 | Alanine and small-residue enrichment | A sequence-composition feature that detects small, non-aromatic residues—especially alanine—and their enrichment, lighting up individual Ala residues and stretches enriched in A/G/S/T/P/V (alanine-rich, low-complexity segments) irrespective of protein function or secondary structure. | +4.27 | 8.70 | 4.43 | 17 | 8 |
| #14866 | Intrinsic disorder and low-complexity regions | Intrinsic disorder/low-complexity signal: the feature marks compositionally biased, non-globular regions enriched in polar/charged and small residues (S/T/E/D/R/K/G/P; often A/L), typically flexible N- or C‑terminal tails, linkers, and propeptides that host short linear motifs or processing sites, while avoiding structured domains. | +4.19 | 8.41 | 4.22 | 37 | 35 |
| #9237 | Glycine detector in disordered regions | Residue-identity detector for glycine (G), with a mild bias toward intrinsically disordered/low-complexity segments (often including very N-terminal residues); pan-taxonomic and function-agnostic, but activation is sparse and frequently absent, becoming detectable mainly in glycine-rich low-complexity contexts; occasional weak responses to other small/polar residues. | +4.17 | 10.06 | 5.89 | 15 | 10 |
| #6384 | Widespread activation in folded proteins | Proteins showing broad, diffuse activation across large fractions of their sequence, with peaks frequently occurring in folded regions of single-domain or multi-domain enzymes and structural proteins. | +3.94 | 6.51 | 2.57 | 68 | 23 |
| #4740 | Disordered loops and linkers | Low-complexity / flexible regions enriched in polar (Ser/Thr/Asn/Gln), basic (Lys/Arg), and Gly/Pro residues, including disordered linkers, propeptides, surface loops, and flexible segments of membrane and globular proteins across taxa (including viral proteins, small secreted peptide precursors, and membrane-protein cytosolic loops) | +3.87 | 6.82 | 2.94 | 13 | 10 |
| #9395 | Short terminal targeting segment | Terminal or short internal targeting/interaction segment: a short, compositionally biased or hydrophobic patch near a protein terminus, or occasionally an internal low‑complexity loop, that functions as a signal peptide or signal‑anchor (Sec/ER/thylakoid/Tat entry) when hydrophobic, or as a basic/Ser/Lys/Arg/Pro‑rich low‑complexity segment that mediates localization or partner binding in soluble proteins; typically the single prominent terminal segment of short secreted, membrane, or regulatory proteins | +3.60 | 6.49 | 2.89 | 9 | 5 |
| #12620 | Disordered low-complexity regions | Intrinsically disordered, low‑complexity, compositionally biased regions/tails (IDRs), typically enriched in Ser/Gly/Pro/Ala/Thr and often occurring as acidic (Asp/Glu) or basic (Arg/Lys) tracts or Gln/Asn‑/Gln‑rich repeats; these segments are common in secreted precursors, viral proteins, micropeptides, and testis‑associated proteins, and can also occur as low‑complexity termini or surface loops appended to otherwise folded enzymes; they frequently coincide with low predicted structural confidence. | +3.56 | 6.38 | 2.83 | 10 | 2 |
| #1260 | Disordered low-complexity basic segments | Broadly distributed feature with notable but non-exclusive enrichment in compositionally biased, intrinsically disordered low-complexity regions (LCRs), including Gly/Ser/Pro-rich and Arg/Lys-rich tracts or simple repeats. | +3.54 | 6.32 | 2.77 | 12 | 7 |
| #6713 | Polar low-complexity disordered repeats | Polar low-complexity intrinsically disordered regions enriched in small/polar residues (Asn, Ser, Thr, Gly, often with Pro), including N-rich repeat tracts of protist proteins, S/T/Gly-rich disordered tails of viral and host proteins, and mucin-like extracellular repeats. | +3.42 | 6.82 | 3.40 | 11 | 8 |
| #7780 | Gly/Pro-rich disordered regions | A generic signature of intrinsically disordered, low‑complexity regions enriched in glycine/proline and charged/polar residues, common in secreted peptide precursors and collagen-like segments as well as flexible terminal/linker regions of diverse proteins; it largely avoids well‑structured catalytic cores and transmembrane helices. | +3.26 | 6.76 | 3.50 | 13 | 11 |
| #583 | Asp-biased disordered region detector | Detector of Asp residues embedded in low-complexity/disordered segments, with a strong bias for Asp over Glu; fires broadly in polar or acidic-enriched tracts and only weakly on isolated Asp in structured domains. | +3.25 | 9.30 | 6.05 | 21 | 21 |
| #3926 | Intrinsic disorder/low-complexity detector | Intrinsic disorder/low-complexity detector: activates on flexible, polar/charged, Q/N/S/T/E/D/K/R- and G/P-enriched segments—often N‑terminal tails—typical of intrinsically disordered regions in small secreted peptides/toxins and viral or eukaryotic regulatory proteins, including Q/N‑rich tracts; signals low-confidence coils or labile helices rather than folded domains. | +3.01 | 5.73 | 2.72 | 8 | 6 |
| #627 | Generic serine detector | Generic serine detector: the feature marks serine residues, with weaker affinity for threonine (and occasionally proline), and is especially prominent in low-complexity S/T-rich stretches; it is context- and structure-agnostic and does not correspond to a specific functional motif. | +2.95 | 8.61 | 5.66 | 18 | 13 |
| #11023 | Low-complexity disordered regions | Intrinsically disordered, low‑complexity segments—often N‑terminal tails or leader/signal regions—enriched in Ser/Pro/Gly/Arg and predicted as coils/low‑confidence structure; the feature marks flexible, poorly structured regions rather than a specific function | +2.91 | 5.51 | 2.60 | 11 | 4 |
| #4358 | Low complexity disordered regions | Compositionally biased, low-complexity segments enriched in small/polar residues (serine, glycine, asparagine, proline) and often basic (lysine/arginine) or acidic tracts; the feature preferentially marks IDRs in N/C-terminal tails and repeat-rich regions, though some activations also occur within folded domains. | +2.90 | 5.60 | 2.70 | 13 | 8 |
| #13897 | Proline-rich IDR detector | Detector of proline residues, with strongest signal in proline-rich, intrinsically disordered, low-complexity segments (often at termini, propeptides, and surface-exposed regions). The feature also activates on more isolated prolines in compact protein contexts, including within transmembrane helices, though typically at lower intensity than in extended Pro-rich disorder. | +2.84 | 7.40 | 4.56 | 14 | 7 |
| #5474 | Disordered Arg–Pro motifs | Detector enriched on basic/polar residues near Proline in flexible or disordered protein segments, with frequent firing on Arg and Pro within Arg-Pro (R-P, RRP) motifs and on residues preceding Pro (e.g., S/T-P) inside low-complexity or disordered stretches. Also fires sporadically in non-disordered regions on similar local sequence contexts. | +2.75 | 4.72 | 1.98 | 8 | 3 |
| #2319 | Threonine residue detector | Residue-identity detector for threonine (Thr): activates on individual Thr residues across diverse proteins, with a mild enrichment in low-complexity/disordered, repeat-rich, and secretory regions; still marks Thr within well-structured domains; occasional weak spillover to serine. | +2.73 | 7.64 | 4.91 | 12 | 5 |
| #6757 | N-terminal disordered propeptide signature | Generic signature of small proteins, intrinsically disordered and low-complexity segments, and precursor/propeptide regions (S/T/P/G/A-enriched), often at N-termini, rather than structured catalytic or metal-binding motifs | +2.73 | 4.72 | 1.99 | 17 | 2 |
| #3240 | Short low-complexity peptides/proteins | Activation across short proteins and peptide-like sequences from diverse taxa, with a tendency to fire in disordered or low-complexity stretches but also extending into structured regions of small chains | +2.67 | 5.02 | 2.35 | 10 | 7 |
| #9005 | Unknown generic feature | Unknown generic feature | +11.73 | 22.28 | 10.55 | 204 | 165 |
| #14534 | Unknown generic feature | Unknown generic feature | +11.59 | 26.19 | 14.60 | 223 | 181 |
| #1803 | Unknown generic feature | Unknown generic feature | +7.36 | 24.66 | 17.30 | 224 | 182 |
Part 2 · Differential coordinates
Features firing on just the isoform-unique (extension / alt-frame) residues.
Unique-region features — 265 distinct (prevalence ≥ 2)
| Feature | Label | Description | Activation | Prevalence |
|---|---|---|---|---|
| #13702 | Regulatory low-complexity IDRs | Eukaryotic intrinsically disordered, low‑complexity regulatory segments (terminal tails and inter‑domain linkers) enriched in Pro/Ser/Thr/acidic composition and short linear motifs for modular protein–protein interactions—frequently including WW‑domain–binding PY/PPxY segments and PTM hotspots—with reduced but not absent activation within folded recognition/catalytic domains (WW, PTB/PID, chromo/chromoshadow, SET). | 11.77 | 40 |
| #9852 | N-terminus activation; hydrophobic helices secondary | Extreme N-terminal regions of proteins, starting at the initiator methionine and extending over a short stretch, with occasional secondary activation at internal hydrophobic/amphipathic helices. | 11.45 | 40 |
| #8545 | N-terminus accessibility sensor | Detector of accessible peptide chain termini—primarily the extreme N-terminus (initiator methionine and immediate neighbors) in flexible, unstructured tails; position-specific rather than residue-specific—with occasional weak recognition of the C-terminus; common but not universal. | 10.86 | 2 |
| #9237 | Glycine detector in disordered regions | Residue-identity detector for glycine (G), with a mild bias toward intrinsically disordered/low-complexity segments (often including very N-terminal residues); pan-taxonomic and function-agnostic, but activation is sparse and frequently absent, becoming detectable mainly in glycine-rich low-complexity contexts; occasional weak responses to other small/polar residues. | 10.06 | 5 |
| #2114 | Chromodomain and integrase-proximal activation | A feature that activates broadly across chromodomain-containing proteins and other chromatin-associated factors, as well as on retrotransposon Gag-Pol polyproteins (in integrase-proximal regions). Activation is distributed widely along the chain with peaks tending to fall within or adjacent to chromodomain folds rather than on catalytic enzyme cores. | 9.76 | 35 |
| #14493 | Short hydrophobic helical runs | Short hydrophobic helical segments rich in I/V/F (sometimes M), recognized in both membrane-targeting/insertion contexts (signal peptides, signal-anchors, transmembrane helices) and in hydrophobic helical patches within otherwise soluble small proteins. The feature flags contiguous aliphatic/hydrophobic runs across diverse proteins, including viral membrane/accessory proteins, secreted peptide/effector precursors, small ORFs/microproteins, and short basic/uncharacterized human proteins. | 9.28 | 3 |
| #13480 | Alanine and small-residue enrichment | A sequence-composition feature that detects small, non-aromatic residues—especially alanine—and their enrichment, lighting up individual Ala residues and stretches enriched in A/G/S/T/P/V (alanine-rich, low-complexity segments) irrespective of protein function or secondary structure. | 8.70 | 9 |
| #627 | Generic serine detector | Generic serine detector: the feature marks serine residues, with weaker affinity for threonine (and occasionally proline), and is especially prominent in low-complexity S/T-rich stretches; it is context- and structure-agnostic and does not correspond to a specific functional motif. | 8.61 | 5 |
| #6125 | Disordered PTM/cleavage SLiMs | Short linear motifs in intrinsically disordered/low-complexity regions, often S/T/P/G- and aromatic-rich, that include proteolytic-processing/PTM hotspots and flexible linkers adjacent to transmembrane segments; the feature favors flexible coils or short amphipathic helices and systematically avoids hydrophobic transmembrane cores | 8.45 | 2 |
| #14866 | Intrinsic disorder and low-complexity regions | Intrinsic disorder/low-complexity signal: the feature marks compositionally biased, non-globular regions enriched in polar/charged and small residues (S/T/E/D/R/K/G/P; often A/L), typically flexible N- or C‑terminal tails, linkers, and propeptides that host short linear motifs or processing sites, while avoiding structured domains. | 8.41 | 2 |
| #6356 | Ser/Gly/Pro-rich low-complexity IDRs | Often fires on intrinsically disordered, low-complexity regions (IDRs) enriched in Ser/Gly/Pro and basic (Lys/Arg) residues, frequently in secreted/viral proteins, but also activates on short peptides and small folded proteins with no clear sequence motif. Includes Ser-rich PTM-prone segments within unstructured regions. | 8.24 | 5 |
| #11159 | Glycine residue identity detector | Residue-identity detector for glycine (G), largely context-agnostic but with a tendency to fire on glycines in flexible/disordered, often N-terminal or membrane-proximal regions; also active on glycines within transmembrane helices; broadly taxon- and function-independent. | 8.19 | 4 |
| #16243 | Structured-region leucine detector | Generic detector of leucine side chains in structured regions of folded domains, with occasional weaker responses to other hydrophobics; largely indifferent to specific functional motifs or protein class. | 8.14 | 6 |
| #7903 | Arginine-rich polybasic patches | Basic polycationic patches enriched in arginine (often with lysine) within low-complexity/disordered regions, frequently N-terminal. These include classical monopartite NLSs, nucleic acid–binding basic tails, signal peptide n-regions and cytosolic juxtamembrane segments (positive-inside rule), and basic clusters in secreted precursors; typically depleted of acidic residues. | 8.12 | 3 |
| #14891 | Disordered termini and cleavage detector | Residue-level detector of intrinsically disordered, flexible termini and proteolytic processing junctions—especially N-termini (including the first residue of the mature chain after propeptide/leader cleavage)—with no strict amino‑acid specificity and a bias toward small/charged residues; common in small, secreted/precursor and viral proteins but also present in disordered regions of diverse proteins. | 7.79 | 8 |
| #5337 | Pro/RGG-rich low-complexity IDRs | Intrinsically disordered, low‑complexity basic segments at termini and long loops, enriched in Pro/Gly and/or Arg/Ser (including RG/RGG or S/T clusters), common across eukaryotic and viral proteins. | 7.68 | 33 |
| #7986 | Intrinsically disordered low-complexity regions | Compositionally biased, intrinsically disordered low‑complexity regions used as flexible interaction/nucleic‑acid–binding modules—enriched in Pro/Ser/Lys/Glu/Gln/Thr/Gly/Arg and encompassing poly‑Pro tracts, Lys/Arg‑rich basic stretches, and acidic Asp/Glu clusters—common in transcription/RNA‑processing factors, viral basic regulators, and the C‑terminal scaffolds of RNase E | 7.68 | 19 |
| #2319 | Threonine residue detector | Residue-identity detector for threonine (Thr): activates on individual Thr residues across diverse proteins, with a mild enrichment in low-complexity/disordered, repeat-rich, and secretory regions; still marks Thr within well-structured domains; occasional weak spillover to serine. | 7.64 | 7 |
| #4900 | Arginine-rich segments and residues | Arginine residue identity/basic-tract feature: primarily marks arginine (R) residues—and dense arginine-rich, low-complexity segments—independent of protein family, fold, or function; shows a strong bias for R over other residues (occasional weak lysine signal), appearing both in disordered/repetitive regions and on individual R within structured domains or near transmembrane boundaries. | 7.62 | 3 |
| #2692 | Compositionally biased disordered LCRs | Feature targets compositionally biased, intrinsically disordered low‑complexity regions with long contiguous runs strongly enriched in small/polar (Gly/Ser/Asn/Thr) or acidic (Asp/Glu) residues; occasional activation on highly basic protamine‑like LCRs. Mere disorder, generic tails, or coiled‑coils are insufficient without such compositional bias. | 7.58 | 5 |
| #13897 | Proline-rich IDR detector | Detector of proline residues, with strongest signal in proline-rich, intrinsically disordered, low-complexity segments (often at termini, propeptides, and surface-exposed regions). The feature also activates on more isolated prolines in compact protein contexts, including within transmembrane helices, though typically at lower intensity than in extended Pro-rich disorder. | 7.40 | 7 |
| #678 | Sparse terminal residue peaks | Sparse residue-level activations distributed across short proteins and protein segments, with frequent firing in disordered N-terminal/C-terminal tails as well as in some short transmembrane or compact domains | 7.39 | 10 |
| #14646 | Short N-terminal leader segment | Short N-terminal segments within the first ~10–40 amino acids, often Lys/Arg-enriched but sometimes purely hydrophobic. These commonly correspond to Sec-type signal peptides, signal anchors, or other short N-terminal segments preceding/overlapping the first hydrophobic helix, and also include disordered N-terminal tails in soluble proteins. | 7.32 | 39 |
| #3670 | Terminal disordered peptide detector | Short unstructured peptides and N-terminal/C-terminal segments of larger proteins; occasional firing near N-terminal leader/signal sequences and basic targeting motifs (e.g., NLS); broadly residue-tolerant with bias toward hydrophobic, Gly/Pro, and basic residues; appears in viral accessory proteins, microproteins, and secreted precursors but also in diverse bacterial/archaeal enzymes, with sparse activation overall. | 7.26 | 6 |
| #14534 | Unknown generic feature | Unknown generic feature | 26.19 | 41 |
| #1803 | Unknown generic feature | Unknown generic feature | 24.66 | 41 |
| #9005 | Unknown generic feature | Unknown generic feature | 22.28 | 41 |
| #14895 | Unknown generic feature | Unknown generic feature | 19.00 | 41 |
| #9214 | Unknown generic feature | Unknown generic feature | 13.60 | 41 |
| #9194 | Unknown generic feature | Unknown generic feature | 11.73 | 41 |
Top-K sparse-autoencoder activations on the ESM-C residual stream. Activation is a feature's peak value over its region; Prevalence is the number of residues it fires on; Δ activation is isoform − canonical. Only the top 30 features per category are saved; generic / "unknown" features are hidden until you expand a table. Read more about SAE features →
Clinical variants
Differential region — N-terminal extension (isoform-unique)
| Pos (iso) | AA change | Consequence | Source | Clin. sig. | AF (gnomAD) | Impact | AlphaMissense | ESM-C ΔLLR | Link |
|---|---|---|---|---|---|---|---|---|---|
| 0 | L→L | synonymous_variant | gnomAD | — | 1.20e-05 | — | N/A | — | chr17-48101351-C-G |
| 1 | G→G | synonymous_variant | gnomAD | — | 2.40e-06 | — | N/A | 0.00 | chr17-48101348-G-A |
| 1 | G→G | synonymous_variant | gnomAD | — | 3.60e-06 | — | N/A | 0.00 | chr17-48101348-G-T |
| 1 | — | frameshift_variant | gnomAD | — | 2.40e-06 | LoF | — | — | chr17-48101348-GC-G |
| 1 | G→R | missense_variant | gnomAD | — | 1.20e-06 | — | N/A | 0.22 | chr17-48101350-C-G |
| 2 | A→D | missense_variant | gnomAD | — | 3.60e-06 | — | N/A | -1.59 | chr17-48101346-G-T |
| 2 | A→S | missense_variant | gnomAD | — | 6.00e-06 | — | N/A | -0.66 | chr17-48101347-C-A |
| 2 | A→T | missense_variant | gnomAD | — | 7.08e-05 | — | N/A | -0.92 | chr17-48101347-C-T |
| 3 | T→N | missense_variant | gnomAD | — | 1.20e-06 | — | N/A | -1.44 | chr17-48101343-G-T |
| 4 | P→P | synonymous_variant | gnomAD | — | 2.40e-06 | — | N/A | 0.00 | chr17-48101339-G-T |
| 4 | P→T | missense_variant | gnomAD | — | 1.20e-06 | — | N/A | -1.11 | chr17-48101341-G-T |
| 5 | P→S | missense_variant | gnomAD | — | 1.20e-06 | — | N/A | -0.39 | chr17-48101338-G-A |
| 6 | G→S | missense_variant | gnomAD | — | 1.20e-06 | — | N/A | -0.14 | chr17-48101335-C-T |
| 8 | P→P | synonymous_variant | gnomAD | — | 2.40e-06 | — | N/A | 0.00 | chr17-48101327-C-G |
| 8 | P→P | synonymous_variant | gnomAD | — | 4.80e-06 | — | N/A | 0.00 | chr17-48101327-C-T |
| 8 | P→Q | missense_variant | gnomAD | — | 2.40e-06 | — | N/A | -1.23 | chr17-48101328-G-T |
| 9 | — | frameshift_variant | gnomAD | — | 3.60e-06 | LoF | — | — | chr17-48101324-C-CGTCGG |
| 9 | T→P | missense_variant | gnomAD | — | 4.80e-06 | — | N/A | 1.19 | chr17-48101326-T-G |
| 10 | R→R | synonymous_variant | gnomAD | — | 1.20e-06 | — | N/A | 0.00 | chr17-48101321-T-A |
| 12 | A→A | synonymous_variant | gnomAD | — | 1.20e-06 | — | N/A | 0.00 | chr17-48101315-G-A |
| 12 | — | frameshift_variant | gnomAD | — | 1.20e-06 | LoF | — | — | chr17-48101316-G-GCGCGT |
| 12 | A→S | missense_variant | gnomAD | — | 7.19e-06 | — | N/A | 0.03 | chr17-48101317-C-A |
| 12 | — | frameshift_variant | gnomAD | — | 4.80e-06 | LoF | — | — | chr17-48101317-C-CA |
| 13 | S→S | synonymous_variant | gnomAD | — | 7.89e-01 | — | N/A | 0.00 | chr17-48101312-G-A |
| 13 | S→R | missense_variant | gnomAD | — | 4.08e-05 | — | N/A | 0.45 | chr17-48101312-G-C |
| 13 | S→T | missense_variant | gnomAD | — | 1.14e-04 | — | N/A | -0.98 | chr17-48101313-C-G |
| 13 | S→N | missense_variant | gnomAD | — | 1.20e-06 | — | N/A | -2.05 | chr17-48101313-C-T |
| 14 | S→S | synonymous_variant | gnomAD | — | 1.20e-06 | — | N/A | 0.00 | chr17-48101309-G-A |
| 14 | S→N | missense_variant | gnomAD | — | 7.19e-06 | — | N/A | -2.97 | chr17-48101310-C-T |
| 14 | S→G | missense_variant | gnomAD | — | 1.20e-06 | — | N/A | 0.05 | chr17-48101311-T-C |
| 15 | A→A | synonymous_variant | gnomAD | — | 1.20e-06 | — | N/A | 0.00 | chr17-48101306-T-C |
| 15 | A→V | missense_variant | gnomAD | — | 8.39e-05 | — | N/A | -1.72 | chr17-48101307-G-A |
| 15 | A→T | missense_variant | gnomAD | — | 2.40e-06 | — | N/A | -1.67 | chr17-48101308-C-T |
| 16 | A→V | missense_variant | gnomAD | — | 1.20e-06 | — | N/A | -1.05 | chr17-48101304-G-A |
| 17 | P→P | synonymous_variant | gnomAD | — | 1.20e-06 | — | N/A | 0.00 | chr17-48101300-C-A |
| 17 | P→L | missense_variant | gnomAD | — | 2.04e-05 | — | N/A | -0.31 | chr17-48101301-G-A |
| 17 | P→R | missense_variant | gnomAD | — | 2.40e-06 | — | N/A | 0.12 | chr17-48101301-G-C |
| 18 | I→V | missense_variant | gnomAD | — | 1.20e-06 | — | N/A | 1.38 | chr17-48101299-T-C |
| 19 | P→L | missense_variant | gnomAD | — | 1.20e-06 | — | N/A | -0.57 | chr17-48101295-G-A |
| 20 | L→L | synonymous_variant | gnomAD | — | 9.59e-06 | — | N/A | 0.00 | chr17-48101291-G-A |
| 20 | L→F | missense_variant | gnomAD | — | 1.20e-06 | — | N/A | -1.69 | chr17-48101293-G-A |
| 21 | G→G | synonymous_variant | gnomAD | — | 2.40e-06 | — | N/A | 0.00 | chr17-48101288-C-A |
| 21 | G→E | missense_variant | gnomAD | — | 1.20e-06 | — | N/A | -1.44 | chr17-48101289-C-T |
| 21 | G→R | missense_variant | gnomAD | — | 1.20e-06 | — | N/A | 0.59 | chr17-48101290-C-G |
| 22 | L→L | synonymous_variant | gnomAD | — | 1.20e-06 | — | N/A | 0.00 | chr17-48101285-G-A |
| 22 | L→L | synonymous_variant | gnomAD | — | 1.20e-06 | — | N/A | 0.00 | chr17-48101285-G-C |
| 22 | L→F | missense_variant | gnomAD | — | 2.40e-06 | — | N/A | -1.19 | chr17-48101287-G-A |
| 23 | L→F | missense_variant | gnomAD | — | 7.19e-06 | — | N/A | -1.48 | chr17-48101282-C-G |
| 24 | G→G | synonymous_variant | gnomAD | — | 1.20e-06 | — | N/A | 0.00 | chr17-48101279-G-A |
| 25 | A→S | missense_variant | gnomAD | — | 1.20e-06 | — | N/A | -0.05 | chr17-48101278-C-A |
| 25 | A→T | missense_variant | gnomAD | — | 1.20e-06 | — | N/A | -1.42 | chr17-48101278-C-T |
| 27 | L→L | synonymous_variant | gnomAD | — | 1.20e-06 | — | N/A | 0.00 | chr17-48101270-C-T |
| 29 | S→S | synonymous_variant | gnomAD | — | 4.57e-04 | — | N/A | 0.00 | chr17-48077038-G-A |
| 29 | S→S | synonymous_variant | COSMIC | — | — | — | N/A | 0.00 | COSV99848140 |
| 30 | V→V | synonymous_variant | gnomAD | — | 6.91e-07 | — | N/A | 0.00 | chr17-48077035-G-C |
| 30 | V→I | missense_variant | gnomAD | — | 1.52e-05 | — | N/A | -3.78 | chr17-48077037-C-T |
| 30 | V→I | missense_variant | COSMIC | — | — | — | N/A | -3.78 | COSV99848059 |
| 31 | T→T | synonymous_variant | gnomAD | — | 2.76e-06 | — | N/A | 0.00 | chr17-48077032-G-A |
| 31 | T→A | missense_variant | gnomAD | — | 6.91e-07 | — | N/A | -3.44 | chr17-48077034-T-C |
| 31 | — | frameshift_variant | gnomAD | — | 6.91e-07 | LoF | — | — | chr17-48077034-T-TCA |
| 32 | L→V | missense_variant | gnomAD | — | 6.91e-07 | — | N/A | -5.82 | chr17-48077031-G-C |
| 33 | — | inframe_insertion | gnomAD | — | 2.06e-06 | — | N/A | — | chr17-48077027-T-TAAA |
| 34 | T→S | missense_variant | gnomAD | — | 6.18e-06 | — | N/A | -0.96 | chr17-48077024-G-C |
| 34 | T→N | missense_variant | gnomAD | — | 6.87e-07 | — | N/A | -3.93 | chr17-48077024-G-T |
| 37 | L→L | synonymous_variant | gnomAD | — | 6.85e-07 | — | N/A | 0.00 | chr17-48077016-G-A |
| 37 | L→V | missense_variant | gnomAD | — | 2.19e-05 | — | N/A | -6.18 | chr17-48077016-G-C |
| 38 | A→A | synonymous_variant | gnomAD | — | 3.70e-05 | — | N/A | 0.00 | chr17-48077011-C-T |
| 38 | A→V | missense_variant | gnomAD | — | 6.85e-07 | — | N/A | -5.89 | chr17-48077012-G-A |
| 38 | A→V | missense_variant | COSMIC | — | — | — | N/A | -5.89 | COSV56682161 |
| 39 | G→G | synonymous_variant | gnomAD | — | 6.85e-07 | — | N/A | 0.00 | chr17-48077008-G-A |
| 39 | G→G | synonymous_variant | COSMIC | — | — | — | N/A | 0.00 | COSV56681939 |
| 40 | T→T | synonymous_variant | gnomAD | — | 6.85e-07 | — | N/A | 0.00 | chr17-48077005-A-G |
72 variants in the differential region.
Shared canonical core — sequence common to canonical and isoform; AlphaMissense applies here
| Pos (iso) | AA change | Consequence | Source | Clin. sig. | AF (gnomAD) | Impact | AlphaMissense | ESM-C ΔLLR | Link |
|---|---|---|---|---|---|---|---|---|---|
| 41 | M→I | missense_variant | gnomAD | — | 1.37e-06 | damaging | — | -10.94 | chr17-48077002-C-G |
| 41 | M→I | missense_variant | gnomAD | — | 6.85e-07 | damaging | — | -10.94 | chr17-48077002-C-T |
| 41 | M→T | missense_variant | gnomAD | — | 6.85e-07 | damaging | — | -11.50 | chr17-48077003-A-G |
| 41 | M→L | missense_variant | gnomAD | — | 6.85e-07 | damaging | — | -11.19 | chr17-48077004-T-G |
| 41 | M→V | missense_variant | COSMIC | — | — | damaging | — | -11.44 | COSV56682591 |
| 42 | G→G | synonymous_variant | gnomAD | — | 6.85e-07 | — | — | 0.00 | chr17-48076999-C-A |
| 42 | G→V | missense_variant | gnomAD | — | 6.85e-07 | damaging | likely_pathogenic (0.88) | -9.25 | chr17-48077000-C-A |
| 42 | G→W | missense_variant | gnomAD | — | 6.85e-07 | damaging | likely_pathogenic (0.97) | -12.31 | chr17-48077001-C-A |
| 43 | — | frameshift_variant | COSMIC | — | — | LoF | — | — | COSV56681712 |
| 43 | — | frameshift_variant | COSMIC | — | — | LoF | — | — | COSV56681499 |
| 44 | K→K | synonymous_variant | gnomAD | — | 2.74e-06 | — | — | 0.00 | chr17-48076993-T-C |
| 45 | Q→Q | synonymous_variant | gnomAD | — | 6.84e-07 | — | — | 0.00 | chr17-48076990-T-C |
| 45 | Q→K | missense_variant | COSMIC | — | — | damaging | ambiguous (0.37) | -10.00 | COSV56682149 |
| 46 | N→K | missense_variant | gnomAD | — | 1.85e-05 | damaging | likely_pathogenic (0.68) | -8.87 | chr17-48076987-G-C |
| 46 | N→K | missense_variant | ClinVar | Uncertain significance | — | damaging | likely_pathogenic (0.68) | -8.87 | ClinVar:3827926 |
| 47 | K→N | missense_variant | gnomAD | — | 6.84e-07 | damaging | likely_pathogenic (0.73) | -9.94 | chr17-48076984-C-G |
| 47 | — | inframe_deletion | gnomAD | — | 6.84e-07 | — | — | — | chr17-48076984-CTTG-C |
| 47 | K→Q | missense_variant | gnomAD | — | 1.37e-06 | damaging | likely_benign (0.24) | -10.25 | chr17-48076986-T-G |
| 48 | K→N | missense_variant | gnomAD | — | 6.84e-07 | damaging | likely_pathogenic (0.84) | -9.87 | chr17-48076981-C-A |
| 48 | K→R | missense_variant | gnomAD | — | 6.84e-07 | damaging | likely_benign (0.12) | -9.06 | chr17-48076982-T-C |
| 49 | K→N | missense_variant | gnomAD | — | 6.84e-07 | damaging | likely_pathogenic (0.58) | -9.00 | chr17-48076978-T-A |
| 49 | — | inframe_deletion | gnomAD | — | 1.37e-06 | — | — | — | chr17-48076978-TTTC-T |
| 49 | K→I | missense_variant | gnomAD | — | 6.84e-07 | damaging | ambiguous (0.56) | -11.69 | chr17-48076979-T-A |
| 49 | K→R | missense_variant | gnomAD | — | 9.58e-06 | damaging | likely_benign (0.12) | -8.94 | chr17-48076979-T-C |
| 49 | K→* | stop_gained | gnomAD | — | 6.84e-07 | LoF | — | — | chr17-48076980-T-A |
| 49 | K→E | missense_variant | gnomAD | — | 6.84e-07 | damaging | ambiguous (0.42) | -9.69 | chr17-48076980-T-C |
| 50 | V→V | synonymous_variant | gnomAD | — | 1.37e-06 | — | — | 0.00 | chr17-48076975-C-T |
| 50 | V→G | missense_variant | gnomAD | — | 6.85e-07 | damaging | likely_benign (0.14) | -8.75 | chr17-48076976-A-C |
| 50 | V→M | missense_variant | gnomAD | — | 7.53e-06 | damaging | likely_benign (0.13) | -8.37 | chr17-48076977-C-T |
| 50 | V→L | missense_variant | COSMIC | — | — | damaging | likely_benign (0.20) | -8.06 | COSV56682280 |
| 53 | V→V | synonymous_variant | gnomAD | — | 1.37e-06 | — | — | 0.00 | chr17-48076966-C-T |
| 53 | V→L | missense_variant | gnomAD | — | 2.05e-06 | damaging | likely_benign (0.30) | -7.78 | chr17-48076968-C-A |
| 53 | V→M | missense_variant | gnomAD | — | 2.74e-06 | damaging | likely_benign (0.24) | -8.68 | chr17-48076968-C-T |
| 54 | L→L | synonymous_variant | gnomAD | — | 2.05e-06 | — | — | 0.00 | chr17-48076963-T-C |
| 55 | E→K | missense_variant | gnomAD | — | 6.84e-07 | damaging | likely_pathogenic (0.83) | -11.12 | chr17-48076962-C-T |
| 56 | E→Q | missense_variant | gnomAD | — | 6.84e-07 | damaging | likely_pathogenic (0.62) | -9.69 | chr17-48076959-C-G |
| 56 | E→K | missense_variant | COSMIC | — | — | damaging | likely_pathogenic (0.83) | -9.56 | COSV56681509 |
| 57 | E→Q | missense_variant | gnomAD | — | 1.37e-06 | damaging | likely_pathogenic (0.61) | -9.81 | chr17-48076956-C-G |
| 58 | E→E | synonymous_variant | gnomAD | — | 1.37e-06 | — | — | 0.00 | chr17-48076951-T-C |
| 58 | E→G | missense_variant | gnomAD | — | 6.84e-07 | damaging | likely_pathogenic (0.77) | -9.56 | chr17-48076952-T-C |
| 59 | E→Q | missense_variant | COSMIC | — | — | damaging | likely_pathogenic (0.95) | -11.25 | COSV99848313 |
| 60 | E→E | synonymous_variant | gnomAD | — | 2.05e-06 | — | — | 0.00 | chr17-48076945-T-C |
| 62 | V→V | synonymous_variant | gnomAD | — | 2.74e-06 | — | — | 0.00 | chr17-48076939-C-T |
| 62 | V→V | synonymous_variant | COSMIC | — | — | — | — | 0.00 | COSV99848141 |
| 63 | V→V | synonymous_variant | gnomAD | — | 6.84e-07 | — | — | 0.00 | chr17-48076936-C-A |
| 63 | V→V | synonymous_variant | gnomAD | — | 6.84e-07 | — | — | 0.00 | chr17-48076936-C-G |
| 64 | E→V | missense_variant | ClinVar | — | — | damaging | likely_pathogenic (1.00) | -11.31 | ClinVar:4423332 |
| 65 | K→N | missense_variant | gnomAD | — | 6.84e-07 | damaging | likely_pathogenic (0.99) | -9.37 | chr17-48076930-T-A |
| 67 | L→L | synonymous_variant | gnomAD | — | 6.84e-06 | — | — | 0.00 | chr17-48076924-G-A |
| 67 | L→I | missense_variant | COSMIC | — | — | damaging | likely_benign (0.28) | -9.56 | COSV56682110 |
| 68 | D→H | missense_variant | gnomAD | — | 1.37e-06 | damaging | likely_pathogenic (0.94) | -7.89 | chr17-48076923-C-G |
| 68 | D→N | missense_variant | gnomAD | — | 1.85e-05 | damaging | likely_pathogenic (0.56) | -4.01 | chr17-48076923-C-T |
| 69 | R→H | missense_variant | gnomAD | — | 2.74e-06 | damaging | likely_pathogenic (0.79) | -7.03 | chr17-48076919-C-T |
| 69 | R→C | missense_variant | gnomAD | — | 1.23e-05 | damaging | likely_pathogenic (0.87) | -7.12 | chr17-48076920-G-A |
| 70 | R→Q | missense_variant | gnomAD | — | 2.05e-06 | damaging | likely_pathogenic (0.99) | -9.25 | chr17-48076916-C-T |
| 70 | R→R | synonymous_variant | gnomAD | — | 6.84e-07 | — | — | 0.00 | chr17-48076917-G-T |
| 70 | R→R | synonymous_variant | COSMIC | — | — | — | — | 0.00 | COSV56682056 |
| 70 | R→* | stop_gained | COSMIC | — | — | LoF | — | — | COSV56681749 |
| 71 | V→V | synonymous_variant | gnomAD | — | 2.05e-06 | — | — | 0.00 | chr17-48076912-C-T |
| 71 | V→A | missense_variant | gnomAD | — | 1.37e-06 | damaging | likely_pathogenic (0.92) | -7.43 | chr17-48076913-A-G |
| 72 | V→V | synonymous_variant | gnomAD | — | 1.37e-06 | — | — | 0.00 | chr17-48076909-T-C |
| 72 | V→L | missense_variant | gnomAD | — | 6.84e-07 | damaging | likely_pathogenic (0.76) | -6.96 | chr17-48076911-C-A |
| 73 | K→K | synonymous_variant | gnomAD | — | 6.84e-07 | — | — | 0.00 | chr17-48076906-C-T |
| 74 | G→G | synonymous_variant | gnomAD | — | 3.42e-06 | — | — | 0.00 | chr17-48076903-G-A |
| 74 | G→S | missense_variant | gnomAD | — | 6.84e-07 | damaging | likely_pathogenic (0.96) | -7.97 | chr17-48076905-C-T |
| 74 | G→S | missense_variant | COSMIC | — | — | damaging | likely_pathogenic (0.96) | -7.97 | COSV99847936 |
| 75 | K→R | missense_variant | gnomAD | — | 2.05e-06 | — | likely_benign (0.09) | -5.18 | chr17-48076901-T-C |
| 77 | E→D | missense_variant | gnomAD | — | 6.85e-07 | damaging | likely_pathogenic (1.00) | -7.72 | chr17-48076894-C-G |
| 77 | E→K | missense_variant | COSMIC | — | — | damaging | likely_pathogenic (1.00) | -10.19 | COSV56682856 |
| 78 | Y→Y | synonymous_variant | gnomAD | — | 1.71e-05 | — | — | 0.00 | chr17-48076891-G-A |
| 79 | L→L | synonymous_variant | gnomAD | — | 3.43e-06 | — | — | 0.00 | chr17-48076888-G-A |
| 79 | L→R | missense_variant | gnomAD | — | 6.85e-07 | damaging | likely_pathogenic (0.84) | -9.56 | chr17-48076889-A-C |
| 79 | L→F | missense_variant | COSMIC | — | — | — | ambiguous (0.35) | -6.53 | COSV56681814 |
| 80 | L→L | synonymous_variant | gnomAD | — | 6.85e-07 | — | — | 0.00 | chr17-48076885-T-G |
| 81 | K→K | synonymous_variant | gnomAD | — | 6.85e-07 | — | — | 0.00 | chr17-48076882-C-T |
| 81 | K→E | missense_variant | gnomAD | — | 6.85e-07 | damaging | likely_pathogenic (1.00) | -11.19 | chr17-48076884-T-C |
| 84 | G→G | synonymous_variant | gnomAD | — | 4.39e-05 | — | — | 0.00 | chr17-48076873-T-C |
| 85 | F→F | synonymous_variant | gnomAD | — | 6.86e-07 | — | — | 0.00 | chr17-48076870-G-A |
| 85 | F→L | missense_variant | COSMIC | — | — | damaging | likely_pathogenic (1.00) | -10.37 | COSV56682209 |
| 88 | E→K | missense_variant | gnomAD | — | 7.10e-07 | damaging | likely_pathogenic (0.84) | -9.50 | chr17-48076177-C-T |
| 88 | E→Q | missense_variant | ClinVar | — | — | damaging | likely_pathogenic (0.60) | -11.12 | ClinVar:4423331 |
| 88 | E→Q | missense_variant | COSMIC | — | — | damaging | likely_pathogenic (0.60) | -11.12 | COSV56681530 |
| 89 | D→H | missense_variant | gnomAD | — | 7.06e-07 | damaging | likely_pathogenic (0.97) | -12.12 | chr17-48076174-C-G |
| 90 | N→D | missense_variant | gnomAD | — | 1.41e-06 | damaging | likely_pathogenic (0.97) | -10.37 | chr17-48076171-T-C |
| 90 | N→S | missense_variant | COSMIC | — | — | damaging | likely_pathogenic (0.59) | -9.00 | COSV56681968 |
| 93 | E→D | missense_variant | COSMIC | — | — | damaging | likely_pathogenic (1.00) | -10.56 | COSV99848274 |
| 95 | E→D | missense_variant | COSMIC | — | — | damaging | likely_pathogenic (0.96) | -8.69 | COSV99848151 |
| 97 | N→N | synonymous_variant | gnomAD | — | 1.38e-06 | — | — | 0.00 | chr17-48076148-G-A |
| 98 | L→L | synonymous_variant | gnomAD | — | 1.38e-06 | — | — | 0.00 | chr17-48076147-G-A |
| 101 | P→P | synonymous_variant | gnomAD | — | 1.51e-05 | — | — | 0.00 | chr17-48076136-G-A |
| 101 | P→P | synonymous_variant | gnomAD | — | 6.88e-07 | — | — | 0.00 | chr17-48076136-G-C |
| 101 | P→P | synonymous_variant | COSMIC | — | — | — | — | 0.00 | COSV99847948 |
| 102 | D→N | missense_variant | COSMIC | — | — | damaging | likely_pathogenic (0.96) | -8.31 | COSV99847954 |
| 103 | L→L | synonymous_variant | gnomAD | — | 1.23e-05 | — | — | 0.00 | chr17-48076130-G-A |
| 105 | A→A | synonymous_variant | gnomAD | — | 5.48e-06 | — | — | 0.00 | chr17-48076124-A-G |
| 105 | A→A | synonymous_variant | gnomAD | — | 6.85e-07 | — | — | 0.00 | chr17-48076124-A-T |
| 105 | A→G | missense_variant | gnomAD | — | 1.37e-06 | damaging | ambiguous (0.38) | -9.94 | chr17-48076125-G-C |
| 108 | L→L | synonymous_variant | gnomAD | — | 1.37e-06 | — | — | 0.00 | chr17-48076115-C-T |
| 108 | L→L | synonymous_variant | gnomAD | — | 6.85e-07 | — | — | 0.00 | chr17-48076117-G-A |
| 108 | L→L | synonymous_variant | COSMIC | — | — | — | — | 0.00 | COSV105045650 |
| 109 | Q→* | stop_gained | COSMIC | — | — | LoF | — | — | COSV99848158 |
| 110 | S→L | missense_variant | gnomAD | — | 2.74e-06 | damaging | likely_benign (0.27) | -9.37 | chr17-48076110-G-A |
| 110 | S→L | missense_variant | ClinVar | Uncertain significance | — | damaging | likely_benign (0.27) | -9.37 | ClinVar:4648788 |
| 111 | Q→R | missense_variant | gnomAD | — | 6.85e-07 | — | likely_benign (0.33) | -7.47 | chr17-48076107-T-C |
| 111 | Q→* | stop_gained | COSMIC | — | — | LoF | — | — | COSV99848270 |
| 112 | K→R | missense_variant | gnomAD | — | 6.85e-07 | — | likely_benign (0.12) | -6.44 | chr17-48076104-T-C |
| 112 | K→Q | missense_variant | ClinVar | Uncertain significance | — | damaging | likely_benign (0.33) | -8.75 | ClinVar:4531663 |
| 113 | T→A | missense_variant | gnomAD | — | 6.85e-07 | — | likely_benign (0.06) | -6.06 | chr17-48076102-T-C |
| 115 | H→P | missense_variant | gnomAD | — | 4.79e-06 | — | likely_benign (0.11) | -7.21 | chr17-48076095-T-G |
| 117 | T→R | missense_variant | gnomAD | — | 2.74e-06 | damaging | likely_benign (0.23) | -7.90 | chr17-48076089-G-C |
| 118 | D→G | missense_variant | gnomAD | — | 6.84e-07 | — | likely_benign (0.10) | -6.27 | chr17-48076086-T-C |
| 118 | D→N | missense_variant | gnomAD | — | 6.84e-07 | — | likely_benign (0.10) | -6.15 | chr17-48076087-C-T |
| 118 | D→G | missense_variant | ClinVar | Uncertain significance | — | — | likely_benign (0.10) | -6.27 | ClinVar:4220041 |
| 118 | D→N | missense_variant | ClinVar | Uncertain significance | — | — | likely_benign (0.10) | -6.15 | ClinVar:4220044 |
| 119 | K→K | synonymous_variant | gnomAD | — | 4.04e-05 | — | — | 0.00 | chr17-48076082-T-C |
| 119 | K→I | missense_variant | gnomAD | — | 6.84e-07 | damaging | ambiguous (0.43) | -10.75 | chr17-48076083-T-A |
| 119 | K→I | missense_variant | ClinVar | Uncertain significance | — | damaging | ambiguous (0.43) | -10.75 | ClinVar:4220042 |
| 120 | S→S | synonymous_variant | gnomAD | — | 4.79e-06 | — | — | 0.00 | chr17-48076079-T-G |
| 121 | E→K | missense_variant | COSMIC | — | — | damaging | likely_benign (0.26) | -8.68 | COSV99848247 |
| 122 | G→R | missense_variant | gnomAD | — | 6.84e-07 | damaging | likely_pathogenic (0.58) | -8.76 | chr17-48076075-C-G |
| 122 | G→R | missense_variant | COSMIC | — | — | damaging | likely_pathogenic (0.58) | -8.76 | COSV56682528 |
| 125 | R→H | missense_variant | gnomAD | — | 7.53e-06 | damaging | likely_pathogenic (0.89) | -9.31 | chr17-48076065-C-T |
| 125 | R→C | missense_variant | gnomAD | — | 6.16e-06 | damaging | likely_pathogenic (0.96) | -9.31 | chr17-48076066-G-A |
| 125 | R→H | missense_variant | ClinVar | Uncertain significance | — | damaging | likely_pathogenic (0.89) | -9.31 | ClinVar:2290145 |
| 125 | R→H | missense_variant | COSMIC | — | — | damaging | likely_pathogenic (0.89) | -9.31 | COSV56682901 |
| 125 | R→C | missense_variant | COSMIC | — | — | damaging | likely_pathogenic (0.96) | -9.31 | COSV56682920 |
| 126 | K→R | missense_variant | gnomAD | — | 8.90e-06 | damaging | likely_benign (0.11) | -7.87 | chr17-48076062-T-C |
| 127 | A→T | missense_variant | gnomAD | — | 1.37e-06 | — | likely_benign (0.08) | -6.34 | chr17-48076060-C-T |
| 127 | A→T | missense_variant | ClinVar | Uncertain significance | — | — | likely_benign (0.08) | -6.34 | ClinVar:4220043 |
| 128 | D→G | missense_variant | gnomAD | — | 6.85e-07 | damaging | likely_benign (0.15) | -8.36 | chr17-48076056-T-C |
| 129 | S→A | missense_variant | gnomAD | — | 6.85e-07 | damaging | likely_benign (0.06) | -9.87 | chr17-48076054-A-C |
| 129 | S→T | missense_variant | gnomAD | — | 1.37e-06 | — | likely_benign (0.05) | -7.31 | chr17-48076054-A-T |
| 133 | D→V | missense_variant | gnomAD | — | 1.37e-06 | damaging | likely_benign (0.13) | -9.56 | chr17-48076041-T-A |
| 133 | D→V | missense_variant | ClinVar | Uncertain significance | — | damaging | likely_benign (0.13) | -9.56 | ClinVar:2307098 |
| 133 | D→N | missense_variant | COSMIC | — | — | damaging | likely_benign (0.11) | -7.93 | COSV56682563 |
| 134 | K→K | synonymous_variant | gnomAD | — | 1.37e-06 | — | — | 0.00 | chr17-48076037-C-T |
| 134 | K→R | missense_variant | gnomAD | — | 2.05e-06 | — | likely_benign (0.09) | -4.90 | chr17-48076038-T-C |
| 134 | K→E | missense_variant | gnomAD | — | 6.85e-07 | damaging | likely_benign (0.18) | -11.05 | chr17-48076039-T-C |
| 134 | K→K | synonymous_variant | COSMIC | — | — | — | — | 0.00 | COSV56682520 |
| 134 | K→N | missense_variant | COSMIC | — | — | damaging | likely_benign (0.29) | -9.30 | COSV99848170 |
| 135 | G→G | synonymous_variant | gnomAD | — | 2.06e-06 | — | — | 0.00 | chr17-48076034-T-C |
| 135 | — | frameshift_variant | gnomAD | — | 1.37e-06 | LoF | — | — | chr17-48076035-CCCTT-C |
| 135 | G→R | missense_variant | gnomAD | — | 2.74e-06 | damaging | ambiguous (0.54) | -7.55 | chr17-48076036-C-T |
| 136 | E→E | synonymous_variant | gnomAD | — | 6.86e-07 | — | — | 0.00 | chr17-48076031-C-T |
| 136 | E→Q | missense_variant | COSMIC | — | — | damaging | likely_benign (0.23) | -8.62 | COSV99848203 |
| 137 | E→D | missense_variant | gnomAD | — | 6.86e-07 | — | likely_benign (0.05) | -5.84 | chr17-48076028-C-G |
| 138 | S→S | synonymous_variant | gnomAD | — | 4.81e-06 | — | — | 0.00 | chr17-48076025-G-A |
| 138 | S→G | missense_variant | gnomAD | — | 6.86e-07 | — | likely_benign (0.06) | -5.94 | chr17-48076027-T-C |
| 138 | S→S | synonymous_variant | COSMIC | — | — | — | — | 0.00 | COSV99848275 |
| 140 | P→P | synonymous_variant | gnomAD | — | 2.07e-06 | — | — | 0.00 | chr17-48076019-T-C |
| 140 | P→L | missense_variant | gnomAD | — | 1.38e-06 | — | likely_benign (0.09) | -6.30 | chr17-48076020-G-A |
| 142 | K→N | missense_variant | gnomAD | — | 6.90e-07 | damaging | likely_pathogenic (0.93) | -10.37 | chr17-48076013-C-A |
| 142 | K→K | synonymous_variant | gnomAD | — | 1.38e-06 | — | — | 0.00 | chr17-48076013-C-T |
| 143 | K→N | missense_variant | gnomAD | — | 6.91e-07 | damaging | likely_pathogenic (0.91) | -10.25 | chr17-48076010-C-G |
| 143 | K→R | missense_variant | gnomAD | — | 1.38e-06 | — | likely_benign (0.09) | -5.37 | chr17-48076011-T-C |
| 143 | K→N | missense_variant | COSMIC | — | — | damaging | likely_pathogenic (0.91) | -10.25 | COSV56681336 |
| 143 | K→N | missense_variant | COSMIC | — | — | damaging | likely_pathogenic (0.91) | -10.25 | COSV99848297 |
| 144 | — | inframe_deletion | gnomAD | — | 2.77e-06 | — | — | — | chr17-48076007-TTTC-T |
| 144 | K→R | missense_variant | gnomAD | — | 6.92e-07 | — | likely_benign (0.09) | -6.75 | chr17-48076008-T-C |
| 145 | E→G | missense_variant | gnomAD | — | 6.92e-07 | damaging | likely_benign (0.33) | -8.81 | chr17-48076005-T-C |
| 146 | E→V | missense_variant | gnomAD | — | 1.39e-06 | damaging | likely_benign (0.30) | -8.24 | chr17-48076002-T-A |
| 146 | E→Q | missense_variant | gnomAD | — | 6.94e-07 | damaging | ambiguous (0.43) | -8.99 | chr17-48076003-C-G |
| 147 | S→S | synonymous_variant | gnomAD | — | 6.85e-07 | — | — | 0.00 | chr17-48075098-T-A |
| 147 | S→L | missense_variant | gnomAD | — | 6.85e-07 | — | likely_benign (0.08) | -6.09 | chr17-48075099-G-A |
| 148 | E→A | missense_variant | gnomAD | — | 6.85e-07 | damaging | ambiguous (0.50) | -9.43 | chr17-48075096-T-G |
| 150 | P→S | missense_variant | gnomAD | — | 6.84e-07 | — | ambiguous (0.53) | -5.76 | chr17-48075091-G-A |
| 150 | P→Q | missense_variant | COSMIC | — | — | damaging | likely_pathogenic (0.61) | -7.20 | COSV106082017 |
| 150 | P→S | missense_variant | COSMIC | — | — | — | ambiguous (0.53) | -5.76 | COSV56682997 |
| 151 | R→Q | missense_variant | gnomAD | — | 6.84e-07 | damaging | likely_pathogenic (0.92) | -8.69 | chr17-48075087-C-T |
| 151 | R→* | stop_gained | gnomAD | — | 2.05e-06 | LoF | — | — | chr17-48075088-G-A |
| 151 | R→* | stop_gained | COSMIC | — | — | LoF | — | — | COSV56682269 |
| 153 | F→S | missense_variant | gnomAD | — | 6.84e-07 | damaging | likely_pathogenic (1.00) | -10.44 | chr17-48075081-A-G |
| 154 | A→A | synonymous_variant | gnomAD | — | 1.37e-06 | — | — | 0.00 | chr17-48075077-A-C |
| 154 | A→A | synonymous_variant | gnomAD | — | 6.84e-07 | — | — | 0.00 | chr17-48075077-A-G |
| 155 | R→Q | missense_variant | gnomAD | — | 6.84e-07 | damaging | likely_pathogenic (0.99) | -8.12 | chr17-48075075-C-T |
| 155 | R→* | stop_gained | gnomAD | — | 6.84e-07 | LoF | — | — | chr17-48075076-G-A |
| 155 | R→R | synonymous_variant | COSMIC | — | — | — | — | 0.00 | COSV107309107 |
| 155 | R→L | missense_variant | COSMIC | — | — | damaging | likely_pathogenic (1.00) | -10.37 | COSV99847942 |
| 155 | R→Q | missense_variant | COSMIC | — | — | damaging | likely_pathogenic (0.99) | -8.12 | COSV56683114 |
| 155 | R→* | stop_gained | COSMIC | — | — | LoF | — | — | COSV104387334 |
| 156 | G→D | missense_variant | COSMIC | — | — | damaging | likely_pathogenic (0.89) | -9.62 | COSV105045685 |
| 158 | E→K | missense_variant | gnomAD | — | 6.84e-07 | damaging | likely_pathogenic (0.87) | -7.54 | chr17-48075067-C-T |
| 158 | E→E | synonymous_variant | COSMIC | — | — | — | — | 0.00 | COSV56681442 |
| 159 | P→P | synonymous_variant | gnomAD | — | 1.97e-04 | — | — | 0.00 | chr17-48075062-C-T |
| 159 | P→L | missense_variant | gnomAD | — | 7.52e-06 | damaging | likely_pathogenic (1.00) | -10.81 | chr17-48075063-G-A |
| 159 | P→P | synonymous_variant | COSMIC | — | — | — | — | 0.00 | COSV99848187 |
| 159 | P→Q | missense_variant | COSMIC | — | — | damaging | likely_pathogenic (1.00) | -12.12 | COSV56682499 |
| 159 | P→S | missense_variant | COSMIC | — | — | damaging | likely_pathogenic (1.00) | -9.69 | COSV99848079 |
| 160 | E→D | missense_variant | gnomAD | — | 1.37e-06 | damaging | likely_pathogenic (0.92) | -9.06 | chr17-48075059-C-A |
| 160 | E→E | synonymous_variant | gnomAD | — | 1.37e-06 | — | — | 0.00 | chr17-48075059-C-T |
| 160 | E→D | missense_variant | ClinVar | Uncertain significance | — | damaging | likely_pathogenic (0.92) | -9.06 | ClinVar:4648787 |
| 160 | E→K | missense_variant | COSMIC | — | — | damaging | likely_pathogenic (1.00) | -12.12 | COSV56681862 |
| 161 | R→R | synonymous_variant | gnomAD | — | 6.84e-07 | — | — | 0.00 | chr17-48075056-C-T |
| 161 | R→Q | missense_variant | COSMIC | — | — | damaging | likely_pathogenic (0.98) | -7.87 | COSV56682101 |
| 161 | R→L | missense_variant | COSMIC | — | — | damaging | likely_pathogenic (0.99) | -10.12 | COSV99848174 |
| 162 | I→V | missense_variant | gnomAD | — | 6.84e-07 | damaging | likely_pathogenic (0.96) | -7.69 | chr17-48075055-T-C |
| 164 | G→A | missense_variant | COSMIC | — | — | damaging | likely_pathogenic (1.00) | -12.44 | COSV56682437 |
| 166 | T→T | synonymous_variant | gnomAD | — | 2.74e-06 | — | — | 0.00 | chr17-48075041-T-C |
| 168 | S→S | synonymous_variant | gnomAD | — | 6.84e-07 | — | — | 0.00 | chr17-48075035-G-A |
| 169 | S→S | synonymous_variant | COSMIC | — | — | — | — | 0.00 | COSV56681832 |
| 171 | E→Q | missense_variant | gnomAD | — | 6.84e-07 | damaging | likely_pathogenic (0.92) | -10.00 | chr17-48075028-C-G |
| 172 | L→L | synonymous_variant | gnomAD | — | 9.59e-06 | — | — | 0.00 | chr17-48075023-G-A |
| 173 | M→L | missense_variant | gnomAD | — | 6.85e-07 | damaging | likely_pathogenic (0.91) | -10.25 | chr17-48075022-T-A |
| 173 | M→V | missense_variant | gnomAD | — | 6.85e-07 | damaging | likely_pathogenic (0.99) | -11.50 | chr17-48075022-T-C |
| 173 | — | inframe_deletion | COSMIC | — | — | — | — | — | COSV56681204 |
| 174 | F→F | synonymous_variant | gnomAD | — | 6.85e-07 | — | — | 0.00 | chr17-48075017-G-A |
| 175 | L→P | missense_variant | gnomAD | — | 6.85e-07 | damaging | likely_pathogenic (1.00) | -13.06 | chr17-48075015-A-G |
| 175 | L→L | synonymous_variant | COSMIC | — | — | — | — | 0.00 | COSV99848148 |
| 175 | L→L | synonymous_variant | COSMIC | — | — | — | — | 0.00 | COSV56681918 |
| 176 | M→I | missense_variant | gnomAD | — | 6.86e-07 | damaging | likely_pathogenic (0.97) | -8.12 | chr17-48075011-C-A |
| 176 | M→T | missense_variant | gnomAD | — | 6.86e-07 | damaging | likely_pathogenic (1.00) | -11.81 | chr17-48075012-A-G |
| 177 | K→K | synonymous_variant | gnomAD | — | 6.86e-07 | — | — | 0.00 | chr17-48075008-T-C |
| 177 | K→T | missense_variant | gnomAD | — | 6.86e-07 | damaging | likely_pathogenic (1.00) | -11.87 | chr17-48075009-T-G |
| 179 | K→R | missense_variant | gnomAD | — | 8.28e-06 | — | likely_benign (0.14) | -6.94 | chr17-48071577-T-C |
| 180 | N→K | missense_variant | COSMIC | — | — | damaging | likely_pathogenic (0.89) | -8.94 | COSV99848229 |
| 181 | S→F | missense_variant | COSMIC | — | — | damaging | likely_pathogenic (0.99) | -10.37 | COSV56681364 |
| 184 | A→A | synonymous_variant | gnomAD | — | 6.87e-07 | — | — | 0.00 | chr17-48071561-A-G |
| 185 | D→D | synonymous_variant | gnomAD | — | 6.87e-07 | — | — | 0.00 | chr17-48071558-G-A |
| 187 | V→L | missense_variant | gnomAD | — | 6.86e-07 | damaging | likely_pathogenic (1.00) | -11.25 | chr17-48071554-C-G |
| 189 | A→A | synonymous_variant | gnomAD | — | 6.85e-07 | — | — | 0.00 | chr17-48071546-G-A |
| 189 | A→D | missense_variant | gnomAD | — | 6.85e-07 | damaging | likely_pathogenic (1.00) | -12.00 | chr17-48071547-G-T |
| 189 | A→V | missense_variant | COSMIC | — | — | damaging | likely_pathogenic (1.00) | -10.62 | COSV99848117 |
| 191 | E→* | stop_gained | gnomAD | — | 6.85e-07 | LoF | — | — | chr17-48071542-C-A |
| 191 | E→K | missense_variant | COSMIC | — | — | damaging | likely_pathogenic (0.99) | -10.94 | COSV56681808 |
| 193 | N→S | missense_variant | gnomAD | — | 6.84e-07 | damaging | likely_pathogenic (0.91) | -12.44 | chr17-48071535-T-C |
| 194 | V→F | missense_variant | gnomAD | — | 2.74e-06 | damaging | likely_pathogenic (0.57) | -8.73 | chr17-48071533-C-A |
| 194 | V→V | synonymous_variant | COSMIC | — | — | — | — | 0.00 | COSV56681933 |
| 195 | K→K | synonymous_variant | COSMIC | — | — | — | — | 0.00 | COSV56682664 |
| 196 | C→C | synonymous_variant | gnomAD | — | 1.37e-06 | — | — | 0.00 | chr17-48071525-G-A |
| 196 | C→S | missense_variant | gnomAD | — | 6.84e-07 | damaging | likely_pathogenic (1.00) | -5.97 | chr17-48071526-C-G |
| 197 | P→L | missense_variant | gnomAD | — | 6.84e-07 | damaging | likely_pathogenic (1.00) | -10.62 | chr17-48071523-G-A |
| 198 | Q→Q | synonymous_variant | gnomAD | — | 6.84e-07 | — | — | 0.00 | chr17-48071519-C-T |
| 198 | Q→* | stop_gained | gnomAD | — | 6.84e-07 | LoF | — | — | chr17-48071521-G-A |
| 199 | V→F | missense_variant | gnomAD | — | 2.05e-06 | damaging | likely_pathogenic (0.99) | -12.05 | chr17-48071518-C-A |
| 201 | I→I | synonymous_variant | COSMIC | — | — | — | — | 0.00 | COSV56681722 |
| 201 | I→T | missense_variant | COSMIC | — | — | damaging | likely_pathogenic (1.00) | -11.06 | COSV56682067 |
| 202 | S→S | synonymous_variant | gnomAD | — | 2.05e-06 | — | — | 0.00 | chr17-48071507-G-T |
| 203 | F→F | synonymous_variant | gnomAD | — | 6.84e-07 | — | — | 0.00 | chr17-48071504-G-A |
| 204 | Y→Y | synonymous_variant | gnomAD | — | 1.37e-06 | — | — | 0.00 | chr17-48071501-A-G |
| 204 | Y→C | missense_variant | gnomAD | — | 6.84e-07 | damaging | likely_pathogenic (1.00) | -11.94 | chr17-48071502-T-C |
| 207 | R→K | missense_variant | COSMIC | — | — | damaging | likely_pathogenic (0.99) | -11.06 | COSV56682552 |
| 209 | T→T | synonymous_variant | gnomAD | — | 4.11e-06 | — | — | 0.00 | chr17-48071486-C-T |
| 209 | T→K | missense_variant | COSMIC | — | — | damaging | likely_pathogenic (0.99) | -13.94 | COSV56681836 |
| 211 | H→H | synonymous_variant | gnomAD | — | 6.84e-07 | — | — | 0.00 | chr17-48071480-A-G |
| 212 | S→S | synonymous_variant | gnomAD | — | 6.84e-07 | — | — | 0.00 | chr17-48071477-G-A |
| 212 | S→F | missense_variant | gnomAD | — | 6.84e-07 | damaging | likely_pathogenic (0.97) | -11.12 | chr17-48071478-G-A |
| 212 | S→A | missense_variant | gnomAD | — | 6.84e-07 | damaging | likely_benign (0.27) | -9.06 | chr17-48071479-A-C |
| 212 | S→P | missense_variant | gnomAD | — | 6.84e-07 | damaging | likely_pathogenic (0.98) | -11.81 | chr17-48071479-A-G |
| 212 | S→T | missense_variant | gnomAD | — | 6.84e-07 | damaging | ambiguous (0.44) | -10.19 | chr17-48071479-A-T |
| 213 | Y→Y | synonymous_variant | gnomAD | — | 2.94e-05 | — | — | 0.00 | chr17-48071474-G-A |
| 213 | Y→* | stop_gained | gnomAD | — | 6.85e-07 | LoF | — | — | chr17-48071474-G-T |
| 214 | P→S | missense_variant | gnomAD | — | 6.85e-07 | damaging | likely_pathogenic (0.58) | -8.75 | chr17-48071473-G-A |
| 214 | P→S | missense_variant | COSMIC | — | — | damaging | likely_pathogenic (0.58) | -8.75 | COSV108798400 |
| 215 | S→S | synonymous_variant | gnomAD | — | 4.11e-06 | — | — | 0.00 | chr17-48071468-C-T |
| 215 | S→L | missense_variant | gnomAD | — | 2.05e-06 | damaging | likely_benign (0.21) | -8.42 | chr17-48071469-G-A |
| 215 | S→L | missense_variant | ClinVar | Uncertain significance | — | damaging | likely_benign (0.21) | -8.42 | ClinVar:3138046 |
| 215 | S→S | synonymous_variant | COSMIC | — | — | — | — | 0.00 | COSV56681272 |
| 215 | S→* | stop_gained | COSMIC | — | — | LoF | — | — | COSV99848014 |
| 216 | E→K | missense_variant | COSMIC | — | — | damaging | likely_pathogenic (0.75) | -12.00 | COSV104555665 |
| 217 | D→E | missense_variant | gnomAD | — | 4.80e-06 | — | likely_benign (0.09) | -5.24 | chr17-48071462-A-C |
| 217 | D→D | synonymous_variant | gnomAD | — | 6.85e-07 | — | — | 0.00 | chr17-48071462-A-G |
| 217 | D→Y | missense_variant | gnomAD | — | 6.85e-07 | damaging | likely_pathogenic (0.71) | -11.56 | chr17-48071464-C-A |
| 217 | D→N | missense_variant | gnomAD | — | 1.37e-06 | damaging | likely_benign (0.31) | -10.18 | chr17-48071464-C-T |
| 217 | D→E | missense_variant | ClinVar | Uncertain significance | — | — | likely_benign (0.09) | -5.24 | ClinVar:2517719 |
| 219 | — | inframe_deletion | gnomAD | — | 2.06e-06 | — | — | — | chr17-48071456-GTCA-G |
| 219 | D→N | missense_variant | COSMIC | — | — | damaging | likely_benign (0.19) | -8.62 | COSV99848218 |
| 220 | K→N | missense_variant | COSMIC | — | — | damaging | likely_pathogenic (0.72) | -9.62 | COSV99848111 |
| 220 | — | frameshift_variant | COSMIC | — | — | LoF | — | — | COSV56682928 |
| 221 | K→* | stop_gained | gnomAD | — | 6.85e-07 | LoF | — | — | chr17-48071452-T-A |
| 222 | D→V | missense_variant | gnomAD | — | 6.86e-07 | damaging | ambiguous (0.45) | -9.47 | chr17-48071448-T-A |
| 222 | — | frameshift_variant | gnomAD | — | 6.86e-07 | LoF | — | — | chr17-48071449-C-CT |
| 222 | — | frameshift_variant | gnomAD | — | 6.86e-07 | LoF | — | — | chr17-48071449-CTTTT-C |
| 222 | D→H | missense_variant | COSMIC | — | — | damaging | likely_pathogenic (0.73) | -10.72 | COSV56682228 |
| 222 | D→Y | missense_variant | COSMIC | — | — | damaging | likely_pathogenic (0.65) | -9.97 | COSV56681987 |
| 223 | — | frameshift_variant | gnomAD | — | 6.86e-07 | LoF | — | — | chr17-48071445-TC-T |
| 223 | D→Y | missense_variant | gnomAD | — | 1.37e-06 | damaging | likely_pathogenic (0.68) | -10.43 | chr17-48071446-C-A |
| 223 | D→N | missense_variant | gnomAD | — | 6.86e-07 | — | likely_benign (0.27) | -7.43 | chr17-48071446-C-T |
| 223 | — | frameshift_variant | gnomAD | — | 6.86e-07 | LoF | — | — | chr17-48071446-CA-C |
| 225 | N→K | missense_variant | gnomAD | — | 6.87e-07 | damaging | likely_pathogenic (0.68) | -7.21 | chr17-48071438-G-C |
| 225 | — | frameshift_variant | gnomAD | — | 1.37e-06 | LoF | — | — | chr17-48071439-TTCTTG-T |
| 225 | N→K | missense_variant | COSMIC | — | — | damaging | likely_pathogenic (0.68) | -7.21 | COSV56682095 |
281 variants in the shared canonical core.
LoF = frameshift / stop-gain / splice-disrupting — inherently loss-of-function, flagged by consequence (AlphaMissense and ESM-C score only missense/substitutions, so they are blank here by design, not by absence of impact). AlphaMissense is computed in the canonical reading frame, so it scores the shared core but reads N/A across an isoform-unique extension — use ESM-C ΔLLR there.