CBX1
EXTENDED 189 aa (canonical 185 aa) · UniProt P83916 · CDLMPS
chr17:48077016:-:CTG:ENST00000225603.9
AI summary A 4-residue N-terminal add-on leaves nuclear targeting, domains, and biophysics essentially untouched despite a locally well-folded tip.
The only notable tier-2 signal is that the added MAGT segment folds locally with decent confidence (P1), but the large apparent shared-region RMSD and PAE data show this segment's placement relative to the HP1-beta body is unresolved rather than a genuine docked element, so it does not qualify as a confidently integrated structured extension. Localization (still nucleus, NLS retained), domain content (no InterPro domain gained/lost), and whole-protein biophysics (negligible GRAVY/charge/disorder shifts) are all unchanged from canonical.
CBX1/HP1-beta's core functions — H3K9me3 reading via its chromodomain, SUV39H1/PRC2 complex assembly, and CK2/Chk2-regulated chromatin mobilization — all depend on nuclear localization and an intact chromodomain/chromoshadow architecture, none of which this extension perturbs: the isoform remains predicted nuclear with its NLS intact and gains no new domain. The seemingly large core-fold RMSD is not trustworthy evidence of altered chromodomain folding given the low pTM and high inter-region PAE, so no functional consequence for heterochromatin reading or complex assembly can be inferred from it.
The 13.9 Å shared-region RMSD and the extension's local pLDDT are undermined by low global pTM (~0.4) and high diff-vs-body PAE (~29 Å), so neither a core refold nor a structured N-terminal extension can be confidently claimed; only D3 (mass spec) is negative among existence evidence, but detection and conservation are otherwise strong, arguing the ORF is real even though no functional mechanism change is well supported.
Folding
Evidence — click any tile for the differential-region detail
LLM reasoning
How well is the isoform-unique region conserved at the amino-acid level across primates (mean percent identity)? High AA identity in close relatives argues the alternative protein is real and translated, not a sequencing or annotation artifact. (Reading-frame intactness is shown as context.)
Comparison
Sequence conservation · primate
mean AA identity over the differential ORF is the score basis; frame intactness across the clade is context
| Conservation metric | Canonical ORF | Differential ORF |
|---|---|---|
| Mean AA identity | 100% | 100% |
| Frame intact (fraction of species) | 84% | 100% |
| Species aligned | 25 | 25 |
| Species frame-intact | 21 | 25 |
| Start codon conserved | 100% | 100% |
| Deepest intact species | Propithecus_coquereli | Microcebus_murinus |
| Phylo depth (MRCA) | 7 | 7 |
The same amino-acid identity test across mammals. Conservation over deeper evolutionary distance is stronger evidence the ORF is under selection to be translated. (Reading-frame intactness is shown as context.)
Comparison
Sequence conservation · mammalian
mean AA identity over the differential ORF is the score basis; frame intactness across the clade is context
| Conservation metric | Canonical ORF | Differential ORF |
|---|---|---|
| Mean AA identity | 98% | 100% |
| Frame intact (fraction of species) | 15% | 100% |
| Species aligned | 20 | 20 |
| Species frame-intact | 3 | 20 |
| Start codon conserved | 100% | 100% |
| Deepest intact species | Callithrix_jacchus | Loxodonta_africana |
| Phylo depth (MRCA) | 6 | 12 |
phyloP / phastCons measure per-base evolutionary constraint. A high absolute phyloP over the differential region means the sequence itself is under strong purifying selection — evidence it is coding and does something. (Enrichment over the shared core is shown as context only.)
Comparison
Per-base conservation · differential region
absolute mean phyloP over the differential region is the score basis (high = strong purifying selection); the shared column and enrichment are context, not the claim
| Track | Differential | Shared | vs shared |
|---|---|---|---|
| phyloP mean | 6.36 | 4.95 | 1.28 |
| phastCons mean | 1 | 0.92 | — |
Start site · canonical vs isoform
initiation context + per-base conservation at each start codon
| Property | Canonical | Isoform |
|---|---|---|
| Start codon | ATG | CTG |
| Kozak context (−9..+4) | GCGGGCACTATGG | ACCAGAAAGCTGG |
| phyloP at start codon | 6.43 | 7.15 |
| phastCons at start codon | 1 | 1 |
| phyloP over Kozak window | 6.31 | 5.31 |
| phastCons over Kozak window | 1 | 1 |
| Kozak mismatch — full consensus | 3 | 7 |
| Kozak window GC content | 0.692 | 0.538 |
LLM reasoning
Is the alternative start used in more than one cell line? Reproducible initiation across independent samples argues against a one-off ribosome artifact.
Comparison
Per-cell-line usage · canonical vs isoform
Fisher q = 2.2e-29
| Cell line | Canonical | Isoform | p-value |
|---|---|---|---|
| HeLa | — | 4.19 | 0.00011 |
| K562 | — | 15.4 | 2.2e-29 |
| U2OS | — | 10.2 | 1.44e-15 |
| RPE1 Async | — | 4.58 | 7.79e-13 |
| RPE1 Que | — | 2.03 | 1.23e-18 |
How efficiently ribosomes initiate at this start (TIS) relative to background, per cell line. Strong, reproducible initiation supports a genuine translation event.
Comparison
Start-Site Usage · canonical vs isoform
ribosome initiation efficiency at the canonical start vs this alternative start, per cell line
| Cell line | Canonical | Isoform |
|---|---|---|
| HeLa | — | 0.114 |
| K562 | — | 0.709 |
| U2OS | — | 0.156 |
| RPE1 Async | — | 0.109 |
| RPE1 Que | — | 0.0998 |
Were tryptic peptides unique to the differential region detected by mass spec? Direct peptide evidence is the strongest proof the isoform protein exists.
Comparison
Peptide Evidence (canonical vs isoform)
isoform-unique peptides are direct evidence the alternative protein exists
| Feature | Canonical | Isoform |
|---|---|---|
| Tryptic peptides (in-silico) | 28 | 4 |
| Validated by mass-spec | 0 | 0 |
| Isoform-unique peptides | — | 4 |
Details
Peptide Evidence (canonical vs isoform)
- peptide MAGTMGK 0–7
- peptide AGTMGK 1–7
- peptide MAGTMGKK 0–8
- peptide AGTMGKK 1–8
LLM reasoning
Do the predicted localization features (DeepLoc prediction, sorting signals, or membrane association) differ between the canonical and the isoform? A protein with changed localization features acts in a different cellular context.
Comparison
Subcellular localization (DeepLoc)
no localization-feature change
| Property | Canonical | Isoform |
|---|---|---|
| Predicted location | Nucleus | Nucleus |
| Sorting signals | Nuclear localization signal | Nuclear localization signal |
| Membrane | Soluble | Soluble |
Do N-terminal targeting signals (SignalP secretion, TargetP mitochondrial/chloroplast) differ between canonical and isoform? N-terminal changes most directly add or remove targeting peptides.
LLM reasoning
Does healthy human germline variation (gnomAD) avoid this region (depletion ratio < 1×), and is it intrinsically constrained (ESM-C constraint delta > 0)? Depletion of population variation plus high sequence constraint mean the region resists change — it is functionally important. Scored only where the unique region is canonical coding sequence, i.e. on truncations. gnomAD is a tolerance catalogue, not a disease one; disease/cancer variants (ClinVar / COSMIC) live in M2.
Comparison
Germline Variants · differential vs shared region
gnomAD (population) variant density per nucleotide — depletion ratio < 1× means healthy human variation avoids the region (constrained), the M1 basis alongside ESM-C constraint
| Variant set | Differential | Shared | Depletion ratio |
|---|---|---|---|
| gnomAD variants | 6 | 192 | 1.4× |
Sequence constraint (ESM-C) · differential vs shared region
lower LLR = more constrained; enrichment = differential / conserved
| Property | Differential | Shared | Enrichment |
|---|---|---|---|
| ESM-C mean LLR | -2.12 | -0.00823 | — |
| Constrained positions | 1 | 0 | — |
Predicted-damaging variants · germline (gnomAD)
AlphaMissense / ESM-C / LoF calls per region (length-normalized)
| Variant set | Differential | Shared | Enrichment |
|---|---|---|---|
| Scorable variants | 4 | 188 | 0.98× |
| Damaging variants | 0 | 103 | 0× |
| — of which loss-of-function | 0 | 13 | 0× |
| AlphaMissense-pathogenic | 0 | 58 | 0× |
Predictor scores · germline (gnomAD)
scored: 260 ESM-C · 168 AlphaMissense
| Property | Differential | Shared |
|---|---|---|
| Mean ΔLLR (ESM-C) | -0.438 | -5.64 |
| Min ΔLLR (ESM-C) | -1.75 | -13.1 |
| Mean AlphaMissense | — | 0.578 |
Are disease (ClinVar / COSMIC) variants enriched per nucleotide in the differential region versus the shared core (disease enrichment ratio ≥ 1×)? Disease variants concentrating in the unique region tie it to phenotype.
Comparison
Clinical-variant burden · differential vs shared region
counts per region; ratio is density-normalized (variants per nucleotide, differential ÷ shared — ≥1× = disease variants concentrate in the differential region, the M2 basis)
| Variant set | Differential | Shared | Enrichment |
|---|---|---|---|
| Disease variants | 2 | 89 | 1× |
| Pathogenic | 0 | 0 | — |
Predicted-damaging variants · disease (ClinVar/COSMIC)
AlphaMissense / ESM-C / LoF calls per region (length-normalized)
| Variant set | Differential | Shared | Enrichment |
|---|---|---|---|
| Scorable variants | 2 | 88 | 1.1× |
| Damaging variants | 0 | 66 | 0× |
| — of which loss-of-function | 0 | 9 | 0× |
| AlphaMissense-pathogenic | 0 | 43 | 0× |
Predictor scores · disease (ClinVar/COSMIC)
scored: 260 ESM-C · 168 AlphaMissense
| Property | Differential | Shared |
|---|---|---|
| Mean ΔLLR (ESM-C) | -0.875 | -7.53 |
| Min ΔLLR (ESM-C) | -1.75 | -13.9 |
| Mean AlphaMissense | — | 0.68 |
LLM reasoning
Does the gained/lost region fold into ordered structure (pLDDT) and shift the protein's biophysical character (charge, hydropathy, disorder)? Structured, biophysically distinct regions are more likely to be functional.
Comparison
Structure (ESMFold2) · canonical vs isoform
TM-score 0.542 · RMSD 3.99 Å · 3 interface contacts
| Metric | Canonical | Isoform |
|---|---|---|
| pLDDT (whole protein) | 0.822 | 0.863 |
Fold confidence · differential vs shared region
mean ESMFold2 pLDDT in each region
| Metric | Differential | Shared | Enrichment |
|---|---|---|---|
| pLDDT | 0.822 | 0.781 | 1.1 |
The shared region is the stretch of protein identical in both the isoform and the canonical (the canonical body for an extension; the post-truncation body for a truncation). Folded in both contexts it normally comes out nearly identical, so its Cα backbone RMSD ≈ 0. A high shared-region RMSD (superposed on the shared residues only, from the ESMFold2 structures) means the extension or truncation reorganizes how the retained region folds — a rare, high-interest functional signal. TM-score is a length-normalized companion; the check only fires when both structures are confidently folded, and uORF/altORF isoforms (no shared region) are not evaluable.
Comparison
Shared-region structural change
Cα RMSD superposed on the shared residues only; TM-score is length-normalized · shared-region Cα RMSD 13.9 Å · shared TM-score 0.547 · shared region 185 aa · min shared pLDDT 0.822 · global TM-score 0.542 · global RMSD 3.99 Å
| Property | Canonical | Isoform |
|---|---|---|
| Shared-region pLDDT | 0.822 | 0.863 |
Does the differential region contain actual secondary structure — a helix or strand — rather than coil? ESMFold2 predicts coordinates but no secondary structure, so elements are assigned from those coordinates (P-SEA, Cα geometry). Because they are derived from a prediction, each element carries its own mean pLDDT: a geometrically clean helix running through a disordered stretch is geometry fitted to a guess, so BOTH length and confidence are required to score. The direction differs by ORF type — an extension GAINS the element (a candidate functional addition), a truncation LOSES one, and there the element is read off the canonical structure because the removed segment exists only in it. This says nothing about whether the element is integrated with the rest of the fold; the contact and PAE evidence in this same category answers that.
Comparison
Secondary structure · differential vs shared region
helices and strands assigned from the predicted coordinates (P-SEA)
| Metric | Differential | Shared |
|---|---|---|
| Alpha helices | 0 | 3 |
| Beta strands | 1 | 2 |
| Longest element (aa) | 7 | 19 |
| Mean pLDDT | 0.79 | 0.88 |
Elements and coordinates
1 in the differential region, 5 in the shared core — residue numbering is 1-based on the protein holding the region
| Isoform-unique | Shared core |
|---|---|
| beta strand 2–8 7 aa · pLDDT 0.79 | alpha helix 65–83 19 aa · pLDDT 0.83 |
| — | beta strand 87–92 6 aa · pLDDT 0.75 |
| — | beta strand 147–152 6 aa · pLDDT 0.93 |
| — | alpha helix 153–159 7 aa · pLDDT 0.95 |
| — | alpha helix 161–171 11 aa · pLDDT 0.96 |
Below threshold
0 in the differential region, 4 in the shared core — shorter than 6 aa or below pLDDT 0.70, so not counted above
| Isoform-unique | Shared core |
|---|---|
| — | beta strand 39–43 5 aa · pLDDT 0.96 |
| — | beta strand 54–57 4 aa · pLDDT 0.97 |
| — | beta strand 99–103 5 aa · pLDDT 0.81 |
| — | beta strand 172–176 5 aa · pLDDT 0.88 |
LLM reasoning
Does the isoform gain or lose a real InterPro functional domain in the differential region (disorder/structural-only signatures excluded)? Gaining or losing a domain changes function directly.
Comparison
Domains & motifs (canonical vs isoform)
gained = only in the isoform; lost = only in the canonical
| Feature | Canonical | Isoform |
|---|---|---|
| InterPro domains | 15 | 16 |
| Short linear motifs | 1 | 1 |
Details
Domains & motifs (canonical vs isoform)
- gained Coil domain
Biophysical character of the isoform-differential region versus the shared canonical core — pI, hydropathy, charge, disorder and related properties. Descriptive; the folding (P1) score keys off the GRAVY / charge / disorder deltas.
Comparison
Biophysics · differential vs shared region
highlighted rows are enriched in the differential region
| Property | Differential | Shared | Enrichment |
|---|---|---|---|
| Isoelectric point (pI) | 6.01 | 4.54 | 1.33 |
| Hydropathy (GRAVY) | 0.65 | -1.28 | -0.508 |
| Fraction charged | 0 | 0.46 | 0 |
| Disorder fraction | -0.028 | 0.232 | -0.121 |
| Disorder-promoting | 0.5 | 0.681 | 0.734 |
| Low-complexity fraction | 0 | 0.222 | 0 |
| Prion-like fraction | 0.25 | 0.2 | 1.25 |
| LLPS score | 0.0555 | 0.192 | 0.289 |
| π–π propensity | 0 | 0.173 | 0 |
| Aromaticity | 0 | 0.0703 | 0 |
| Instability index | 17.1 | 53 | 0.323 |
| Shannon entropy | 2 | 3.9 | 0.513 |
| Normalized complexity | 1 | 0.902 | 1.11 |
Sparse-autoencoder (SAE) interpretability features on the ESM-C residual stream, comparing the isoform against the canonical protein. This is the S3 criterion — a presence check that fires when any interpretable feature is gained or lost. Only the top 30 features per category (ranked by activation, prevalence ≥ 2) are saved; each table shows the top 10 interpretable features, and generic / "unknown" features are hidden unless you expand it. What are SAE features? →
Part 1 · Global feature comparison
Which SAE features fire in the isoform vs the canonical protein, across the whole sequence.
Isoform-only features — 26 total (prevalence ≥ 2)
| Feature | Label | Description | Activation | Prevalence |
|---|---|---|---|---|
| #4992 | Small-residue disordered N-terminus | Context-dependent free N-terminus signature: short, disordered amino-terminal stretches beginning with the initiator Met followed by residues compatible with MAP cleavage or N-acetylation (e.g., M–G/A/S/T/V, and other contexts including M–D/E/L); strongest emphasis on small residues at positions 2–3; activation is absent when the native N-terminus is structured or starts with a signal peptide/transmembrane segment. | 13.42 | 3 |
| #448 | N-terminal leader/targeting segments | N-terminal leader/targeting segments: the feature activates on the extreme N-terminus (first 2–5 residues) of proteins, especially the N-region of signal peptides and organelle transit peptides, and more generally on short, disordered/low-complexity N-terminal tails; strongest at residue ~3 and largely independent of amino-acid identity, occurring across all taxa and functions. | 11.94 | 2 |
| #3609 | Position-2 N-terminus residue detector | Detector of the identity of the second residue at the extreme N-terminus of proteins, with strongest preference for small residues (especially Ser, Ala, and Gly; also Thr/Pro), i.e., an N-terminal position-2 signal commonly associated with generic, disordered starts and co-translational N-terminus processing | 10.78 | 2 |
| #11504 | Initiator methionine at M1 | Initiator methionine at the very start of the polypeptide chain (M1), i.e., the translation start residue, independent of protein family, taxonomy, membrane association, or presence of signal/leader/propeptide regions; often annotated as post-translationally removed and typically situated in a flexible, coil-like N-terminus. | 10.41 | 2 |
| #9998 | Weak N+3 positional cue | Absolute N-terminal positional cue centered near the fourth residue (N+3), independent of amino-acid identity; most often within signal/transit peptides of precursors but also present in non-secreted proteins; the effect is weak and often does not produce a discrete detected site. | 8.61 | 2 |
| #5482 | N-terminal targeting presequences | Universal eukaryotic N-terminal targeting presequences: the feature detects short, cleavable leader regions at the extreme N-terminus that direct proteins to organelles or the secretory pathway—especially chloroplast/apicoplast transit peptides and thylakoid lumen signals, but also mitochondrial targeting peptides and classical signal peptides. These segments are Ser/Thr- and small/hydrophobic–rich, enriched in Lys/Arg and depleted of acidic residues, typically low-structure/low-confidence and ending at the maturation cleavage site. | 4.93 | 2 |
| #5241 | Early N-terminal positional signal | Generic early N-terminus positional signal peaking at residue ~5–7, found broadly within the disordered N-tail of proteins, including both cleavable leader/targeting segments (signal peptides, propeptides) of precursor proteins and disordered N-terminal regions of mature proteins; residue identity is not specific | 4.76 | 2 |
| #14646 | Short N-terminal leader segment | Short N-terminal segments within the first ~10–40 amino acids, often Lys/Arg-enriched but sometimes purely hydrophobic. These commonly correspond to Sec-type signal peptides, signal anchors, or other short N-terminal segments preceding/overlapping the first hydrophobic helix, and also include disordered N-terminal tails in soluble proteins. | 2.92 | 4 |
| #13772 | CCT QREAAL; signal peptide core | Detector with two recurring activation modes — (1) a conserved motif within the CCT domain of plant pseudo-response regulators (peaks on the A of a "QREAAL" stretch, with secondary peaks at a downstream Pro and at a Glu/Ser within the upstream Response Regulatory domain), and (2) the hydrophobic core of secretory signal peptides in short peptide hormone precursors (neuropeptide W, ghrelin). | 2.64 | 2 |
| #11448 | VIT domain signature | VIT (Vault protein Inter-alpha-Trypsin) domain feature, firing within the ~120-residue β-sheet VIT module found N-terminal to VWA domains in inter-alpha-trypsin inhibitor heavy chains and related proteins. | 2.28 | 2 |
| #15799 | Aliphatic β-strand core detector | Generic detector of short, aliphatic-rich β‑strand segments that form the cores of β‑sheets in diverse domain scaffolds and repeats (e.g., PDZ, SRCR, LysM, KH/R3H, ubiquitin/UBL β‑grasp), rather than a function-specific motif; sequence windows show alternating small/polar and hydrophobic residues, strongly enriched in V/L/I/A and often with nearby Ser/Thr, and in secreted modules may be flanked by cysteines. | 2.17 | 2 |
| #16041 | L27-like helical interaction modules | Eukaryotic protein–protein interaction modules built from short α-helical bundles, in particular L27 domains and L27-like helical interaction regions found in MAGUK/MPP scaffolds and in LUBAC/Sharpin-type adaptors. | 1.86 | 3 |
| #9096 | Intrinsically disordered low-complexity regions | Compositionally biased, intrinsically disordered low‑complexity segments (often short tandem‑repeat regions) enriched in small/flexible residues and sometimes cysteine/glutamine, commonly occurring in N‑terminal/propeptide stretches of secreted precursors and small unstructured proteins; peaks can also occur in flexible terminal/loop segments within otherwise structured proteins (including enzymes); peaks can fall at residues within simple repeats or immediately flanking processed peptide segments, but the unifying concept is IDR/low‑complexity or flexibly disordered segments rather than a specific domain. | 1.85 | 2 |
| #4850 | RRM-disordered linker activation | RRM-containing eukaryotic RNA-binding proteins, with broad activation spanning both the RRM domains and adjacent/inter-RRM disordered regions; covers splicing factors, hnRNP family, SR-type shuttling mRNA-binding proteins, cyclophilin-type RRM proteins, and plant glycine-rich RBPs. | 1.77 | 2 |
| #10487 | Nuclear regulatory low-complexity IDRs | Intrinsically disordered, low‑complexity regulatory regions of large metazoan regulatory and scaffolding proteins—enriched in nuclear gene‑regulatory factors—especially transactivation domains, flexible linkers, and long regulatory tails that are S/P/Q/G/T- and acidic‑rich, harbor SP/TP phospho‑motifs and short homorepeats; the feature preferentially avoids compact folded DNA-/ligand-/zinc‑binding domains. | 1.76 | 2 |
| #11222 | N-terminal transcriptional regulatory IDRs | Intrinsically disordered, low-complexity N-terminal regulatory regions of eukaryotic transcription factors—segments enriched for polar and basic/acidic residues with frequent Ser/Thr and Pro content—positioned immediately upstream of structured DNA-binding domains (e.g., homeobox, bHLH) and corresponding to transcriptional regulatory tails (activation/repression domains) subject to phosphorylation-dependent control. | 1.66 | 2 |
| #2092 | AlaRS N-terminal FFxxxGF/Y motif | A conserved short N-terminal motif in the aminoacylation domain of alanyl-tRNA ligases (AlaRS), centered on a hydrophobic/aromatic stretch with mixed charged residues (consensus loosely matching F-F-x-x-x-G-F/Y). | 1.64 | 3 |
| #6961 | PDZ-flanking low-complexity IDRs | Intrinsically disordered linker and tail regions that flank or connect folded domains—most prominently between and after PDZ domains—in large metazoan scaffold/adaptor proteins (synaptic/junctional PDZ scaffolds, MAGUK-related adaptors, exocytosis regulators); the feature largely avoids the structured PDZ cores and highlights compositionally biased regulatory segments rich in polar, basic/acidic, and proline residues. | 1.64 | 2 |
| #11427 | Short domain-boundary docking motif | A signal for short, well-structured docking/interaction elements (often at helix→β-strand junctions or at boundaries between structured segments) that mediate protein–protein interactions in trafficking and signaling proteins. | 1.62 | 2 |
| #4002 | Disordered low-complexity termini/linkers | Intrinsically disordered, low-complexity terminal and linker regions enriched in polar/acidic and proline-rich content—covering transcriptional activation/repression segments and targeting peptides (signal peptides, mitochondrial transit peptides)—with depletion in compact folded domains such as HMG boxes and other stable globular folds | 1.58 | 2 |
| #8426 | Acidic serine-rich IDRs | Phospho-dense, intrinsically disordered low‑complexity tracts enriched in Ser/Pro and acidic residues (often interspersed with Lys/Arg; mixed‑charge “polyampholyte” segments) that serve as flexible regulatory/interaction regions in eukaryotic proteins, frequently near SAP/LEM modules and in other nuclear or membrane-associated scaffolds (variable across SUN/PEX14 families). | 1.58 | 2 |
| #5193 | S/T-Pro-rich disordered activation tails | Intrinsically disordered, low‑complexity serine/threonine‑ and proline‑rich regulatory segments (typically N‑terminal) that contain clustered proline‑directed phosphorylation motifs (e.g., SP/TP) and serve as phospho‑regulated activation/interaction tails in transcription factors and chromatin proteins; the feature avoids folded DNA‑binding cores and other structured domains | 1.55 | 2 |
| #13401 | Pre-J leader motif activation | Short “pre-J” leader segments immediately N‑terminal to J/J‑like domains in DnaJ‑family and J‑like proteins—typically low‑complexity, charged or Gly/Ser/Pro‑rich coils or short amphipathic helices (often organelle transit peptides or coiled‑coil stubs) that precede the folded J domain; the J/J‑like core itself is not favored. A secondary signal appears at analogous short N‑terminal motifs at the start of other domains (e.g., the YRG start of AP2/ERF). | 1.52 | 2 |
| #3978 | Short N-terminal targeting signals | N-terminal short leaders/motifs: classical Sec-type signal peptides plus other short N-terminal export/targeting/assembly signals (signal anchors, organellar transit peptides, type III/flagellar export signals, and amphipathic/basic N-termini used for macromolecular binding). | 1.52 | 2 |
| #5571 | Periplasmic beta-strand boundary | Short segments within the periplasmic/extracellular domains of bacterial secretion-system membrane components, typically corresponding to N-terminal β-strand–containing regions of HR/PDZ-like subdomains near the membrane-proximal end, sometimes adjacent to disordered linkers. | 1.41 | 2 |
| #6701 | Kinase αF and analogous helices | Conserved long hydrophobic α-helices within structured protein cores—most prominently the kinase C-lobe αF helix—plus analogous helices in ankyrin-repeat scaffolds and other folded helical bundles; strongest in kinases but broadly present across diverse families. | 1.40 | 2 |
Canonical-only features — 10 total (prevalence ≥ 2)
| Feature | Label | Description | Activation | Prevalence |
|---|---|---|---|---|
| #2164 | Conserved protein kinase catalytic core | Conserved protein kinase catalytic core (Hanks-type eukaryotic-like Ser/Thr and dual-specificity kinases), responding to the Pkinase domain across cellular and viral kinases, including kinase-like/pseudokinase variants | 2.16 | 2 |
| #1218 | Low-complexity disordered regions | Compositionally biased, intrinsically disordered low‑complexity regions (LCRs), especially repetitive S/T/P/G‑ and cysteine‑rich segments in coils and unstructured tails, rather than structured/catalytic domains. | 2.10 | 2 |
| #9446 | Amphipathic surface interaction helix | A short, contiguous, amphipathic alpha-helix enriched in charged residues (often Lys/Glu/Gly) that serves as a generic interaction helix on the surface of large NTP-dependent enzymes and RNA–protein assemblies—most commonly nucleic‑acid–processing machineries, but also observed in certain membrane-associated metabolic ATPases (e.g., ABC1/COQ8)—frequently corresponding to a helical element of helix–hairpin modules (e.g., H2TH/HhH/HTH-like) or a transducer/adjacent helix near NTP/nucleotide-binding cores, rather than a catalytic loop or active-site motif. | 2.05 | 2 |
| #14431 | Ser/Thr-rich disordered motifs | Low-complexity, intrinsically disordered segments with a bias for serine/threonine (often with proline and nearby basic/acidic residues) that encode short linear motifs for regulation or targeting (e.g., phosphosite motifs, RxLR-dEER), found in flexible regulatory tails, disordered viral protein regions, and short secreted peptide precursors. | 1.95 | 2 |
| #7295 | N-terminal leader/linker segments | Short N-terminal leader/linker segments preceding the catalytic core, including helix N-caps, adjacent flexible tails, and short strand or coil segments at the very start of the mature chain; these regions often act as targeting/regulatory tails or flexible connectors rather than catalytic motifs, with similar helix–coil boundary motifs also captured internally. | 1.81 | 2 |
| #5824 | Pro/Gly-rich beta-strand caps | Short Pro/Gly-enriched β-strand edge/turn motifs at strand–loop (and loop–helix) junctions—often at the N-terminus of a domain—i.e., β-hairpin turns and strand N-caps used as structural connectors rather than catalytic sites | 1.71 | 2 |
| #7907 | Domain-edge mixed secondary surface segment | A broad surface segment at domain boundaries—often a region just after a signal peptide or N-terminal cap in secreted/periplasmic carbohydrate-active enzymes, and analogous regions in other proteins—that mixes disordered coils with short beta-strands/loops and lies outside catalytic residues but can host substrate-contact or modification sites. | 1.67 | 2 |
| #1914 | DE-rich acidic IDR tails | DE-rich, low-complexity intrinsically disordered acidic tracts—typically terminal tails—characterized by long Asp/Glu‑enriched stretches (often with G/S/P) in proteins from large macromolecular assemblies (e.g., transcription/translation/proteostasis complexes). | 1.67 | 2 |
| #3180 | Noncatalytic C-terminal interaction regions | Non-catalytic C-terminal interaction regions—long, charged/polar terminal domains or tails (often low-complexity or repeat-based) that mediate assembly, trafficking, and RNA/protein binding, rather than enzymatic catalysis; includes structured terminal modules (e.g., NTF2-like) and flexible extensions, with a strong bias toward the final domain/segment of the protein. | 1.48 | 2 |
| #10265 | C-terminal effector modules | C-terminal functional regions that serve as terminal interaction/effector modules—typically the last domain or tail of a protein, often intrinsically disordered activation/sorting segments in regulators and secretory proteins, but also the final repeat/domain in multi-repeat scaffolds. | 1.44 | 2 |
Shared features by |Δ| activation — 945 total (prevalence ≥ 2)
| Feature | Label | Description | Δ activation | Iso act. | Canon act. | Iso prev. | Canon prev. |
|---|---|---|---|---|---|---|---|
| #7996 | Generic extreme N-terminus detector | Generic extreme N-terminus detector: marks residues around position 5 (occasionally 6–7) at the very start of the polypeptide, most often within the N-region of signal/transit peptides of secreted or organelle-targeted precursors, but also in generally disordered N-termini of cytosolic/nuclear proteins; not motif- or residue-specific | -2.80 | 5.82 | 8.62 | 3 | 2 |
| #10977 | Early N-terminal disordered patch | Short N-terminal patch at the extreme N-terminus (around residues 5–10), capturing the first distinctive segment of preproteins/precursors and N-terminal disordered tails; residue composition is heterogeneous but frequently polar (Ser/Thr) and/or charged. | -2.42 | 4.41 | 6.83 | 3 | 3 |
| #15815 | Basic N-terminal targeting patch | Short, basic/polar N‑terminal leader/transit segment immediately after the initiator methionine (positions ~3–6) — i.e., the positively charged/polar start of signal peptides and organelle transit peptides, and more generally low‑acidic, disordered N‑terminal patches (including occasional cysteine‑rich starts) used for targeting, export, rRNA binding, or membrane topology cues | -1.99 | 7.71 | 9.70 | 6 | 4 |
| #8063 | Disordered N-terminal activation window | Short N-terminal disordered segment found in a variety of proteins, including plant chloroplast transit peptides, bacterial prokaryotic ubiquitin-like protein Pup N-termini, anti-sigma factor N-tails, and some small RiPP precursor leaders. The activating region is Ser/Thr/Ala/Pro/Gly/Gln-rich, low in acidic residues, and typically lies in a disordered N-terminal stretch; the feature does not generally mark all transit peptides or all RiPP leaders. | -1.35 | 3.99 | 5.34 | 8 | 7 |
| #6484 | Prion-like low-complexity IDRs | Polar low-complexity intrinsically disordered regions (IDRs)—especially Q/N-rich (prion-like) segments and S/T- and glycine-rich tracts with simple SG/TG/PG repeats—typically located in inter-domain linkers and terminal tails of eukaryotic proteins and often harboring Ser/Thr phosphorylation sites. | +1.00 | 5.12 | 4.13 | 6 | 6 |
| #13480 | Alanine and small-residue enrichment | A sequence-composition feature that detects small, non-aromatic residues—especially alanine—and their enrichment, lighting up individual Ala residues and stretches enriched in A/G/S/T/P/V (alanine-rich, low-complexity segments) irrespective of protein function or secondary structure. | +0.69 | 5.12 | 4.43 | 9 | 8 |
| #13334 | N-terminal IDR boundary | Recognition of short, low-complexity, intrinsically disordered N-terminal tails immediately upstream of the first folded domain; these segments are often Ser/Thr/Pro/Gly/Ala-biased and sometimes extend a few residues into the initial helix or the very beginning of the first domain (domain boundary/disorder-to-order transition), especially in transcription regulators and co-regulators. | +0.66 | 5.66 | 5.00 | 28 | 25 |
| #654 | Long low-complexity IDRs | Long, low-complexity intrinsically disordered regions (IDRs), typically enriched in S/P/Q/E/A/G and frequently occupying extended N‑terminal segments that serve as propeptides, transcriptional activation domains, flexible linkers, or receptor tails/stalks; activation drops sharply in adjacent folded motifs (coiled-coils, globular domains, zinc-finger/Ig-like modules, BH3 helix). | +0.57 | 2.30 | 1.73 | 6 | 6 |
| #6384 | Widespread activation in folded proteins | Proteins showing broad, diffuse activation across large fractions of their sequence, with peaks frequently occurring in folded regions of single-domain or multi-domain enzymes and structural proteins. | +0.54 | 3.11 | 2.57 | 28 | 23 |
| #6205 | Broad domain-level activation | A broad, low-amplitude sensor that activates across most of the sequence in folded proteins, with particular prominence in DEAD-box / SF1-SF2 RNA helicase ATP-binding domains. Activation is broadly distributed rather than localized to specific catalytic residues, and tends to be weaker on N-terminal targeting/transit peptides and cleavable leaders than on the mature chain. | -0.49 | 3.23 | 3.72 | 16 | 17 |
| #14891 | Disordered termini and cleavage detector | Residue-level detector of intrinsically disordered, flexible termini and proteolytic processing junctions—especially N-termini (including the first residue of the mature chain after propeptide/leader cleavage)—with no strict amino‑acid specificity and a bias toward small/charged residues; common in small, secreted/precursor and viral proteins but also present in disordered regions of diverse proteins. | +0.49 | 3.02 | 2.53 | 4 | 3 |
| #4358 | Low complexity disordered regions | Compositionally biased, low-complexity segments enriched in small/polar residues (serine, glycine, asparagine, proline) and often basic (lysine/arginine) or acidic tracts; the feature preferentially marks IDRs in N/C-terminal tails and repeat-rich regions, though some activations also occur within folded domains. | +0.43 | 3.13 | 2.70 | 10 | 8 |
| #512 | Short functional domain segments | Short sequence segments distributed across diverse proteins, with peaks that often fall within structured catalytic, transporter, or fold-defining domains as well as occasional acidic/charged stretches and flexible linkers. | +0.41 | 2.22 | 1.82 | 8 | 5 |
| #13204 | SRA/YDG DNA methylation reader | YDG/SRA domain of UHRF1-family and related RING-type E3 ubiquitin ligases involved in DNA methylation reading, with activation concentrated on the structured base-flipping/5-methylcytosine-binding region. | +0.40 | 4.12 | 3.73 | 26 | 23 |
| #1370 | Short low-complexity N-terminal segment | Short, low‑complexity, intrinsically disordered N‑terminal segments (~10–20 aa) that frequently correspond to leader/transit/propeptide presequences (e.g., chloroplast transit peptides, RiPP/bacteriocin leaders) or generic unstructured N‑tails; typically non‑hydrophobic and enriched in Ser/Thr/Pro but tolerant of basic (poly‑Arg/Lys) or Cys‑rich compositions | -0.39 | 3.02 | 3.42 | 7 | 8 |
| #9586 | NAC pre-core regulatory IDRs | Leading subregion of plant NAC-family DNA-binding domains, corresponding to the pre-core/flexible/dimerization-leaning portion of the fold that precedes the DNA-contacting residues; the feature targets this internal NAC sub-segment rather than the downstream DNA-contacting core. Also active on related N-terminal regulatory segments of eukaryotic transcription factors and small viral regulatory proteins. | +0.31 | 3.07 | 2.75 | 33 | 29 |
| #10704 | N-terminal PEST-like IDRs | Low-complexity, often acidic/proline-rich intrinsically disordered N-terminal tails of intracellular eukaryotic proteins, frequently overlapping annotated PEST-like segments. | -0.28 | 8.59 | 8.87 | 14 | 10 |
| #2319 | Threonine residue detector | Residue-identity detector for threonine (Thr): activates on individual Thr residues across diverse proteins, with a mild enrichment in low-complexity/disordered, repeat-rich, and secretory regions; still marks Thr within well-structured domains; occasional weak spillover to serine. | +0.28 | 5.19 | 4.91 | 6 | 5 |
| #6016 | Helix-biased internal methionine detector | Detector for methionine residues, firing on internal Met across diverse proteins with somewhat enhanced response when Met occurs in helical or low-complexity contexts; occasional hits at the initiator Met when embedded in a locally Met- or hydrophobic-rich N-terminus; weak cross-reactivity to other bulky hydrophobics (notably tryptophan). | +0.27 | 6.10 | 5.83 | 4 | 3 |
| #12567 | N-terminal amphipathic leader helix | A short, compositionally biased N-terminal segment around positions ~18–40 that is enriched in charged/polar residues (often with glycine runs) and behaves as an intrinsically disordered patch with strong propensity to form a single short amphipathic alpha-helix; this generic “leader-like” interface is used across proteins as a targeting/secretion/assembly or interaction module (rather than a specific sequence motif). | -0.27 | 2.04 | 2.30 | 2 | 3 |
| #11402 | Juxta-kinase regulatory segment | Juxta-kinase regulatory segment immediately upstream of eukaryotic protein kinase catalytic domains: a kinase-proximal linker that precedes the conserved core, often disordered or in a helix→coil transition that interfaces with the N-lobe; the feature avoids the catalytic motifs (P-loop, VAIK Lys, HRD/DFG) and distant disordered tails. | -0.25 | 1.79 | 2.04 | 4 | 2 |
| #11245 | Disordered KR-rich NLS-like motifs | Arg/Lys-rich low-complexity patches in intrinsically disordered regions that function as generic basic, nucleic-acid–contacting and nuclear-targeting motifs, encompassing K-rich PKKK/AAPKKK tracts, mixed KR repeats, and classical Lys/Arg-rich nuclear localization signals. | +0.25 | 3.29 | 3.04 | 12 | 10 |
| #6356 | Ser/Gly/Pro-rich low-complexity IDRs | Often fires on intrinsically disordered, low-complexity regions (IDRs) enriched in Ser/Gly/Pro and basic (Lys/Arg) residues, frequently in secreted/viral proteins, but also activates on short peptides and small folded proteins with no clear sequence motif. Includes Ser-rich PTM-prone segments within unstructured regions. | +0.25 | 3.99 | 3.75 | 9 | 8 |
| #6360 | N-terminal low-complexity IDRs | Eukaryotic intrinsically disordered, low‑complexity N‑terminal tails enriched in Lys/Arg/Ser/Asp/Glu that precede structured domains of RNA biogenesis/translation and chromatin‑remodeling factors (nucleolar/ssu‑processome, spliceosome/DEAD‑box helicases, H2A.Z chaperones); these flexible segments are typical interaction/assembly regions. | +0.24 | 5.01 | 4.77 | 27 | 23 |
| #13 | N-terminal β-strand block | N-terminal structural elements—particularly the first β-strand(s) of small folded domains and the N-terminal segment immediately preceding/following them—across phage capsid proteins, SecB-family chaperones, signal peptidase complex subunits, and other small bacterial/archaeal proteins. | -0.22 | 1.88 | 2.10 | 2 | 3 |
| #5076 | N-terminal disordered pre-sequences | Intrinsic N-terminal pre-sequences and regulatory tails: extended, low-complexity, intrinsically disordered segments at protein termini—especially cleavable signal/transit peptides and propeptides, as well as noncleavable acidic/PST-rich regulatory tails that harbor PTM sites and linear binding motifs—marking the boundary before structured mature domains. | +0.21 | 3.99 | 3.78 | 66 | 64 |
| #9632 | Disordered acidic/mixed-charge interaction modules | Intrinsically disordered segments in diverse phosphoproteins and chromatin-associated factors—often with mixed basic/acidic or acidic-leaning composition—that act as flexible interaction modules; widely used in chromatin/transcription factors and also present in diverse acidic phosphoproteins across taxa. | +0.21 | 4.33 | 4.12 | 10 | 6 |
| #1803 | Unknown generic feature | Unknown generic feature | +2.08 | 19.38 | 17.30 | 186 | 182 |
| #14534 | Unknown generic feature | Unknown generic feature | +1.24 | 15.84 | 14.60 | 185 | 181 |
| #9005 | Unknown generic feature | Unknown generic feature | +0.38 | 10.93 | 10.55 | 168 | 165 |
Part 2 · Differential coordinates
Features firing on just the isoform-unique (extension / alt-frame) residues.
Unique-region features — 60 distinct (prevalence ≥ 2)
| Feature | Label | Description | Activation | Prevalence |
|---|---|---|---|---|
| #4992 | Small-residue disordered N-terminus | Context-dependent free N-terminus signature: short, disordered amino-terminal stretches beginning with the initiator Met followed by residues compatible with MAP cleavage or N-acetylation (e.g., M–G/A/S/T/V, and other contexts including M–D/E/L); strongest emphasis on small residues at positions 2–3; activation is absent when the native N-terminus is structured or starts with a signal peptide/transmembrane segment. | 13.42 | 2 |
| #13702 | Regulatory low-complexity IDRs | Eukaryotic intrinsically disordered, low‑complexity regulatory segments (terminal tails and inter‑domain linkers) enriched in Pro/Ser/Thr/acidic composition and short linear motifs for modular protein–protein interactions—frequently including WW‑domain–binding PY/PPxY segments and PTM hotspots—with reduced but not absent activation within folded recognition/catalytic domains (WW, PTB/PID, chromo/chromoshadow, SET). | 11.11 | 4 |
| #2114 | Chromodomain and integrase-proximal activation | A feature that activates broadly across chromodomain-containing proteins and other chromatin-associated factors, as well as on retrotransposon Gag-Pol polyproteins (in integrase-proximal regions). Activation is distributed widely along the chain with peaks tending to fall within or adjacent to chromodomain folds rather than on catalytic enzyme cores. | 10.14 | 4 |
| #10704 | N-terminal PEST-like IDRs | Low-complexity, often acidic/proline-rich intrinsically disordered N-terminal tails of intracellular eukaryotic proteins, frequently overlapping annotated PEST-like segments. | 8.59 | 4 |
| #15815 | Basic N-terminal targeting patch | Short, basic/polar N‑terminal leader/transit segment immediately after the initiator methionine (positions ~3–6) — i.e., the positively charged/polar start of signal peptides and organelle transit peptides, and more generally low‑acidic, disordered N‑terminal patches (including occasional cysteine‑rich starts) used for targeting, export, rRNA binding, or membrane topology cues | 7.71 | 2 |
| #7986 | Intrinsically disordered low-complexity regions | Compositionally biased, intrinsically disordered low‑complexity regions used as flexible interaction/nucleic‑acid–binding modules—enriched in Pro/Ser/Lys/Glu/Gln/Thr/Gly/Arg and encompassing poly‑Pro tracts, Lys/Arg‑rich basic stretches, and acidic Asp/Glu clusters—common in transcription/RNA‑processing factors, viral basic regulators, and the C‑terminal scaffolds of RNase E | 6.22 | 4 |
| #13334 | N-terminal IDR boundary | Recognition of short, low-complexity, intrinsically disordered N-terminal tails immediately upstream of the first folded domain; these segments are often Ser/Thr/Pro/Gly/Ala-biased and sometimes extend a few residues into the initial helix or the very beginning of the first domain (domain boundary/disorder-to-order transition), especially in transcription regulators and co-regulators. | 4.69 | 4 |
| #3298 | Chromatin reader N-terminal motifs | Eukaryotic chromatin- and chromosome-associated nuclear factors, including histone methyltransferases/demethylases and chromatin remodelers. Recognition is primarily via structured reader/auxiliary modules (notably Tudor domains) and short conserved motifs in N-terminal regions, with secondary activation over adjacent disordered charged segments. | 4.66 | 3 |
| #8038 | Disordered chromatin regulatory tails | Intrinsically disordered, low‑complexity terminal regions of eukaryotic chromatin/chromosome–maintenance proteins, enriched for Ser/Thr–Pro dipeptides and alternating acidic (E/D) and Lys-rich patches—i.e., regulatory tails that host proline‑directed phosphorylation and chromatin‑association segments | 4.62 | 4 |
| #9632 | Disordered acidic/mixed-charge interaction modules | Intrinsically disordered segments in diverse phosphoproteins and chromatin-associated factors—often with mixed basic/acidic or acidic-leaning composition—that act as flexible interaction modules; widely used in chromatin/transcription factors and also present in diverse acidic phosphoproteins across taxa. | 4.33 | 4 |
| #5578 | Long low-complexity IDRs | Long, non-globular low‑complexity/IDR segments in eukaryotic proteins—often proline/Ser‑rich regions that host linear motifs (frequently SH3‑recognition PXXP), flexible linkers, or targeting pre‑sequences—while excluding compact folded domains (e.g., SH3, PWWP, catalytic cores). Includes N‑terminal signal peptides or TM‑adjacent stretches when embedded within otherwise disordered regions. | 4.29 | 4 |
| #6360 | N-terminal low-complexity IDRs | Eukaryotic intrinsically disordered, low‑complexity N‑terminal tails enriched in Lys/Arg/Ser/Asp/Glu that precede structured domains of RNA biogenesis/translation and chromatin‑remodeling factors (nucleolar/ssu‑processome, spliceosome/DEAD‑box helicases, H2A.Z chaperones); these flexible segments are typical interaction/assembly regions. | 4.27 | 4 |
| #13204 | SRA/YDG DNA methylation reader | YDG/SRA domain of UHRF1-family and related RING-type E3 ubiquitin ligases involved in DNA methylation reading, with activation concentrated on the structured base-flipping/5-methylcytosine-binding region. | 4.12 | 4 |
| #5076 | N-terminal disordered pre-sequences | Intrinsic N-terminal pre-sequences and regulatory tails: extended, low-complexity, intrinsically disordered segments at protein termini—especially cleavable signal/transit peptides and propeptides, as well as noncleavable acidic/PST-rich regulatory tails that harbor PTM sites and linear binding motifs—marking the boundary before structured mature domains. | 3.76 | 4 |
| #9627 | Short Pro/Gly-rich strand-turn patches | Short, compositionally biased strand/turn segments that nucleate or flank brief secondary‑structure elements (typically beta‑strands with adjacent loops/turns, occasionally short helices), enriched in Pro and Gly together with aliphatic residues (Leu/Ile/Val) and frequent Ser/Thr; these patches occur both within folded beta‑sheet domains and in intrinsically disordered regions (often at coil↔secondary‑structure transitions), and can host linear motifs such as leucine‑rich nuclear export signals or mark the first structured residues after signal‑peptide cleavage. | 3.52 | 4 |
| #15150 | Charged disordered terminal tails | Intrinsically disordered, low‑complexity, charge‑biased tails (long N- or C‑terminal IDRs enriched in Lys/Arg and/or Asp/Glu, often with K–A–P repeats and Ser/Thr phosphorylation sites). The feature primarily recognizes the sequence/structural property of these polyelectrolyte IDRs—which are especially common in chromatin/nuclear proteins—while excluding folded domains. | 3.16 | 4 |
| #4358 | Low complexity disordered regions | Compositionally biased, low-complexity segments enriched in small/polar residues (serine, glycine, asparagine, proline) and often basic (lysine/arginine) or acidic tracts; the feature preferentially marks IDRs in N/C-terminal tails and repeat-rich regions, though some activations also occur within folded domains. | 3.13 | 2 |
| #6384 | Widespread activation in folded proteins | Proteins showing broad, diffuse activation across large fractions of their sequence, with peaks frequently occurring in folded regions of single-domain or multi-domain enzymes and structural proteins. | 3.11 | 4 |
| #8694 | Acidic S/T-rich IDRs | Acidic, serine/threonine-rich low-complexity intrinsically disordered regions (often with S/T-P motifs) in eukaryotic nuclear/chromatin-associated proteins, typically residing in long N- or C-terminal regulatory tails enriched for phosphorylation sites. | 3.09 | 3 |
| #8093 | Diffuse activation with motif peaks | Broadly distributed activation across many protein regions, including structured immunoglobulin-like domains, DNA-binding HTH motifs, and N-terminal chloroplast transit peptides, with characteristically high fraction of active residues per protein | 3.08 | 2 |
| #7095 | N-terminal start-of-chain detector | N-terminal start-of-chain detector that recognizes signal peptides and the immediate post-cleavage beginning of the mature polypeptide, with strong emphasis on conserved N-terminal motifs of secreted, disulfide-rich peptides (e.g., paired/first cysteines in chemokines/defensins); more generally, it marks the extreme N-terminal tail/first helix across diverse proteins. | 3.01 | 4 |
| #6854 | Low complexity helical repeats | A general, composition-driven signal for non-globular sequence regions: it prefers low-complexity, compositionally biased tracts and repetitive helical polymers, emphasizing charged/polar (E, D, K, R, S, Q) or Gly/Pro-rich content, and residues at helix–coil transitions; it is largely muted in compact, well-packed helical cores. | 3.00 | 4 |
| #14646 | Short N-terminal leader segment | Short N-terminal segments within the first ~10–40 amino acids, often Lys/Arg-enriched but sometimes purely hydrophobic. These commonly correspond to Sec-type signal peptides, signal anchors, or other short N-terminal segments preceding/overlapping the first hydrophobic helix, and also include disordered N-terminal tails in soluble proteins. | 2.92 | 4 |
| #9586 | NAC pre-core regulatory IDRs | Leading subregion of plant NAC-family DNA-binding domains, corresponding to the pre-core/flexible/dimerization-leaning portion of the fold that precedes the DNA-contacting residues; the feature targets this internal NAC sub-segment rather than the downstream DNA-contacting core. Also active on related N-terminal regulatory segments of eukaryotic transcription factors and small viral regulatory proteins. | 2.91 | 4 |
| #14895 | Unknown generic feature | Unknown generic feature | 20.69 | 4 |
| #1803 | Unknown generic feature | Unknown generic feature | 19.38 | 4 |
| #9214 | Unknown generic feature | Unknown generic feature | 16.88 | 4 |
| #14534 | Unknown generic feature | Unknown generic feature | 15.84 | 4 |
| #9194 | Unknown generic feature | Unknown generic feature | 12.12 | 4 |
| #9005 | Unknown generic feature | Unknown generic feature | 10.93 | 4 |
Top-K sparse-autoencoder activations on the ESM-C residual stream. Activation is a feature's peak value over its region; Prevalence is the number of residues it fires on; Δ activation is isoform − canonical. Only the top 30 features per category are saved; generic / "unknown" features are hidden until you expand a table. Read more about SAE features →
Clinical variants
Differential region — N-terminal extension (isoform-unique)
| Pos (iso) | AA change | Consequence | Source | Clin. sig. | AF (gnomAD) | Impact | AlphaMissense | ESM-C ΔLLR | Link |
|---|---|---|---|---|---|---|---|---|---|
| 0 | L→L | synonymous_variant | gnomAD | — | 6.85e-07 | — | N/A | — | chr17-48077016-G-A |
| 0 | L→V | missense_variant | gnomAD | — | 2.19e-05 | — | N/A | — | chr17-48077016-G-C |
| 1 | A→A | synonymous_variant | gnomAD | — | 3.70e-05 | — | N/A | 0.00 | chr17-48077011-C-T |
| 1 | A→V | missense_variant | gnomAD | — | 6.85e-07 | — | N/A | -1.75 | chr17-48077012-G-A |
| 1 | A→V | missense_variant | COSMIC | — | — | — | N/A | -1.75 | COSV56682161 |
| 2 | G→G | synonymous_variant | gnomAD | — | 6.85e-07 | — | N/A | 0.00 | chr17-48077008-G-A |
| 2 | G→G | synonymous_variant | COSMIC | — | — | — | N/A | 0.00 | COSV56681939 |
| 3 | T→T | synonymous_variant | gnomAD | — | 6.85e-07 | — | N/A | 0.00 | chr17-48077005-A-G |
8 variants in the differential region.
Shared canonical core — sequence common to canonical and isoform; AlphaMissense applies here
| Pos (iso) | AA change | Consequence | Source | Clin. sig. | AF (gnomAD) | Impact | AlphaMissense | ESM-C ΔLLR | Link |
|---|---|---|---|---|---|---|---|---|---|
| 4 | M→I | missense_variant | gnomAD | — | 1.37e-06 | damaging | — | -10.94 | chr17-48077002-C-G |
| 4 | M→I | missense_variant | gnomAD | — | 6.85e-07 | damaging | — | -10.94 | chr17-48077002-C-T |
| 4 | M→T | missense_variant | gnomAD | — | 6.85e-07 | damaging | — | -11.50 | chr17-48077003-A-G |
| 4 | M→L | missense_variant | gnomAD | — | 6.85e-07 | damaging | — | -11.19 | chr17-48077004-T-G |
| 4 | M→V | missense_variant | COSMIC | — | — | damaging | — | -11.44 | COSV56682591 |
| 5 | G→G | synonymous_variant | gnomAD | — | 6.85e-07 | — | — | 0.00 | chr17-48076999-C-A |
| 5 | G→V | missense_variant | gnomAD | — | 6.85e-07 | damaging | likely_pathogenic (0.88) | -9.25 | chr17-48077000-C-A |
| 5 | G→W | missense_variant | gnomAD | — | 6.85e-07 | damaging | likely_pathogenic (0.97) | -12.31 | chr17-48077001-C-A |
| 6 | — | frameshift_variant | COSMIC | — | — | LoF | — | — | COSV56681712 |
| 6 | — | frameshift_variant | COSMIC | — | — | LoF | — | — | COSV56681499 |
| 7 | K→K | synonymous_variant | gnomAD | — | 2.74e-06 | — | — | 0.00 | chr17-48076993-T-C |
| 8 | Q→Q | synonymous_variant | gnomAD | — | 6.84e-07 | — | — | 0.00 | chr17-48076990-T-C |
| 8 | Q→K | missense_variant | COSMIC | — | — | damaging | ambiguous (0.37) | -10.00 | COSV56682149 |
| 9 | N→K | missense_variant | gnomAD | — | 1.85e-05 | damaging | likely_pathogenic (0.68) | -8.87 | chr17-48076987-G-C |
| 9 | N→K | missense_variant | ClinVar | Uncertain significance | — | damaging | likely_pathogenic (0.68) | -8.87 | ClinVar:3827926 |
| 10 | K→N | missense_variant | gnomAD | — | 6.84e-07 | damaging | likely_pathogenic (0.73) | -9.94 | chr17-48076984-C-G |
| 10 | — | inframe_deletion | gnomAD | — | 6.84e-07 | — | — | — | chr17-48076984-CTTG-C |
| 10 | K→Q | missense_variant | gnomAD | — | 1.37e-06 | damaging | likely_benign (0.24) | -10.25 | chr17-48076986-T-G |
| 11 | K→N | missense_variant | gnomAD | — | 6.84e-07 | damaging | likely_pathogenic (0.84) | -9.87 | chr17-48076981-C-A |
| 11 | K→R | missense_variant | gnomAD | — | 6.84e-07 | damaging | likely_benign (0.12) | -9.06 | chr17-48076982-T-C |
| 12 | K→N | missense_variant | gnomAD | — | 6.84e-07 | damaging | likely_pathogenic (0.58) | -9.00 | chr17-48076978-T-A |
| 12 | — | inframe_deletion | gnomAD | — | 1.37e-06 | — | — | — | chr17-48076978-TTTC-T |
| 12 | K→I | missense_variant | gnomAD | — | 6.84e-07 | damaging | ambiguous (0.56) | -11.69 | chr17-48076979-T-A |
| 12 | K→R | missense_variant | gnomAD | — | 9.58e-06 | damaging | likely_benign (0.12) | -8.94 | chr17-48076979-T-C |
| 12 | K→* | stop_gained | gnomAD | — | 6.84e-07 | LoF | — | — | chr17-48076980-T-A |
| 12 | K→E | missense_variant | gnomAD | — | 6.84e-07 | damaging | ambiguous (0.42) | -9.69 | chr17-48076980-T-C |
| 13 | V→V | synonymous_variant | gnomAD | — | 1.37e-06 | — | — | 0.00 | chr17-48076975-C-T |
| 13 | V→G | missense_variant | gnomAD | — | 6.85e-07 | damaging | likely_benign (0.14) | -8.75 | chr17-48076976-A-C |
| 13 | V→M | missense_variant | gnomAD | — | 7.53e-06 | damaging | likely_benign (0.13) | -8.37 | chr17-48076977-C-T |
| 13 | V→L | missense_variant | COSMIC | — | — | damaging | likely_benign (0.20) | -8.06 | COSV56682280 |
| 16 | V→V | synonymous_variant | gnomAD | — | 1.37e-06 | — | — | 0.00 | chr17-48076966-C-T |
| 16 | V→L | missense_variant | gnomAD | — | 2.05e-06 | damaging | likely_benign (0.30) | -7.78 | chr17-48076968-C-A |
| 16 | V→M | missense_variant | gnomAD | — | 2.74e-06 | damaging | likely_benign (0.24) | -8.68 | chr17-48076968-C-T |
| 17 | L→L | synonymous_variant | gnomAD | — | 2.05e-06 | — | — | 0.00 | chr17-48076963-T-C |
| 18 | E→K | missense_variant | gnomAD | — | 6.84e-07 | damaging | likely_pathogenic (0.83) | -11.12 | chr17-48076962-C-T |
| 19 | E→Q | missense_variant | gnomAD | — | 6.84e-07 | damaging | likely_pathogenic (0.62) | -9.69 | chr17-48076959-C-G |
| 19 | E→K | missense_variant | COSMIC | — | — | damaging | likely_pathogenic (0.83) | -9.56 | COSV56681509 |
| 20 | E→Q | missense_variant | gnomAD | — | 1.37e-06 | damaging | likely_pathogenic (0.61) | -9.81 | chr17-48076956-C-G |
| 21 | E→E | synonymous_variant | gnomAD | — | 1.37e-06 | — | — | 0.00 | chr17-48076951-T-C |
| 21 | E→G | missense_variant | gnomAD | — | 6.84e-07 | damaging | likely_pathogenic (0.77) | -9.56 | chr17-48076952-T-C |
| 22 | E→Q | missense_variant | COSMIC | — | — | damaging | likely_pathogenic (0.95) | -11.25 | COSV99848313 |
| 23 | E→E | synonymous_variant | gnomAD | — | 2.05e-06 | — | — | 0.00 | chr17-48076945-T-C |
| 25 | V→V | synonymous_variant | gnomAD | — | 2.74e-06 | — | — | 0.00 | chr17-48076939-C-T |
| 25 | V→V | synonymous_variant | COSMIC | — | — | — | — | 0.00 | COSV99848141 |
| 26 | V→V | synonymous_variant | gnomAD | — | 6.84e-07 | — | — | 0.00 | chr17-48076936-C-A |
| 26 | V→V | synonymous_variant | gnomAD | — | 6.84e-07 | — | — | 0.00 | chr17-48076936-C-G |
| 27 | E→V | missense_variant | ClinVar | — | — | damaging | likely_pathogenic (1.00) | -11.31 | ClinVar:4423332 |
| 28 | K→N | missense_variant | gnomAD | — | 6.84e-07 | damaging | likely_pathogenic (0.99) | -9.37 | chr17-48076930-T-A |
| 30 | L→L | synonymous_variant | gnomAD | — | 6.84e-06 | — | — | 0.00 | chr17-48076924-G-A |
| 30 | L→I | missense_variant | COSMIC | — | — | damaging | likely_benign (0.28) | -9.56 | COSV56682110 |
| 31 | D→H | missense_variant | gnomAD | — | 1.37e-06 | damaging | likely_pathogenic (0.94) | -7.89 | chr17-48076923-C-G |
| 31 | D→N | missense_variant | gnomAD | — | 1.85e-05 | damaging | likely_pathogenic (0.56) | -4.01 | chr17-48076923-C-T |
| 32 | R→H | missense_variant | gnomAD | — | 2.74e-06 | damaging | likely_pathogenic (0.79) | -7.03 | chr17-48076919-C-T |
| 32 | R→C | missense_variant | gnomAD | — | 1.23e-05 | damaging | likely_pathogenic (0.87) | -7.12 | chr17-48076920-G-A |
| 33 | R→Q | missense_variant | gnomAD | — | 2.05e-06 | damaging | likely_pathogenic (0.99) | -9.25 | chr17-48076916-C-T |
| 33 | R→R | synonymous_variant | gnomAD | — | 6.84e-07 | — | — | 0.00 | chr17-48076917-G-T |
| 33 | R→R | synonymous_variant | COSMIC | — | — | — | — | 0.00 | COSV56682056 |
| 33 | R→* | stop_gained | COSMIC | — | — | LoF | — | — | COSV56681749 |
| 34 | V→V | synonymous_variant | gnomAD | — | 2.05e-06 | — | — | 0.00 | chr17-48076912-C-T |
| 34 | V→A | missense_variant | gnomAD | — | 1.37e-06 | damaging | likely_pathogenic (0.92) | -7.43 | chr17-48076913-A-G |
| 35 | V→V | synonymous_variant | gnomAD | — | 1.37e-06 | — | — | 0.00 | chr17-48076909-T-C |
| 35 | V→L | missense_variant | gnomAD | — | 6.84e-07 | damaging | likely_pathogenic (0.76) | -6.96 | chr17-48076911-C-A |
| 36 | K→K | synonymous_variant | gnomAD | — | 6.84e-07 | — | — | 0.00 | chr17-48076906-C-T |
| 37 | G→G | synonymous_variant | gnomAD | — | 3.42e-06 | — | — | 0.00 | chr17-48076903-G-A |
| 37 | G→S | missense_variant | gnomAD | — | 6.84e-07 | damaging | likely_pathogenic (0.96) | -7.97 | chr17-48076905-C-T |
| 37 | G→S | missense_variant | COSMIC | — | — | damaging | likely_pathogenic (0.96) | -7.97 | COSV99847936 |
| 38 | K→R | missense_variant | gnomAD | — | 2.05e-06 | — | likely_benign (0.09) | -5.18 | chr17-48076901-T-C |
| 40 | E→D | missense_variant | gnomAD | — | 6.85e-07 | damaging | likely_pathogenic (1.00) | -7.72 | chr17-48076894-C-G |
| 40 | E→K | missense_variant | COSMIC | — | — | damaging | likely_pathogenic (1.00) | -10.19 | COSV56682856 |
| 41 | Y→Y | synonymous_variant | gnomAD | — | 1.71e-05 | — | — | 0.00 | chr17-48076891-G-A |
| 42 | L→L | synonymous_variant | gnomAD | — | 3.43e-06 | — | — | 0.00 | chr17-48076888-G-A |
| 42 | L→R | missense_variant | gnomAD | — | 6.85e-07 | damaging | likely_pathogenic (0.84) | -9.56 | chr17-48076889-A-C |
| 42 | L→F | missense_variant | COSMIC | — | — | — | ambiguous (0.35) | -6.53 | COSV56681814 |
| 43 | L→L | synonymous_variant | gnomAD | — | 6.85e-07 | — | — | 0.00 | chr17-48076885-T-G |
| 44 | K→K | synonymous_variant | gnomAD | — | 6.85e-07 | — | — | 0.00 | chr17-48076882-C-T |
| 44 | K→E | missense_variant | gnomAD | — | 6.85e-07 | damaging | likely_pathogenic (1.00) | -11.19 | chr17-48076884-T-C |
| 47 | G→G | synonymous_variant | gnomAD | — | 4.39e-05 | — | — | 0.00 | chr17-48076873-T-C |
| 48 | F→F | synonymous_variant | gnomAD | — | 6.86e-07 | — | — | 0.00 | chr17-48076870-G-A |
| 48 | F→L | missense_variant | COSMIC | — | — | damaging | likely_pathogenic (1.00) | -10.37 | COSV56682209 |
| 51 | E→K | missense_variant | gnomAD | — | 7.10e-07 | damaging | likely_pathogenic (0.84) | -9.50 | chr17-48076177-C-T |
| 51 | E→Q | missense_variant | ClinVar | — | — | damaging | likely_pathogenic (0.60) | -11.12 | ClinVar:4423331 |
| 51 | E→Q | missense_variant | COSMIC | — | — | damaging | likely_pathogenic (0.60) | -11.12 | COSV56681530 |
| 52 | D→H | missense_variant | gnomAD | — | 7.06e-07 | damaging | likely_pathogenic (0.97) | -12.12 | chr17-48076174-C-G |
| 53 | N→D | missense_variant | gnomAD | — | 1.41e-06 | damaging | likely_pathogenic (0.97) | -10.37 | chr17-48076171-T-C |
| 53 | N→S | missense_variant | COSMIC | — | — | damaging | likely_pathogenic (0.59) | -9.00 | COSV56681968 |
| 56 | E→D | missense_variant | COSMIC | — | — | damaging | likely_pathogenic (1.00) | -10.56 | COSV99848274 |
| 58 | E→D | missense_variant | COSMIC | — | — | damaging | likely_pathogenic (0.96) | -8.69 | COSV99848151 |
| 60 | N→N | synonymous_variant | gnomAD | — | 1.38e-06 | — | — | 0.00 | chr17-48076148-G-A |
| 61 | L→L | synonymous_variant | gnomAD | — | 1.38e-06 | — | — | 0.00 | chr17-48076147-G-A |
| 64 | P→P | synonymous_variant | gnomAD | — | 1.51e-05 | — | — | 0.00 | chr17-48076136-G-A |
| 64 | P→P | synonymous_variant | gnomAD | — | 6.88e-07 | — | — | 0.00 | chr17-48076136-G-C |
| 64 | P→P | synonymous_variant | COSMIC | — | — | — | — | 0.00 | COSV99847948 |
| 65 | D→N | missense_variant | COSMIC | — | — | damaging | likely_pathogenic (0.96) | -8.31 | COSV99847954 |
| 66 | L→L | synonymous_variant | gnomAD | — | 1.23e-05 | — | — | 0.00 | chr17-48076130-G-A |
| 68 | A→A | synonymous_variant | gnomAD | — | 5.48e-06 | — | — | 0.00 | chr17-48076124-A-G |
| 68 | A→A | synonymous_variant | gnomAD | — | 6.85e-07 | — | — | 0.00 | chr17-48076124-A-T |
| 68 | A→G | missense_variant | gnomAD | — | 1.37e-06 | damaging | ambiguous (0.38) | -9.94 | chr17-48076125-G-C |
| 71 | L→L | synonymous_variant | gnomAD | — | 1.37e-06 | — | — | 0.00 | chr17-48076115-C-T |
| 71 | L→L | synonymous_variant | gnomAD | — | 6.85e-07 | — | — | 0.00 | chr17-48076117-G-A |
| 71 | L→L | synonymous_variant | COSMIC | — | — | — | — | 0.00 | COSV105045650 |
| 72 | Q→* | stop_gained | COSMIC | — | — | LoF | — | — | COSV99848158 |
| 73 | S→L | missense_variant | gnomAD | — | 2.74e-06 | damaging | likely_benign (0.27) | -9.37 | chr17-48076110-G-A |
| 73 | S→L | missense_variant | ClinVar | Uncertain significance | — | damaging | likely_benign (0.27) | -9.37 | ClinVar:4648788 |
| 74 | Q→R | missense_variant | gnomAD | — | 6.85e-07 | — | likely_benign (0.33) | -7.47 | chr17-48076107-T-C |
| 74 | Q→* | stop_gained | COSMIC | — | — | LoF | — | — | COSV99848270 |
| 75 | K→R | missense_variant | gnomAD | — | 6.85e-07 | — | likely_benign (0.12) | -6.44 | chr17-48076104-T-C |
| 75 | K→Q | missense_variant | ClinVar | Uncertain significance | — | damaging | likely_benign (0.33) | -8.75 | ClinVar:4531663 |
| 76 | T→A | missense_variant | gnomAD | — | 6.85e-07 | — | likely_benign (0.06) | -6.06 | chr17-48076102-T-C |
| 78 | H→P | missense_variant | gnomAD | — | 4.79e-06 | — | likely_benign (0.11) | -7.21 | chr17-48076095-T-G |
| 80 | T→R | missense_variant | gnomAD | — | 2.74e-06 | damaging | likely_benign (0.23) | -7.90 | chr17-48076089-G-C |
| 81 | D→G | missense_variant | gnomAD | — | 6.84e-07 | — | likely_benign (0.10) | -6.27 | chr17-48076086-T-C |
| 81 | D→N | missense_variant | gnomAD | — | 6.84e-07 | — | likely_benign (0.10) | -6.15 | chr17-48076087-C-T |
| 81 | D→G | missense_variant | ClinVar | Uncertain significance | — | — | likely_benign (0.10) | -6.27 | ClinVar:4220041 |
| 81 | D→N | missense_variant | ClinVar | Uncertain significance | — | — | likely_benign (0.10) | -6.15 | ClinVar:4220044 |
| 82 | K→K | synonymous_variant | gnomAD | — | 4.04e-05 | — | — | 0.00 | chr17-48076082-T-C |
| 82 | K→I | missense_variant | gnomAD | — | 6.84e-07 | damaging | ambiguous (0.43) | -10.75 | chr17-48076083-T-A |
| 82 | K→I | missense_variant | ClinVar | Uncertain significance | — | damaging | ambiguous (0.43) | -10.75 | ClinVar:4220042 |
| 83 | S→S | synonymous_variant | gnomAD | — | 4.79e-06 | — | — | 0.00 | chr17-48076079-T-G |
| 84 | E→K | missense_variant | COSMIC | — | — | damaging | likely_benign (0.26) | -8.68 | COSV99848247 |
| 85 | G→R | missense_variant | gnomAD | — | 6.84e-07 | damaging | likely_pathogenic (0.58) | -8.76 | chr17-48076075-C-G |
| 85 | G→R | missense_variant | COSMIC | — | — | damaging | likely_pathogenic (0.58) | -8.76 | COSV56682528 |
| 88 | R→H | missense_variant | gnomAD | — | 7.53e-06 | damaging | likely_pathogenic (0.89) | -9.31 | chr17-48076065-C-T |
| 88 | R→C | missense_variant | gnomAD | — | 6.16e-06 | damaging | likely_pathogenic (0.96) | -9.31 | chr17-48076066-G-A |
| 88 | R→H | missense_variant | ClinVar | Uncertain significance | — | damaging | likely_pathogenic (0.89) | -9.31 | ClinVar:2290145 |
| 88 | R→H | missense_variant | COSMIC | — | — | damaging | likely_pathogenic (0.89) | -9.31 | COSV56682901 |
| 88 | R→C | missense_variant | COSMIC | — | — | damaging | likely_pathogenic (0.96) | -9.31 | COSV56682920 |
| 89 | K→R | missense_variant | gnomAD | — | 8.90e-06 | damaging | likely_benign (0.11) | -7.87 | chr17-48076062-T-C |
| 90 | A→T | missense_variant | gnomAD | — | 1.37e-06 | — | likely_benign (0.08) | -6.34 | chr17-48076060-C-T |
| 90 | A→T | missense_variant | ClinVar | Uncertain significance | — | — | likely_benign (0.08) | -6.34 | ClinVar:4220043 |
| 91 | D→G | missense_variant | gnomAD | — | 6.85e-07 | damaging | likely_benign (0.15) | -8.36 | chr17-48076056-T-C |
| 92 | S→A | missense_variant | gnomAD | — | 6.85e-07 | damaging | likely_benign (0.06) | -9.87 | chr17-48076054-A-C |
| 92 | S→T | missense_variant | gnomAD | — | 1.37e-06 | — | likely_benign (0.05) | -7.31 | chr17-48076054-A-T |
| 96 | D→V | missense_variant | gnomAD | — | 1.37e-06 | damaging | likely_benign (0.13) | -9.56 | chr17-48076041-T-A |
| 96 | D→V | missense_variant | ClinVar | Uncertain significance | — | damaging | likely_benign (0.13) | -9.56 | ClinVar:2307098 |
| 96 | D→N | missense_variant | COSMIC | — | — | damaging | likely_benign (0.11) | -7.93 | COSV56682563 |
| 97 | K→K | synonymous_variant | gnomAD | — | 1.37e-06 | — | — | 0.00 | chr17-48076037-C-T |
| 97 | K→R | missense_variant | gnomAD | — | 2.05e-06 | — | likely_benign (0.09) | -4.90 | chr17-48076038-T-C |
| 97 | K→E | missense_variant | gnomAD | — | 6.85e-07 | damaging | likely_benign (0.18) | -11.05 | chr17-48076039-T-C |
| 97 | K→K | synonymous_variant | COSMIC | — | — | — | — | 0.00 | COSV56682520 |
| 97 | K→N | missense_variant | COSMIC | — | — | damaging | likely_benign (0.29) | -9.30 | COSV99848170 |
| 98 | G→G | synonymous_variant | gnomAD | — | 2.06e-06 | — | — | 0.00 | chr17-48076034-T-C |
| 98 | — | frameshift_variant | gnomAD | — | 1.37e-06 | LoF | — | — | chr17-48076035-CCCTT-C |
| 98 | G→R | missense_variant | gnomAD | — | 2.74e-06 | damaging | ambiguous (0.54) | -7.55 | chr17-48076036-C-T |
| 99 | E→E | synonymous_variant | gnomAD | — | 6.86e-07 | — | — | 0.00 | chr17-48076031-C-T |
| 99 | E→Q | missense_variant | COSMIC | — | — | damaging | likely_benign (0.23) | -8.62 | COSV99848203 |
| 100 | E→D | missense_variant | gnomAD | — | 6.86e-07 | — | likely_benign (0.05) | -5.84 | chr17-48076028-C-G |
| 101 | S→S | synonymous_variant | gnomAD | — | 4.81e-06 | — | — | 0.00 | chr17-48076025-G-A |
| 101 | S→G | missense_variant | gnomAD | — | 6.86e-07 | — | likely_benign (0.06) | -5.94 | chr17-48076027-T-C |
| 101 | S→S | synonymous_variant | COSMIC | — | — | — | — | 0.00 | COSV99848275 |
| 103 | P→P | synonymous_variant | gnomAD | — | 2.07e-06 | — | — | 0.00 | chr17-48076019-T-C |
| 103 | P→L | missense_variant | gnomAD | — | 1.38e-06 | — | likely_benign (0.09) | -6.30 | chr17-48076020-G-A |
| 105 | K→N | missense_variant | gnomAD | — | 6.90e-07 | damaging | likely_pathogenic (0.93) | -10.37 | chr17-48076013-C-A |
| 105 | K→K | synonymous_variant | gnomAD | — | 1.38e-06 | — | — | 0.00 | chr17-48076013-C-T |
| 106 | K→N | missense_variant | gnomAD | — | 6.91e-07 | damaging | likely_pathogenic (0.91) | -10.25 | chr17-48076010-C-G |
| 106 | K→R | missense_variant | gnomAD | — | 1.38e-06 | — | likely_benign (0.09) | -5.37 | chr17-48076011-T-C |
| 106 | K→N | missense_variant | COSMIC | — | — | damaging | likely_pathogenic (0.91) | -10.25 | COSV56681336 |
| 106 | K→N | missense_variant | COSMIC | — | — | damaging | likely_pathogenic (0.91) | -10.25 | COSV99848297 |
| 107 | — | inframe_deletion | gnomAD | — | 2.77e-06 | — | — | — | chr17-48076007-TTTC-T |
| 107 | K→R | missense_variant | gnomAD | — | 6.92e-07 | — | likely_benign (0.09) | -6.75 | chr17-48076008-T-C |
| 108 | E→G | missense_variant | gnomAD | — | 6.92e-07 | damaging | likely_benign (0.33) | -8.81 | chr17-48076005-T-C |
| 109 | E→V | missense_variant | gnomAD | — | 1.39e-06 | damaging | likely_benign (0.30) | -8.24 | chr17-48076002-T-A |
| 109 | E→Q | missense_variant | gnomAD | — | 6.94e-07 | damaging | ambiguous (0.43) | -8.99 | chr17-48076003-C-G |
| 110 | S→S | synonymous_variant | gnomAD | — | 6.85e-07 | — | — | 0.00 | chr17-48075098-T-A |
| 110 | S→L | missense_variant | gnomAD | — | 6.85e-07 | — | likely_benign (0.08) | -6.09 | chr17-48075099-G-A |
| 111 | E→A | missense_variant | gnomAD | — | 6.85e-07 | damaging | ambiguous (0.50) | -9.43 | chr17-48075096-T-G |
| 113 | P→S | missense_variant | gnomAD | — | 6.84e-07 | — | ambiguous (0.53) | -5.76 | chr17-48075091-G-A |
| 113 | P→Q | missense_variant | COSMIC | — | — | damaging | likely_pathogenic (0.61) | -7.20 | COSV106082017 |
| 113 | P→S | missense_variant | COSMIC | — | — | — | ambiguous (0.53) | -5.76 | COSV56682997 |
| 114 | R→Q | missense_variant | gnomAD | — | 6.84e-07 | damaging | likely_pathogenic (0.92) | -8.69 | chr17-48075087-C-T |
| 114 | R→* | stop_gained | gnomAD | — | 2.05e-06 | LoF | — | — | chr17-48075088-G-A |
| 114 | R→* | stop_gained | COSMIC | — | — | LoF | — | — | COSV56682269 |
| 116 | F→S | missense_variant | gnomAD | — | 6.84e-07 | damaging | likely_pathogenic (1.00) | -10.44 | chr17-48075081-A-G |
| 117 | A→A | synonymous_variant | gnomAD | — | 1.37e-06 | — | — | 0.00 | chr17-48075077-A-C |
| 117 | A→A | synonymous_variant | gnomAD | — | 6.84e-07 | — | — | 0.00 | chr17-48075077-A-G |
| 118 | R→Q | missense_variant | gnomAD | — | 6.84e-07 | damaging | likely_pathogenic (0.99) | -8.12 | chr17-48075075-C-T |
| 118 | R→* | stop_gained | gnomAD | — | 6.84e-07 | LoF | — | — | chr17-48075076-G-A |
| 118 | R→R | synonymous_variant | COSMIC | — | — | — | — | 0.00 | COSV107309107 |
| 118 | R→L | missense_variant | COSMIC | — | — | damaging | likely_pathogenic (1.00) | -10.37 | COSV99847942 |
| 118 | R→Q | missense_variant | COSMIC | — | — | damaging | likely_pathogenic (0.99) | -8.12 | COSV56683114 |
| 118 | R→* | stop_gained | COSMIC | — | — | LoF | — | — | COSV104387334 |
| 119 | G→D | missense_variant | COSMIC | — | — | damaging | likely_pathogenic (0.89) | -9.62 | COSV105045685 |
| 121 | E→K | missense_variant | gnomAD | — | 6.84e-07 | damaging | likely_pathogenic (0.87) | -7.54 | chr17-48075067-C-T |
| 121 | E→E | synonymous_variant | COSMIC | — | — | — | — | 0.00 | COSV56681442 |
| 122 | P→P | synonymous_variant | gnomAD | — | 1.97e-04 | — | — | 0.00 | chr17-48075062-C-T |
| 122 | P→L | missense_variant | gnomAD | — | 7.52e-06 | damaging | likely_pathogenic (1.00) | -10.81 | chr17-48075063-G-A |
| 122 | P→P | synonymous_variant | COSMIC | — | — | — | — | 0.00 | COSV99848187 |
| 122 | P→Q | missense_variant | COSMIC | — | — | damaging | likely_pathogenic (1.00) | -12.12 | COSV56682499 |
| 122 | P→S | missense_variant | COSMIC | — | — | damaging | likely_pathogenic (1.00) | -9.69 | COSV99848079 |
| 123 | E→D | missense_variant | gnomAD | — | 1.37e-06 | damaging | likely_pathogenic (0.92) | -9.06 | chr17-48075059-C-A |
| 123 | E→E | synonymous_variant | gnomAD | — | 1.37e-06 | — | — | 0.00 | chr17-48075059-C-T |
| 123 | E→D | missense_variant | ClinVar | Uncertain significance | — | damaging | likely_pathogenic (0.92) | -9.06 | ClinVar:4648787 |
| 123 | E→K | missense_variant | COSMIC | — | — | damaging | likely_pathogenic (1.00) | -12.12 | COSV56681862 |
| 124 | R→R | synonymous_variant | gnomAD | — | 6.84e-07 | — | — | 0.00 | chr17-48075056-C-T |
| 124 | R→Q | missense_variant | COSMIC | — | — | damaging | likely_pathogenic (0.98) | -7.87 | COSV56682101 |
| 124 | R→L | missense_variant | COSMIC | — | — | damaging | likely_pathogenic (0.99) | -10.12 | COSV99848174 |
| 125 | I→V | missense_variant | gnomAD | — | 6.84e-07 | damaging | likely_pathogenic (0.96) | -7.69 | chr17-48075055-T-C |
| 127 | G→A | missense_variant | COSMIC | — | — | damaging | likely_pathogenic (1.00) | -12.44 | COSV56682437 |
| 129 | T→T | synonymous_variant | gnomAD | — | 2.74e-06 | — | — | 0.00 | chr17-48075041-T-C |
| 131 | S→S | synonymous_variant | gnomAD | — | 6.84e-07 | — | — | 0.00 | chr17-48075035-G-A |
| 132 | S→S | synonymous_variant | COSMIC | — | — | — | — | 0.00 | COSV56681832 |
| 134 | E→Q | missense_variant | gnomAD | — | 6.84e-07 | damaging | likely_pathogenic (0.92) | -10.00 | chr17-48075028-C-G |
| 135 | L→L | synonymous_variant | gnomAD | — | 9.59e-06 | — | — | 0.00 | chr17-48075023-G-A |
| 136 | M→L | missense_variant | gnomAD | — | 6.85e-07 | damaging | likely_pathogenic (0.91) | -10.25 | chr17-48075022-T-A |
| 136 | M→V | missense_variant | gnomAD | — | 6.85e-07 | damaging | likely_pathogenic (0.99) | -11.50 | chr17-48075022-T-C |
| 136 | — | inframe_deletion | COSMIC | — | — | — | — | — | COSV56681204 |
| 137 | F→F | synonymous_variant | gnomAD | — | 6.85e-07 | — | — | 0.00 | chr17-48075017-G-A |
| 138 | L→P | missense_variant | gnomAD | — | 6.85e-07 | damaging | likely_pathogenic (1.00) | -13.06 | chr17-48075015-A-G |
| 138 | L→L | synonymous_variant | COSMIC | — | — | — | — | 0.00 | COSV99848148 |
| 138 | L→L | synonymous_variant | COSMIC | — | — | — | — | 0.00 | COSV56681918 |
| 139 | M→I | missense_variant | gnomAD | — | 6.86e-07 | damaging | likely_pathogenic (0.97) | -8.12 | chr17-48075011-C-A |
| 139 | M→T | missense_variant | gnomAD | — | 6.86e-07 | damaging | likely_pathogenic (1.00) | -11.81 | chr17-48075012-A-G |
| 140 | K→K | synonymous_variant | gnomAD | — | 6.86e-07 | — | — | 0.00 | chr17-48075008-T-C |
| 140 | K→T | missense_variant | gnomAD | — | 6.86e-07 | damaging | likely_pathogenic (1.00) | -11.87 | chr17-48075009-T-G |
| 142 | K→R | missense_variant | gnomAD | — | 8.28e-06 | — | likely_benign (0.14) | -6.94 | chr17-48071577-T-C |
| 143 | N→K | missense_variant | COSMIC | — | — | damaging | likely_pathogenic (0.89) | -8.94 | COSV99848229 |
| 144 | S→F | missense_variant | COSMIC | — | — | damaging | likely_pathogenic (0.99) | -10.37 | COSV56681364 |
| 147 | A→A | synonymous_variant | gnomAD | — | 6.87e-07 | — | — | 0.00 | chr17-48071561-A-G |
| 148 | D→D | synonymous_variant | gnomAD | — | 6.87e-07 | — | — | 0.00 | chr17-48071558-G-A |
| 150 | V→L | missense_variant | gnomAD | — | 6.86e-07 | damaging | likely_pathogenic (1.00) | -11.25 | chr17-48071554-C-G |
| 152 | A→A | synonymous_variant | gnomAD | — | 6.85e-07 | — | — | 0.00 | chr17-48071546-G-A |
| 152 | A→D | missense_variant | gnomAD | — | 6.85e-07 | damaging | likely_pathogenic (1.00) | -12.00 | chr17-48071547-G-T |
| 152 | A→V | missense_variant | COSMIC | — | — | damaging | likely_pathogenic (1.00) | -10.62 | COSV99848117 |
| 154 | E→* | stop_gained | gnomAD | — | 6.85e-07 | LoF | — | — | chr17-48071542-C-A |
| 154 | E→K | missense_variant | COSMIC | — | — | damaging | likely_pathogenic (0.99) | -10.94 | COSV56681808 |
| 156 | N→S | missense_variant | gnomAD | — | 6.84e-07 | damaging | likely_pathogenic (0.91) | -12.44 | chr17-48071535-T-C |
| 157 | V→F | missense_variant | gnomAD | — | 2.74e-06 | damaging | likely_pathogenic (0.57) | -8.73 | chr17-48071533-C-A |
| 157 | V→V | synonymous_variant | COSMIC | — | — | — | — | 0.00 | COSV56681933 |
| 158 | K→K | synonymous_variant | COSMIC | — | — | — | — | 0.00 | COSV56682664 |
| 159 | C→C | synonymous_variant | gnomAD | — | 1.37e-06 | — | — | 0.00 | chr17-48071525-G-A |
| 159 | C→S | missense_variant | gnomAD | — | 6.84e-07 | damaging | likely_pathogenic (1.00) | -5.97 | chr17-48071526-C-G |
| 160 | P→L | missense_variant | gnomAD | — | 6.84e-07 | damaging | likely_pathogenic (1.00) | -10.62 | chr17-48071523-G-A |
| 161 | Q→Q | synonymous_variant | gnomAD | — | 6.84e-07 | — | — | 0.00 | chr17-48071519-C-T |
| 161 | Q→* | stop_gained | gnomAD | — | 6.84e-07 | LoF | — | — | chr17-48071521-G-A |
| 162 | V→F | missense_variant | gnomAD | — | 2.05e-06 | damaging | likely_pathogenic (0.99) | -12.05 | chr17-48071518-C-A |
| 164 | I→I | synonymous_variant | COSMIC | — | — | — | — | 0.00 | COSV56681722 |
| 164 | I→T | missense_variant | COSMIC | — | — | damaging | likely_pathogenic (1.00) | -11.06 | COSV56682067 |
| 165 | S→S | synonymous_variant | gnomAD | — | 2.05e-06 | — | — | 0.00 | chr17-48071507-G-T |
| 166 | F→F | synonymous_variant | gnomAD | — | 6.84e-07 | — | — | 0.00 | chr17-48071504-G-A |
| 167 | Y→Y | synonymous_variant | gnomAD | — | 1.37e-06 | — | — | 0.00 | chr17-48071501-A-G |
| 167 | Y→C | missense_variant | gnomAD | — | 6.84e-07 | damaging | likely_pathogenic (1.00) | -11.94 | chr17-48071502-T-C |
| 170 | R→K | missense_variant | COSMIC | — | — | damaging | likely_pathogenic (0.99) | -11.06 | COSV56682552 |
| 172 | T→T | synonymous_variant | gnomAD | — | 4.11e-06 | — | — | 0.00 | chr17-48071486-C-T |
| 172 | T→K | missense_variant | COSMIC | — | — | damaging | likely_pathogenic (0.99) | -13.94 | COSV56681836 |
| 174 | H→H | synonymous_variant | gnomAD | — | 6.84e-07 | — | — | 0.00 | chr17-48071480-A-G |
| 175 | S→S | synonymous_variant | gnomAD | — | 6.84e-07 | — | — | 0.00 | chr17-48071477-G-A |
| 175 | S→F | missense_variant | gnomAD | — | 6.84e-07 | damaging | likely_pathogenic (0.97) | -11.12 | chr17-48071478-G-A |
| 175 | S→A | missense_variant | gnomAD | — | 6.84e-07 | damaging | likely_benign (0.27) | -9.06 | chr17-48071479-A-C |
| 175 | S→P | missense_variant | gnomAD | — | 6.84e-07 | damaging | likely_pathogenic (0.98) | -11.81 | chr17-48071479-A-G |
| 175 | S→T | missense_variant | gnomAD | — | 6.84e-07 | damaging | ambiguous (0.44) | -10.19 | chr17-48071479-A-T |
| 176 | Y→Y | synonymous_variant | gnomAD | — | 2.94e-05 | — | — | 0.00 | chr17-48071474-G-A |
| 176 | Y→* | stop_gained | gnomAD | — | 6.85e-07 | LoF | — | — | chr17-48071474-G-T |
| 177 | P→S | missense_variant | gnomAD | — | 6.85e-07 | damaging | likely_pathogenic (0.58) | -8.75 | chr17-48071473-G-A |
| 177 | P→S | missense_variant | COSMIC | — | — | damaging | likely_pathogenic (0.58) | -8.75 | COSV108798400 |
| 178 | S→S | synonymous_variant | gnomAD | — | 4.11e-06 | — | — | 0.00 | chr17-48071468-C-T |
| 178 | S→L | missense_variant | gnomAD | — | 2.05e-06 | damaging | likely_benign (0.21) | -8.42 | chr17-48071469-G-A |
| 178 | S→L | missense_variant | ClinVar | Uncertain significance | — | damaging | likely_benign (0.21) | -8.42 | ClinVar:3138046 |
| 178 | S→S | synonymous_variant | COSMIC | — | — | — | — | 0.00 | COSV56681272 |
| 178 | S→* | stop_gained | COSMIC | — | — | LoF | — | — | COSV99848014 |
| 179 | E→K | missense_variant | COSMIC | — | — | damaging | likely_pathogenic (0.75) | -12.00 | COSV104555665 |
| 180 | D→E | missense_variant | gnomAD | — | 4.80e-06 | — | likely_benign (0.09) | -5.24 | chr17-48071462-A-C |
| 180 | D→D | synonymous_variant | gnomAD | — | 6.85e-07 | — | — | 0.00 | chr17-48071462-A-G |
| 180 | D→Y | missense_variant | gnomAD | — | 6.85e-07 | damaging | likely_pathogenic (0.71) | -11.56 | chr17-48071464-C-A |
| 180 | D→N | missense_variant | gnomAD | — | 1.37e-06 | damaging | likely_benign (0.31) | -10.18 | chr17-48071464-C-T |
| 180 | D→E | missense_variant | ClinVar | Uncertain significance | — | — | likely_benign (0.09) | -5.24 | ClinVar:2517719 |
| 182 | — | inframe_deletion | gnomAD | — | 2.06e-06 | — | — | — | chr17-48071456-GTCA-G |
| 182 | D→N | missense_variant | COSMIC | — | — | damaging | likely_benign (0.19) | -8.62 | COSV99848218 |
| 183 | K→N | missense_variant | COSMIC | — | — | damaging | likely_pathogenic (0.72) | -9.62 | COSV99848111 |
| 183 | — | frameshift_variant | COSMIC | — | — | LoF | — | — | COSV56682928 |
| 184 | K→* | stop_gained | gnomAD | — | 6.85e-07 | LoF | — | — | chr17-48071452-T-A |
| 185 | D→V | missense_variant | gnomAD | — | 6.86e-07 | damaging | ambiguous (0.45) | -9.47 | chr17-48071448-T-A |
| 185 | — | frameshift_variant | gnomAD | — | 6.86e-07 | LoF | — | — | chr17-48071449-C-CT |
| 185 | — | frameshift_variant | gnomAD | — | 6.86e-07 | LoF | — | — | chr17-48071449-CTTTT-C |
| 185 | D→H | missense_variant | COSMIC | — | — | damaging | likely_pathogenic (0.73) | -10.72 | COSV56682228 |
| 185 | D→Y | missense_variant | COSMIC | — | — | damaging | likely_pathogenic (0.65) | -9.97 | COSV56681987 |
| 186 | — | frameshift_variant | gnomAD | — | 6.86e-07 | LoF | — | — | chr17-48071445-TC-T |
| 186 | D→Y | missense_variant | gnomAD | — | 1.37e-06 | damaging | likely_pathogenic (0.68) | -10.43 | chr17-48071446-C-A |
| 186 | D→N | missense_variant | gnomAD | — | 6.86e-07 | — | likely_benign (0.27) | -7.43 | chr17-48071446-C-T |
| 186 | — | frameshift_variant | gnomAD | — | 6.86e-07 | LoF | — | — | chr17-48071446-CA-C |
| 188 | N→K | missense_variant | gnomAD | — | 6.87e-07 | damaging | likely_pathogenic (0.68) | -7.21 | chr17-48071438-G-C |
| 188 | — | frameshift_variant | gnomAD | — | 1.37e-06 | LoF | — | — | chr17-48071439-TTCTTG-T |
| 188 | N→K | missense_variant | COSMIC | — | — | damaging | likely_pathogenic (0.68) | -7.21 | COSV56682095 |
281 variants in the shared canonical core.
LoF = frameshift / stop-gain / splice-disrupting — inherently loss-of-function, flagged by consequence (AlphaMissense and ESM-C score only missense/substitutions, so they are blank here by design, not by absence of impact). AlphaMissense is computed in the canonical reading frame, so it scores the shared core but reads N/A across an isoform-unique extension — use ESM-C ΔLLR there.