CBX1
TRUNCATED 163 aa (canonical 185 aa) · UniProt P83916 · CDLMPS
chr17:48076938:-:GTG:ENST00000225603.9
AI summary N-terminal 23-aa loss leaves the chromodomain, localization, and biophysics essentially untouched despite the region's conservation.
The truncation removes a 23-residue N-terminal tail entirely outside the chromo and chromo-shadow domains (which begin at residue ~20 in the canonical numbering and are unaffected), and no real InterPro domain, localization signal, or whole-protein biophysical property changes as a result. DeepLoc calls both canonical and isoform Nucleus with matching high confidence and an intact NLS, so there is no localization conflict; the moderate pLDDT of the removed tail (0.74) suggests weak local structure but it is unintegrated with the rest of the fold (PAE ~25 Å) and carries no annotated function.
CBX1/HP1-beta's known activities — H3K9me3 reading via the chromodomain, SUV39H1/PRC2 partnership, and CK2/Chk2-regulated chromatin mobilization — all map to the chromo and chromo-shadow domains and their phosphorylation sites, none of which lie in this removed N-terminal stretch. With domain architecture, nuclear localization, and biophysical character all preserved, this truncation has no clear mechanistic bearing on any of CBX1's documented functions.
No mass-spec validation of the truncated peptide and the shared-region structural model is low-confidence (pTM <0.5, shared pLDDT 0.63), so structural claims about the retained core are not usable evidence either way.
Folding
Evidence — click any tile for the differential-region detail
LLM reasoning
How well is the isoform-unique region conserved at the amino-acid level across primates (mean percent identity)? High AA identity in close relatives argues the alternative protein is real and translated, not a sequencing or annotation artifact. (Reading-frame intactness is shown as context.)
Comparison
Sequence conservation · primate
mean AA identity over the differential ORF is the score basis; frame intactness across the clade is context
| Conservation metric | Canonical ORF | Differential ORF |
|---|---|---|
| Mean AA identity | 100% | 100% |
| Frame intact (fraction of species) | 84% | 100% |
| Species aligned | 25 | 25 |
| Species frame-intact | 21 | 25 |
| Start codon conserved | 100% | 100% |
| Deepest intact species | Propithecus_coquereli | Microcebus_murinus |
| Phylo depth (MRCA) | 7 | 7 |
The same amino-acid identity test across mammals. Conservation over deeper evolutionary distance is stronger evidence the ORF is under selection to be translated. (Reading-frame intactness is shown as context.)
Comparison
Sequence conservation · mammalian
mean AA identity over the differential ORF is the score basis; frame intactness across the clade is context
| Conservation metric | Canonical ORF | Differential ORF |
|---|---|---|
| Mean AA identity | 98% | 100% |
| Frame intact (fraction of species) | 15% | 100% |
| Species aligned | 20 | 20 |
| Species frame-intact | 3 | 20 |
| Start codon conserved | 100% | 100% |
| Deepest intact species | Callithrix_jacchus | Loxodonta_africana |
| Phylo depth (MRCA) | 6 | 12 |
phyloP / phastCons measure per-base evolutionary constraint. A high absolute phyloP over the differential region means the sequence itself is under strong purifying selection — evidence it is coding and does something. (Enrichment over the shared core is shown as context only.)
Comparison
Per-base conservation · differential region
absolute mean phyloP over the differential region is the score basis (high = strong purifying selection); the shared column and enrichment are context, not the claim
| Track | Differential | Shared | vs shared |
|---|---|---|---|
| phyloP mean | 4.95 | 4.94 | 1 |
| phastCons mean | 0.966 | 0.913 | — |
Start site · canonical vs isoform
initiation context + per-base conservation at each start codon
| Property | Canonical | Isoform |
|---|---|---|
| Start codon | ATG | GTG |
| Kozak context (−9..+4) | GCGGGCACTATGG | GAATATGTGGTGG |
| phyloP at start codon | 6.43 | 4.52 |
| phastCons at start codon | 1 | 1 |
| phyloP over Kozak window | 6.31 | 5.02 |
| phastCons over Kozak window | 1 | 0.994 |
| Kozak mismatch — full consensus | 3 | 8 |
| Kozak window GC content | 0.692 | 0.462 |
LLM reasoning
Is the alternative start used in more than one cell line? Reproducible initiation across independent samples argues against a one-off ribosome artifact.
Comparison
Per-cell-line usage · canonical vs isoform
Fisher q = 0.000687
| Cell line | Canonical | Isoform | p-value |
|---|---|---|---|
| HeLa | — | 3.65 | 0.000687 |
| RPE1 Async | — | 0.676 | 0.00865 |
How efficiently ribosomes initiate at this start (TIS) relative to background, per cell line. Strong, reproducible initiation supports a genuine translation event.
Comparison
Start-Site Usage · canonical vs isoform
ribosome initiation efficiency at the canonical start vs this alternative start, per cell line
| Cell line | Canonical | Isoform |
|---|---|---|
| HeLa | — | 0.0995 |
| RPE1 Async | — | 0.0161 |
Were tryptic peptides unique to the differential region detected by mass spec? Direct peptide evidence is the strongest proof the isoform protein exists.
Comparison
Peptide Evidence (canonical vs isoform)
isoform-unique peptides are direct evidence the alternative protein exists
| Feature | Canonical | Isoform |
|---|---|---|
| Tryptic peptides (in-silico) | 28 | 1 |
| Validated by mass-spec | 0 | 0 |
| Isoform-unique peptides | — | 1 |
Details
Peptide Evidence (canonical vs isoform)
- peptide MEKVLDR 0–7
LLM reasoning
Do the predicted localization features (DeepLoc prediction, sorting signals, or membrane association) differ between the canonical and the isoform? A protein with changed localization features acts in a different cellular context.
Comparison
Subcellular localization (DeepLoc)
no localization-feature change
| Property | Canonical | Isoform |
|---|---|---|
| Predicted location | Nucleus | Nucleus |
| Sorting signals | Nuclear localization signal | Nuclear localization signal |
| Membrane | Soluble | Soluble |
Do N-terminal targeting signals (SignalP secretion, TargetP mitochondrial/chloroplast) differ between canonical and isoform? N-terminal changes most directly add or remove targeting peptides.
LLM reasoning
Does healthy human germline variation (gnomAD) avoid this region (depletion ratio < 1×), and is it intrinsically constrained (ESM-C constraint delta > 0)? Depletion of population variation plus high sequence constraint mean the region resists change — it is functionally important. Scored only where the unique region is canonical coding sequence, i.e. on truncations. gnomAD is a tolerance catalogue, not a disease one; disease/cancer variants (ClinVar / COSMIC) live in M2.
Comparison
Germline Variants · differential vs shared region
gnomAD (population) variant density per nucleotide — depletion ratio < 1× means healthy human variation avoids the region (constrained), the M1 basis alongside ESM-C constraint
| Variant set | Differential | Shared | Depletion ratio |
|---|---|---|---|
| gnomAD variants | 35 | 157 | 1.7× |
Sequence constraint (ESM-C) · differential vs shared region
lower LLR = more constrained; enrichment = differential / conserved
| Property | Differential | Shared | Enrichment |
|---|---|---|---|
| ESM-C mean LLR | -0.00177 | -0.0078 | — |
| Constrained positions | 0 | 0 | — |
Predicted-damaging variants · germline (gnomAD)
AlphaMissense / ESM-C / LoF calls per region (length-normalized)
| Variant set | Differential | Shared | Enrichment |
|---|---|---|---|
| Scorable variants | 33 | 155 | 1.5× |
| Damaging variants | 24 | 79 | 2.2× |
| — of which loss-of-function | 1 | 12 | 0.59× |
| AlphaMissense-pathogenic | 10 | 48 | 1.5× |
Predictor scores · germline (gnomAD)
scored: 254 ESM-C · 168 AlphaMissense
| Property | Differential | Shared |
|---|---|---|
| Mean ΔLLR (ESM-C) | -7.1 | -5.32 |
| Min ΔLLR (ESM-C) | -12.3 | -13.1 |
| Mean AlphaMissense | 0.516 | 0.592 |
Are disease (ClinVar / COSMIC) variants enriched per nucleotide in the differential region versus the shared core (disease enrichment ratio ≥ 1×)? Disease variants concentrating in the unique region tie it to phenotype.
Comparison
Clinical-variant burden · differential vs shared region
counts per region; ratio is density-normalized (variants per nucleotide, differential ÷ shared — ≥1× = disease variants concentrate in the differential region, the M2 basis)
| Variant set | Differential | Shared | Enrichment |
|---|---|---|---|
| Disease variants | 9 | 80 | 0.83× |
| Pathogenic | 0 | 0 | — |
Predicted-damaging variants · disease (ClinVar/COSMIC)
AlphaMissense / ESM-C / LoF calls per region (length-normalized)
| Variant set | Differential | Shared | Enrichment |
|---|---|---|---|
| Scorable variants | 9 | 79 | 0.81× |
| Damaging variants | 8 | 58 | 0.98× |
| — of which loss-of-function | 2 | 7 | 2× |
| AlphaMissense-pathogenic | 3 | 40 | 0.53× |
Predictor scores · disease (ClinVar/COSMIC)
scored: 254 ESM-C · 168 AlphaMissense
| Property | Differential | Shared |
|---|---|---|
| Mean ΔLLR (ESM-C) | -8.45 | -7.44 |
| Min ΔLLR (ESM-C) | -11.4 | -13.9 |
| Mean AlphaMissense | 0.604 | 0.686 |
LLM reasoning
Does the gained/lost region fold into ordered structure (pLDDT) and shift the protein's biophysical character (charge, hydropathy, disorder)? Structured, biophysically distinct regions are more likely to be functional.
Comparison
Structure (ESMFold2) · canonical vs isoform
TM-score 0.391 · RMSD 2.85 Å
| Metric | Canonical | Isoform |
|---|---|---|
| pLDDT (whole protein) | 0.822 | 0.634 |
Fold confidence · differential vs shared region
mean ESMFold2 pLDDT in each region
| Metric | Differential | Shared | Enrichment |
|---|---|---|---|
| pLDDT | 0.739 | 0.938 | 0.79 |
The shared region is the stretch of protein identical in both the isoform and the canonical (the canonical body for an extension; the post-truncation body for a truncation). Folded in both contexts it normally comes out nearly identical, so its Cα backbone RMSD ≈ 0. A high shared-region RMSD (superposed on the shared residues only, from the ESMFold2 structures) means the extension or truncation reorganizes how the retained region folds — a rare, high-interest functional signal. TM-score is a length-normalized companion; the check only fires when both structures are confidently folded, and uORF/altORF isoforms (no shared region) are not evaluable.
Comparison
Shared-region structural change
Cα RMSD superposed on the shared residues only; TM-score is length-normalized · shared-region Cα RMSD 22.4 Å · shared TM-score 0.417 · shared region 162 aa · min shared pLDDT 0.635 · global TM-score 0.391 · global RMSD 2.85 Å
| Property | Canonical | Isoform |
|---|---|---|
| Shared-region pLDDT | 0.834 | 0.635 |
Does the differential region contain actual secondary structure — a helix or strand — rather than coil? ESMFold2 predicts coordinates but no secondary structure, so elements are assigned from those coordinates (P-SEA, Cα geometry). Because they are derived from a prediction, each element carries its own mean pLDDT: a geometrically clean helix running through a disordered stretch is geometry fitted to a guess, so BOTH length and confidence are required to score. The direction differs by ORF type — an extension GAINS the element (a candidate functional addition), a truncation LOSES one, and there the element is read off the canonical structure because the removed segment exists only in it. This says nothing about whether the element is integrated with the rest of the fold; the contact and PAE evidence in this same category answers that.
Comparison
Secondary structure · differential vs shared region
helices and strands assigned from the predicted coordinates (P-SEA)
| Metric | Differential | Shared |
|---|---|---|
| Alpha helices | 0 | 3 |
| Beta strands | 0 | 2 |
| Longest element (aa) | 0 | 15 |
| Mean pLDDT | — | 0.88 |
Elements and coordinates
0 in the differential region, 5 in the shared core — residue numbering is 1-based on the protein holding the region
| Removed (canonical) | Shared core |
|---|---|
| — | alpha helix 61–75 15 aa · pLDDT 0.82 |
| — | beta strand 86–92 7 aa · pLDDT 0.79 |
| — | beta strand 143–148 6 aa · pLDDT 0.90 |
| — | alpha helix 149–154 6 aa · pLDDT 0.93 |
| — | alpha helix 157–166 10 aa · pLDDT 0.94 |
Below threshold
0 in the differential region, 4 in the shared core — shorter than 6 aa or below pLDDT 0.70, so not counted above
| Removed (canonical) | Shared core |
|---|---|
| — | beta strand 35–39 5 aa · pLDDT 0.92 |
| — | beta strand 50–53 4 aa · pLDDT 0.92 |
| — | beta strand 96–100 5 aa · pLDDT 0.74 |
| — | beta strand 168–172 5 aa · pLDDT 0.83 |
LLM reasoning
Does the isoform gain or lose a real InterPro functional domain in the differential region (disorder/structural-only signatures excluded)? Gaining or losing a domain changes function directly.
Comparison
Domains & motifs (canonical vs isoform)
gained = only in the isoform; lost = only in the canonical
| Feature | Canonical | Isoform |
|---|---|---|
| InterPro domains | 15 | 15 |
| Short linear motifs | 1 | 1 |
Biophysical character of the isoform-differential region versus the shared canonical core — pI, hydropathy, charge, disorder and related properties. Descriptive; the folding (P1) score keys off the GRAVY / charge / disorder deltas.
Comparison
Biophysics · differential vs shared region
highlighted rows are enriched in the differential region
| Property | Differential | Shared | Enrichment |
|---|---|---|---|
| Isoelectric point (pI) | 4.3 | 4.6 | 0.935 |
| Hydropathy (GRAVY) | -1.47 | -1.23 | 1.19 |
| Fraction charged | 0.565 | 0.442 | 1.28 |
| Disorder fraction | 0.33 | 0.214 | 1.54 |
| Disorder-promoting | 0.652 | 0.681 | 0.958 |
| Low-complexity fraction | 0.783 | 0.117 | 6.71 |
| Prion-like fraction | 0.174 | 0.203 | 0.859 |
| LLPS score | 0.27 | 0.176 | 1.53 |
| π–π propensity | 0.13 | 0.178 | 0.733 |
| Aromaticity | 0.0435 | 0.0736 | 0.591 |
| Instability index | 108 | 45 | 2.4 |
| Shannon entropy | 2.63 | 3.96 | 0.664 |
| Normalized complexity | 0.608 | 0.915 | 0.664 |
Sparse-autoencoder (SAE) interpretability features on the ESM-C residual stream, comparing the isoform against the canonical protein. This is the S3 criterion — a presence check that fires when any interpretable feature is gained or lost. Only the top 30 features per category (ranked by activation, prevalence ≥ 2) are saved; each table shows the top 10 interpretable features, and generic / "unknown" features are hidden unless you expand it. What are SAE features? →
Part 1 · Global feature comparison
Which SAE features fire in the isoform vs the canonical protein, across the whole sequence.
Isoform-only features — 20 total (prevalence ≥ 2)
| Feature | Label | Description | Activation | Prevalence |
|---|---|---|---|---|
| #15915 | C2B and RabBD loop signal | Signal for loops within C2 domains, particularly the second C2 domain (C2B) of Rab3-interacting and synaptotagmin-like proteins, with additional firing in the N-terminal Rab-binding domain (RabBD) region of these same proteins. Activation patches contain Ser/Thr, acidic residues (E/D), and frequent Pro/Gly, consistent with C2-domain inter-strand loops. | 3.76 | 2 |
| #7903 | Arginine-rich polybasic patches | Basic polycationic patches enriched in arginine (often with lysine) within low-complexity/disordered regions, frequently N-terminal. These include classical monopartite NLSs, nucleic acid–binding basic tails, signal peptide n-regions and cytosolic juxtamembrane segments (positive-inside rule), and basic clusters in secreted precursors; typically depleted of acidic residues. | 3.04 | 2 |
| #11943 | Active-site proximal loops | Active-site–proximal flexible loops/turns (beta-strand turns and coil-to-helix junctions) that flank catalytic residues or cofactor/metal-binding pockets across diverse enzyme folds; these segments often contain or neighbor acidic (Asp/Glu) and glycine residues (with occasional Lys/Arg and aromatics) that position substrates, metals, or nucleotides for catalysis. In secreted peptide precursors, the feature also marks short acidic/Gly-rich turns at proteolytic processing sites that enable maturation. | 2.91 | 2 |
| #13152 | Active-site pocket-lining turns | Short, ordered coil/turn segments lining or gating enzyme active/ligand-binding pockets, typically sparing the catalytic residue itself. | 2.87 | 2 |
| #1846 | Glycine-rich active-site loops | Glycine-containing, flexible loop/turn motifs (often with adjacent polar/charged residues) that form catalytic and cofactor/substrate-binding loops in enzymes—covering classic glycine-rich binding loops (e.g., NAD(P)-, ATP/CoA-binding loops), oxyanion/metal-binding loops, and electrostatic substrate-guidance loops—across diverse taxa and cellular locales. | 2.40 | 2 |
| #8665 | Conserved domain-edge hydrophobic loop | Short hydrophobic/polar sequence patches located at domain edges, recurring N-terminal to catalytic Rab-GAP TBC domains and in cytoplasmic regions of small membrane-associated subunits, likely marking a conserved structural segment adjacent to ordered β-strand or short helix | 2.33 | 2 |
| #16372 | Lysine-rich electrostatic/PTM hotspots | A general lysine-centric signal: the feature fires on lysine residues in basic patches and K‑rich stretches that mediate electrostatic interactions with nucleic acids or acidic partners, and on functionally reactive lysines that undergo lysine-specific chemistry (acetyl/methyl/ubiquitin/SUMO in regulators; hydroxylation/oxidative crosslinking in ECM), yielding broad activation across transcriptional regulators and extracellular matrix proteins | 2.20 | 2 |
| #16283 | Short His/Gly-rich active-site loops | Short polar catalytic/cofactor-binding loops, most characteristically the RHG motif of the histidine phosphatase superfamily (RHGXRXP), but broadly recognizing analogous His/Gly–rich segments (HxG, HGG, HxH/HxE, HGTGT) and related short glycine-/acidic-/cysteine-enriched motifs that form active-site or cofactor/metal-binding loops across diverse enzyme classes (phosphoryl transfer, porphyrin chelation, redox oxygenases/monooxygenases, electron transfer, and carbohydrate hydrolases). | 2.17 | 2 |
| #5045 | N-terminal IDR and signal-peptide activation | N-terminal pre-domain segments used for targeting or regulation: either intrinsically disordered, low‑complexity, Ser/Thr/Pro/Gly- and acidic‑rich regulatory tails with clustered phosphosites, or short hydrophobic signal peptides; activation consistently fades before downstream structured domains. | 2.15 | 2 |
| #10487 | Nuclear regulatory low-complexity IDRs | Intrinsically disordered, low‑complexity regulatory regions of large metazoan regulatory and scaffolding proteins—enriched in nuclear gene‑regulatory factors—especially transactivation domains, flexible linkers, and long regulatory tails that are S/P/Q/G/T- and acidic‑rich, harbor SP/TP phospho‑motifs and short homorepeats; the feature preferentially avoids compact folded DNA-/ligand-/zinc‑binding domains. | 1.99 | 2 |
| #14887 | Start of mature domain | Start-of-mature-domain segments, often the post‑signal‑peptide stretch or the first folded subdomain following an N‑terminal region: windows that frequently form the first helix or β‑strand of the mature domain, sometimes preceded by disordered low‑complexity tails enriched in small/polar residues (P/G/S/T/A). These regions mark the onset of a structured domain rather than catalytic cores, and often carry localization/export or assembly cues (e.g., secretion signals, transit/NLS motifs). | 1.95 | 2 |
| #6836 | N-terminus and beta-strand boundary peaks | Residue-level feature whose strongest peaks fall near N-terminal/chain-start positions of bacterial and mitochondrial proteins, with weaker secondary peaks at beta-strand starts/ends and adjacent coils in beta-rich and alpha/beta domains, generally away from catalytic centers. | 1.91 | 2 |
| #10290 | Disordered phospho-rich regulatory tails | Eukaryotic intrinsically disordered, low‑complexity regulatory linkers and tails — often enriched in polar and acidic residues, with embedded phospho‑Ser/Thr/Tyr clusters — in scaffold/adaptor and signaling proteins; well‑folded domains such as PTB are de‑emphasized. | 1.91 | 2 |
| #4673 | NA-binding SLiMs and cystine-knots | Short conserved interaction motifs and broader disordered/low-complexity binding segments in nucleic-acid/chromatin-associated proteins, with peaks in disordered acidic/basic and Pro/polar-rich tails of retroelement Pol proteins, key DNA/RNA-contacting loops/β-strands (OB-folds, polymerase loops), and disordered acidic/trafficking SLiMs (e.g., AHA, NES/NLS) that mediate regulatory binding; also includes disulfide-bonded cystine-knot cores in secreted TGF-beta/BMP/GDF growth factors. | 1.91 | 2 |
| #3660 | Low-complexity basic IDRs and micro-TMs | Generic low-complexity segments that are intrinsically disordered, proline‑rich and/or Lys/Arg‑biased (often also enriched in Ser/Thr) with intermittent hydrophobic residues, together with short hydrophobic helices in small membrane proteins; these include propeptide/processing regions of secreted precursors, basic disordered segments in viral proteins, mucin‑like S/T‑rich patches in glycoproteins, surface loops in enzymes, and single‑pass transmembrane microproteins (e.g., organellar gene products) rather than a specific folded domain. | 1.82 | 2 |
| #4850 | RRM-disordered linker activation | RRM-containing eukaryotic RNA-binding proteins, with broad activation spanning both the RRM domains and adjacent/inter-RRM disordered regions; covers splicing factors, hnRNP family, SR-type shuttling mRNA-binding proteins, cyclophilin-type RRM proteins, and plant glycine-rich RBPs. | 1.76 | 3 |
| #9537 | N-terminal low-hydrophobic presequence detector | N-terminal low-hydrophobic presequence detector: preferentially marks low-complexity, small/polar-rich leaders (e.g., chloroplast transit peptides and some mitochondrial transit peptides), but generally not classical hydrophobic signal peptides or strongly Arg-rich amphipathic leaders | 1.75 | 2 |
| #2588 | Exposed beta-loop interfaces | A cross-kingdom feature marking solvent-exposed beta-strand/loop segments within repeated, beta-rich binding/scaffold domains and processivity modules, recurring in nucleic-acid–binding S1/OB-fold proteins (rRNA biogenesis), DNA-binding plant factors (AP2/ERF, B3), replication clamps, extracellular adhesion modules (Ig/FN3/Laminin-G) and Ly6/uPAR three-finger proteins, as well as cupredoxin/cupin enzyme scaffolds; also present in nuclear transcriptional regulators. | 1.69 | 2 |
| #437 | Short linear helical motifs | Short alpha-helical linear motifs (∼8-20 aa) used for assembly/targeting and interface formation—often amphipathic, mixed charged/hydrophobic (E/D/K/R with hydrophobics) and frequently proline-flanked in flexible regions, but also including short hydrophobic helices such as brief transmembrane segments. | 1.59 | 2 |
| #15733 | Eukaryotic signal peptides and IDRs | Eukaryotic compositionally biased low‑complexity/IDR segments—including N‑terminal signal peptides and long Ser/Thr/Pro/Lys/Glu‑rich tracts—that serve as regulatory and scaffolding regions (phospho‑regulated, proline‑rich) in cell‑cycle/mitotic and microtubule/ciliary proteins, and as targeting/propeptide regions in secreted peptide precursors; also present as basic/acidic tails in RNA/translation factors and organellar ribosomal proteins. | 1.52 | 2 |
Canonical-only features — 96 total (prevalence ≥ 2)
| Feature | Label | Description | Activation | Prevalence |
|---|---|---|---|---|
| #9666 | Alpha-helix caps and anchors | Residue-level recognition of alpha-helix termini/interfacial anchor residues: polar/aromatic sites (e.g., Asn/His/Tyr/Gly and nearby Y/F/W often following Lys/Arg) that cap or stabilize helices, including membrane–interface aromatic anchors and soluble helix caps, occurring broadly across proteins | 10.72 | 2 |
| #15815 | Basic N-terminal targeting patch | Short, basic/polar N‑terminal leader/transit segment immediately after the initiator methionine (positions ~3–6) — i.e., the positively charged/polar start of signal peptides and organelle transit peptides, and more generally low‑acidic, disordered N‑terminal patches (including occasional cysteine‑rich starts) used for targeting, export, rRNA binding, or membrane topology cues | 9.70 | 4 |
| #10704 | N-terminal PEST-like IDRs | Low-complexity, often acidic/proline-rich intrinsically disordered N-terminal tails of intracellular eukaryotic proteins, frequently overlapping annotated PEST-like segments. | 8.87 | 10 |
| #7996 | Generic extreme N-terminus detector | Generic extreme N-terminus detector: marks residues around position 5 (occasionally 6–7) at the very start of the polypeptide, most often within the N-region of signal/transit peptides of secreted or organelle-targeted precursors, but also in generally disordered N-termini of cytosolic/nuclear proteins; not motif- or residue-specific | 8.62 | 2 |
| #10977 | Early N-terminal disordered patch | Short N-terminal patch at the extreme N-terminus (around residues 5–10), capturing the first distinctive segment of preproteins/precursors and N-terminal disordered tails; residue composition is heterogeneous but frequently polar (Ser/Thr) and/or charged. | 6.83 | 3 |
| #8021 | Extracellular exposed-loop activations | Extracytoplasmic/secreted proteins and extracellular or luminal domains (secretory pathway, periplasm, outer membrane, virion surface), with enrichment for carbohydrate-active enzymes and Ca2+-dependent recognition modules; the feature highlights solvent-exposed loop/turn residues across these regions rather than a specific catalytic motif. A minority of cytosolic glycosidases with similar loop features can also activate. | 6.22 | 2 |
| #8063 | Disordered N-terminal activation window | Short N-terminal disordered segment found in a variety of proteins, including plant chloroplast transit peptides, bacterial prokaryotic ubiquitin-like protein Pup N-termini, anti-sigma factor N-tails, and some small RiPP precursor leaders. The activating region is Ser/Thr/Ala/Pro/Gly/Gln-rich, low in acidic residues, and typically lies in a disordered N-terminal stretch; the feature does not generally mark all transit peptides or all RiPP leaders. | 5.34 | 7 |
| #5506 | Cysteine-rich extracellular repeat modules | Cysteine-rich, disulfide‑stabilized extracellular repeat modules—principally EGF‑like repeats and closely related small domains (CCP/Sushi, WAP/TIL, disintegrin, LDLR class‑A, and von Willebrand–type modules)—found across secreted proteins, extracellular matrix components, proteases/inhibitors, and receptor ectodomains. | 5.00 | 2 |
| #11990 | Paired hydrophobic anchor motif | A broad, structural micro‑motif: a pair of hydrophobic “anchor” residues that sit on short β‑strand or adjacent core positions stabilizing compact helix‑based modules (winged‑helix/HTH, histone‑fold handshakes, EF‑hand pairs, HhH/base‑flipping elements). These anchors lie next to, but not at, the primary functional sites (DNA‑ or Ca2+‑binding loops), and recur early in proteins across diverse taxa. | 4.80 | 2 |
| #3505 | Extreme N-terminus recognition | Recognition of the extreme N-terminus of proteins, especially the signal/transit-peptide cleavage junction and the first residues of the mature chain; more generally, short disordered, polar/charged N‑terminal segments (often including the initiator Met) across diverse enzymes and taxa. | 4.24 | 4 |
| #8262 | Charged coil-to-beta cap motif | A short, polar/charged loop/turn motif at coil→β‑strand junctions and β‑hairpin connectors—typically capping or initiating β‑strands at domain boundaries or adjacent to catalytic/binding sites. The sequence signal is a 3–10 residue patch enriched in K/R, D/E, S/T, and small residues (G/P), and recurs across many folds rather than a family‑specific consensus. | 4.21 | 3 |
| #9145 | Helix–loop boundary/linker signal | A generic helix–loop boundary/linker signal: the feature prefers short coil segments and helix-capping/hinge positions that connect or terminate alpha-helices, especially interhelical loops in multi-pass membrane proteins and structured active-site/cofactor-binding loops in soluble enzymes; it tolerates diverse amino acids but often includes small helix-modulating or interfacial aromatic residues. | 4.16 | 2 |
| #9632 | Disordered acidic/mixed-charge interaction modules | Intrinsically disordered segments in diverse phosphoproteins and chromatin-associated factors—often with mixed basic/acidic or acidic-leaning composition—that act as flexible interaction modules; widely used in chromatin/transcription factors and also present in diverse acidic phosphoproteins across taxa. | 4.12 | 6 |
| #9980 | Phosphate-binding loop-adjacent coils | Short coil/loop residues immediately preceding or within conserved phosphate-binding loops of NTP-utilizing enzyme domains, primarily the Walker A/G1 P-loop of P-loop NTPases (AAA+ ATPases, SF1/SF2 helicases, translational GTPases), and also analogous ATP-binding loops in related kinase folds (e.g., APS kinase P-loop kinases, Bergerat-fold histidine kinases); occasional secondary activation on gly/lys-enriched loops adjacent to functional sites in non-NTP enzymes and small translation factors (e.g., glyoxalase I, IF-1). | 4.11 | 2 |
| #3860 | GH5 TIM-barrel stabilizing residues | A signal that fires at conserved residue positions within the catalytic (GH5) domain of fungal endo-β-1,4-mannanases, marking stabilizing positions in the well-ordered TIM-barrel scaffold rather than catalytic or substrate-binding residues. | 4.08 | 2 |
| #693 | Helix-capping repeat boundary motifs | Short helix-capping and inter-helix turn/linker motifs at the boundaries of alpha-helical repeat units and alpha‑solenoid scaffolds (e.g., HEAT/ARM/ANK/CATCHR); the feature also fires at similar helix–loop junctions in other helical domains and in small helical RNA‑binding proteins. | 3.89 | 2 |
| #792 | Short solvent exposed loops | Short, solvent‑exposed coil/turn segments (ordered loops) that connect secondary‑structure elements, most prominently in extracellular/periplasmic/lumenal ectodomains but also in soluble enzymes and nucleic‑acid–binding proteins; compositionally enriched in Gly and polar/acidic residues with frequent Ser/Thr and Lys/Arg; not tied to specific catalytic motifs but marking surface loops/turns (e.g., turns between beta strands) commonly used for flexibility and binding interface shaping | 3.86 | 3 |
| #13538 | Pocket-lining strand-edge loops | Structured loop/turn residues at beta-strand edges in well-structured enzyme cores (including Rossmann-like, TIM-barrel, alpha/beta hydrolase, beta-propeller, and P-loop NTPase folds) that line cofactor/substrate-binding pockets, typically enriched in Gly and polar/charged residues, and flanking rather than constituting the catalytic residues | 3.76 | 2 |
| #14778 | FBPA N-terminal helix motif | Short N-terminal alpha-helical patch enriched in small hydrophobics (A/V/I) with frequent Ser/Thr, corresponding to the start of the first alpha-helix of Class II fructose-1,6-bisphosphate aldolases (FBPA). | 3.59 | 2 |
| #10556 | Helix termini and linkers | Structural motif: alpha-helical segments with a strong preference for helix termini and adjacent coil/turns (helix-capping positions and short helix–loop linkers), largely independent of protein function | 3.57 | 2 |
| #15526 | Polar low-complexity disordered regions | Intrinsically disordered, low‑complexity regions enriched in polar/acidic and amide residues—especially Q/E/S/T/P/G—often including proline‑rich tracts and S/T clusters; these segments are flexible, cysteine‑poor, and typically lie outside structured catalytic cores (e.g., disordered tails, propeptides, and regulatory regions). | 3.56 | 2 |
| #8396 | Amphipathic TPR helix starts | Amphipathic alpha‑helical repeat elements characteristic of tetratricopeptide repeat (TPR) and related alpha‑solenoid scaffolds, with a bias toward the N‑terminal boundary of individual TPR helices used for protein–protein docking; the feature also weakly generalizes to analogous helices in other coiled‑coil/alpha‑helical complex subunits. | 3.49 | 2 |
| #1370 | Short low-complexity N-terminal segment | Short, low‑complexity, intrinsically disordered N‑terminal segments (~10–20 aa) that frequently correspond to leader/transit/propeptide presequences (e.g., chloroplast transit peptides, RiPP/bacteriocin leaders) or generic unstructured N‑tails; typically non‑hydrophobic and enriched in Ser/Thr/Pro but tolerant of basic (poly‑Arg/Lys) or Cys‑rich compositions | 3.42 | 8 |
| #7372 | Helix capping transition sites | Alpha-helix termination/capping residues and helix–transition junctions (particularly C-caps and the immediately following turn/loop or helix–β connectors), which often stabilize helix ends and frequently flank catalytic or ligand-binding pockets across diverse proteins | 3.23 | 2 |
| #11393 | Small-residue coil/turn micro-motifs | Short, flexible coil/turn micro‑motifs enriched in small/turn‑prone residues (G, A, S, T, P and small hydrophobics) that demarcate functional segment boundaries—e.g., cores of processed small peptides and structured loops flanking catalytic or metal‑binding sites—often near N‑termini but also occurring mid‑sequence; catalytic residues themselves are typically not the peaks. | 3.16 | 2 |
| #10406 | β-strand hydrophobic anchor detector | Structural detector for hydrophobic anchor residues on β‑strands—most often the earliest (N‑terminal) strand—of β‑sheet/β‑barrel domains, especially OB/cold‑shock folds, but also other β‑rich domains across diverse proteins; it highlights packing positions rather than catalytic or basic loop/helix sites. | 3.09 | 2 |
| #3103 | SPT/LCB2 DFY motif detector | Residue-level detector that fires predominantly on a small set of conserved positions within the PLP-dependent aminotransferase fold of serine palmitoyltransferase / LCB2-family enzymes, with a single dominant peak on a conserved aromatic (Phe) in a "D/E-F-Y" motif upstream of the catalytic core. | 3.05 | 2 |
| #127 | Proline-rich Ser/Thr-biased IDRs | Intrinsically disordered, low‑complexity proline‑rich segments enriched in Ser/Thr (polyproline, PxxP, and SP/TP motifs) that form eukaryotic regulatory tails/linkers/activation domains; the feature marks these S/T/P‑biased IDRs rather than folded domains. | 3.02 | 2 |
| #14927 | Structural and domain boundary residues | Residues at structural and domain junctions: starts/ends of secondary-structure elements (especially helix starts and transmembrane-helix termini) and positions immediately preceding or following annotated domains, often adjacent to ligand/cofactor-binding regions | 3.01 | 2 |
| #6626 | Catalytic-adjacent helix–strand loops | Short loop/turn motifs at secondary-structure junctions—especially helix–strand (alpha–beta or beta–alpha) connectors and occasional strand–strand loops—often centered on glycine or other small/polar residues, typically within catalytic domains. These positions commonly flank (but rarely coincide with) catalytic or substrate-binding sites, making them particularly prominent in glycoside hydrolases and other carbohydrate-active enzymes, yet they also appear in unrelated enzymes. | 2.98 | 2 |
Shared features by |Δ| activation — 859 total (prevalence ≥ 2)
| Feature | Label | Description | Δ activation | Iso act. | Canon act. | Iso prev. | Canon prev. |
|---|---|---|---|---|---|---|---|
| #1135 | Homopolymeric low-complexity tracts | Compositionally biased, low-complexity sequence segments characterized by homopolymeric residue runs; the feature detects local stretches of repeated single amino acids regardless of chemistry (polar, basic, or hydrophobic), most often within disordered or unstructured contexts but also occasionally within folded domains where such runs occur. | -2.62 | 2.34 | 4.96 | 8 | 18 |
| #6360 | N-terminal low-complexity IDRs | Eukaryotic intrinsically disordered, low‑complexity N‑terminal tails enriched in Lys/Arg/Ser/Asp/Glu that precede structured domains of RNA biogenesis/translation and chromatin‑remodeling factors (nucleolar/ssu‑processome, spliceosome/DEAD‑box helicases, H2A.Z chaperones); these flexible segments are typical interaction/assembly regions. | -2.55 | 2.21 | 4.77 | 7 | 23 |
| #4977 | Strand-edge turn motifs | Short coil-to-β-strand transition motifs—beta-turns/strand-edge loops (often the N-termini of β-strands)—recurring across diverse folds and functions. | -2.33 | 6.12 | 8.45 | 11 | 19 |
| #13334 | N-terminal IDR boundary | Recognition of short, low-complexity, intrinsically disordered N-terminal tails immediately upstream of the first folded domain; these segments are often Ser/Thr/Pro/Gly/Ala-biased and sometimes extend a few residues into the initial helix or the very beginning of the first domain (domain boundary/disorder-to-order transition), especially in transcription regulators and co-regulators. | -2.31 | 2.68 | 5.00 | 6 | 25 |
| #6484 | Prion-like low-complexity IDRs | Polar low-complexity intrinsically disordered regions (IDRs)—especially Q/N-rich (prion-like) segments and S/T- and glycine-rich tracts with simple SG/TG/PG repeats—typically located in inter-domain linkers and terminal tails of eukaryotic proteins and often harboring Ser/Thr phosphorylation sites. | -1.81 | 2.32 | 4.13 | 9 | 6 |
| #6016 | Helix-biased internal methionine detector | Detector for methionine residues, firing on internal Met across diverse proteins with somewhat enhanced response when Met occurs in helical or low-complexity contexts; occasional hits at the initiator Met when embedded in a locally Met- or hydrophobic-rich N-terminus; weak cross-reactivity to other bulky hydrophobics (notably tryptophan). | +1.78 | 7.61 | 5.83 | 3 | 3 |
| #14712 | SPS N-terminal regulatory motif | A conserved N-terminal regulatory motif in plant sucrose-phosphate synthases (SPS), centered on a short ATR(N/S/P)xRERS / SPQERN-type sequence upstream of the catalytic glycosyltransferase core. | -1.68 | 4.98 | 6.66 | 9 | 11 |
| #9332 | Hydrophobic aliphatic core packing | Hydrophobic aliphatic residue packing: the feature marks individual Val/Ile (with Ala/Leu/Thr) at nonpolar, well-ordered positions in secondary structure (alpha-helices and beta-strands), frequently forming the hydrophobic core or lining/co-supporting ligand/nucleotide pockets across diverse folds and taxa. | -1.66 | 3.20 | 4.86 | 4 | 4 |
| #512 | Short functional domain segments | Short sequence segments distributed across diverse proteins, with peaks that often fall within structured catalytic, transporter, or fold-defining domains as well as occasional acidic/charged stretches and flexible linkers. | +1.61 | 3.43 | 1.82 | 10 | 5 |
| #5022 | Disordered zinc-finger linkers | Short, intrinsically disordered linker segments that flank or connect zinc‑binding domains—especially tandem zinc‑finger repeats (C2H2, RanBP2/CCCH, matrin‑type)—with selectivity for the immediate N‑terminal side or inter‑finger spacers, rather than the folded zinc‑binding cores or other structured domains; most common in eukaryotic chromatin/nuclear regulatory proteins. | -1.58 | 3.27 | 4.85 | 7 | 16 |
| #13016 | Acidic low-complexity IDRs | Acidic intrinsically disordered low-complexity regions (IDRs)—often terminal tails or linkers—i.e., long D/E-rich acidic tracts commonly found in nuclear/chromatin proteins, viral tegument/IE phosphoproteins, and cytoskeletal regulators | -1.57 | 4.74 | 6.31 | 13 | 21 |
| #9627 | Short Pro/Gly-rich strand-turn patches | Short, compositionally biased strand/turn segments that nucleate or flank brief secondary‑structure elements (typically beta‑strands with adjacent loops/turns, occasionally short helices), enriched in Pro and Gly together with aliphatic residues (Leu/Ile/Val) and frequent Ser/Thr; these patches occur both within folded beta‑sheet domains and in intrinsically disordered regions (often at coil↔secondary‑structure transitions), and can host linear motifs such as leucine‑rich nuclear export signals or mark the first structured residues after signal‑peptide cleavage. | -1.55 | 2.10 | 3.65 | 7 | 15 |
| #13080 | Phospho-acidic disordered tails | Acidic, serine-rich low-complexity intrinsically disordered regions (IDRs)—typically long N- or C-terminal tails enriched in Asp/Glu and Ser/Thr-Pro motifs and phosphoserine sites—i.e., phospho-acidic regulatory stretches rather than structured domains. | -1.42 | 4.12 | 5.54 | 5 | 15 |
| #5076 | N-terminal disordered pre-sequences | Intrinsic N-terminal pre-sequences and regulatory tails: extended, low-complexity, intrinsically disordered segments at protein termini—especially cleavable signal/transit peptides and propeptides, as well as noncleavable acidic/PST-rich regulatory tails that harbor PTM sites and linear binding motifs—marking the boundary before structured mature domains. | -1.39 | 2.38 | 3.78 | 43 | 64 |
| #1501 | Acidic phosphosite-rich IDRs | Low-complexity intrinsically disordered segments enriched in polar residues with embedded acidic/phospho-regulatory clusters, serving as interaction modules across diverse proteins (nuclear/chromatin/RNA-associated and also secreted) | -1.37 | 5.18 | 6.55 | 10 | 21 |
| #1584 | Acidic serine-rich disordered tails | Long acidic, serine-enriched intrinsically disordered low‑complexity regions (DE/S‑rich IDRs), typically found as internal or C‑terminal tails of chromatin- and transcription-associated eukaryotic proteins (and mimicked by large DNA viruses), often densely phosphorylated and depleted of bulky hydrophobics. | -1.36 | 2.59 | 3.95 | 12 | 25 |
| #6205 | Broad domain-level activation | A broad, low-amplitude sensor that activates across most of the sequence in folded proteins, with particular prominence in DEAD-box / SF1-SF2 RNA helicase ATP-binding domains. Activation is broadly distributed rather than localized to specific catalytic residues, and tends to be weaker on N-terminal targeting/transit peptides and cleavable leaders than on the mature chain. | -1.33 | 2.40 | 3.72 | 8 | 17 |
| #13982 | Bromodomain and HTH core signature | A eukaryotic nuclear recognition-module signature that targets compact all-alpha binding cores—most strongly bromodomains (acetyl-lysine reader pocket and adjoining helices) but also helix-turn-helix/winged-helix nucleic-acid-binding cores (e.g., La/ARID/ETS/IRF/HSF). The shared signal is a conserved hydrophobic/basic recognition element within these reader cores. | -1.24 | 2.11 | 3.36 | 4 | 7 |
| #13122 | Charged polar low-hydrophobicity segments | Compositionally biased, low-hydrophobicity segments enriched in charged and small polar residues (E, D, K, R, S, T; often with P/G), most often found as intrinsically disordered N‑terminal stretches or propeptides but also occurring as surface-exposed polar helices/β-strands within folded domains | -1.24 | 2.44 | 3.68 | 5 | 19 |
| #10941 | Amphipathic alpha-helical interaction motifs | Generic preference for well-ordered alpha-helical elements, with strongest response to amphipathic, interaction-mediating helices used across eukaryotic regulatory proteins (chromatin/transcription factors and signaling/scaffold proteins), including both short binding motifs (e.g., IQ and AKAP RII-binding helices) and multi-helix domains (e.g., SEC7, bromodomains, coiled-coils, EF-hand helices). | -1.24 | 3.58 | 4.82 | 12 | 14 |
| #3970 | TAFA conserved residue motifs | Residues at conserved positions within small secreted cysteine-containing proteins, particularly the TAFA/chemokine-like family, with additional sparse activations in folded enzymatic domains away from annotated catalytic residues. | -1.22 | 3.71 | 4.93 | 3 | 4 |
| #14076 | GARIN central LxLTR motif | A conserved central segment in Golgi-associated RAB2 interactor (GARIN) family proteins, characterized by a short "(L/V)xLTR"-like motif (e.g., LELTR, LVLTR) preceded by a basic/polar stretch (often K/R/S-rich). | -1.19 | 4.88 | 6.07 | 14 | 31 |
| #8115 | GNAT beta-strand core hydrophobics | Hydrophobic residues on conserved beta-strands that form the core of α/β folds, most characteristically the central beta-sheet of GNAT/NAT acetyltransferases lining the acetyl‑CoA/substrate-binding groove; the same signal appears at structurally analogous beta‑strand positions adjacent to catalytic or other functionally critical sites in diverse proteins, including non-enzymes such as ribosomal protein uL22. | -1.16 | 2.42 | 3.58 | 2 | 2 |
| #5310 | Catalytic Asp N-terminal strand | Short beta-strand or tight-loop elements that sit immediately N-terminal to catalytic metal-binding acidic residues (Asp/Glu) in nuclease active sites—especially RecB/RecB-like and VRR-NUC folds—and, more broadly, analogous beta-edge segments lining nucleotide/metal-binding pockets in other DNA-processing enzymes; peaks frequently center on glycine at a coil-to-beta transition. | -1.16 | 6.18 | 7.34 | 7 | 10 |
| #10142 | Transsulfuration N-terminal P[IVL][Y/F] motif | Family-specific sequence detector that fires on a conserved N-terminal "P-[IVL]-[Y/F]" stretch characteristic of PLP-dependent transsulfuration enzymes, with a secondary site near a C-terminal D/E-L-[IV]-R-[IL] segment. | -1.14 | 2.96 | 4.10 | 3 | 5 |
| #13204 | SRA/YDG DNA methylation reader | YDG/SRA domain of UHRF1-family and related RING-type E3 ubiquitin ligases involved in DNA methylation reading, with activation concentrated on the structured base-flipping/5-methylcytosine-binding region. | -1.06 | 2.67 | 3.73 | 13 | 23 |
| #8939 | N-terminal regulatory tail | N-terminal sequence tracts (often disordered or low-complexity, sometimes acidic/polar) within the first ~20-60 residues of diverse eukaryotic and viral proteins; a generic "regulatory/tethering" N-terminal tail feature. | -1.04 | 1.80 | 2.84 | 2 | 4 |
| #10973 | PDZ beta-sandwich GLGF loop | PDZ domains—specifically the β‑sandwich core and the conserved carboxylate‑binding loop (GLGF/GFGF/GGx) that engages C‑terminal PDZ‑binding motifs | -1.00 | 2.48 | 3.49 | 10 | 13 |
| #14534 | Unknown generic feature | Unknown generic feature | -4.91 | 9.70 | 14.60 | 160 | 181 |
| #1803 | Unknown generic feature | Unknown generic feature | -2.43 | 14.87 | 17.30 | 161 | 182 |
Part 2 · Differential coordinates
Features firing on just the canonical-lost (truncated) residues.
Unique-region features — 203 distinct (prevalence ≥ 2)
| Feature | Label | Description | Activation | Prevalence |
|---|---|---|---|---|
| #2114 | Chromodomain and integrase-proximal activation | A feature that activates broadly across chromodomain-containing proteins and other chromatin-associated factors, as well as on retrotransposon Gag-Pol polyproteins (in integrase-proximal regions). Activation is distributed widely along the chain with peaks tending to fall within or adjacent to chromodomain folds rather than on catalytic enzyme cores. | 13.35 | 23 |
| #5780 | Conserved N-terminal beta-strand cores | Conserved β‑strand residues within β‑sheet cores—especially in OB/Cold‑shock–like β‑barrels and analogous β‑rich domains—most often at the first β‑strands near domain N‑termini; captures the structured strand face used widely in nucleic‑acid–binding modules and in diverse β‑sheet proteins | 11.30 | 3 |
| #13702 | Regulatory low-complexity IDRs | Eukaryotic intrinsically disordered, low‑complexity regulatory segments (terminal tails and inter‑domain linkers) enriched in Pro/Ser/Thr/acidic composition and short linear motifs for modular protein–protein interactions—frequently including WW‑domain–binding PY/PPxY segments and PTM hotspots—with reduced but not absent activation within folded recognition/catalytic domains (WW, PTB/PID, chromo/chromoshadow, SET). | 10.05 | 23 |
| #15815 | Basic N-terminal targeting patch | Short, basic/polar N‑terminal leader/transit segment immediately after the initiator methionine (positions ~3–6) — i.e., the positively charged/polar start of signal peptides and organelle transit peptides, and more generally low‑acidic, disordered N‑terminal patches (including occasional cysteine‑rich starts) used for targeting, export, rRNA binding, or membrane topology cues | 9.70 | 4 |
| #10704 | N-terminal PEST-like IDRs | Low-complexity, often acidic/proline-rich intrinsically disordered N-terminal tails of intracellular eukaryotic proteins, frequently overlapping annotated PEST-like segments. | 8.87 | 10 |
| #4977 | Strand-edge turn motifs | Short coil-to-β-strand transition motifs—beta-turns/strand-edge loops (often the N-termini of β-strands)—recurring across diverse folds and functions. | 8.45 | 9 |
| #5895 | PDZ and chromatin readers | Eukaryote-biased feature marking scaffold/signaling PDZ-domain proteins and nuclear chromatin regulators (readers/writers/remodelers), with additional coverage of ubiquitin/SUMO-pathway E3s, polarity/Rho-pathway components, PP1 glycogen-targeting CBM21 proteins, and ATAT1 acetyltransferases; broad functional/subcellular coverage and low specificity in SwissProt. | 7.82 | 2 |
| #300 | AT-hook-like basic IDRs | Arg/Lys-rich, glycine/proline-spaced intrinsically disordered segments in nuclear chromatin/transcription regulators—especially AT-hook–like RGR cores and RGG/GRG tracts—that provide nonspecific DNA minor-groove/chromatin association and targeting, typically in tails and linkers flanking structured domains. | 7.36 | 7 |
| #5310 | Catalytic Asp N-terminal strand | Short beta-strand or tight-loop elements that sit immediately N-terminal to catalytic metal-binding acidic residues (Asp/Glu) in nuclease active sites—especially RecB/RecB-like and VRR-NUC folds—and, more broadly, analogous beta-edge segments lining nucleotide/metal-binding pockets in other DNA-processing enzymes; peaks frequently center on glycine at a coil-to-beta transition. | 7.34 | 4 |
| #8462 | Acidic phospho-regulated IDR SLiMs | Short linear motifs embedded in intrinsically disordered, acidic Ser/Thr‑rich regions of nuclear proteins—often phospho‑regulated and frequently including SUMO‑interaction motifs—captured as hydrophobic cores (I/V/L/M) within low‑complexity acidic tracts that mediate transient interactions in DNA/RNA metabolism. | 6.91 | 9 |
| #3541 | Flexible solvent-exposed loops/linkers | Generic structural signal for short, flexible, solvent‑exposed coil/loop and linker residues—especially at secondary‑structure transitions and inter‑domain boundaries—and for disordered signal/propeptide segments; not tied to a specific sequence motif or function, explaining its ubiquity across taxa and protein families | 6.84 | 6 |
| #10977 | Early N-terminal disordered patch | Short N-terminal patch at the extreme N-terminus (around residues 5–10), capturing the first distinctive segment of preproteins/precursors and N-terminal disordered tails; residue composition is heterogeneous but frequently polar (Ser/Thr) and/or charged. | 6.83 | 3 |
| #8412 | Valine-centered residue detector | Detector that fires on valine residues across diverse sequence contexts, including hydrophobic transmembrane helices and signal peptides as well as Val-containing tandem repeats (elastomeric Gly-rich lamprin-like repeats, Pro/Arg-rich disordered repeats) and isolated Val positions in soluble globular proteins. | 6.82 | 4 |
| #14712 | SPS N-terminal regulatory motif | A conserved N-terminal regulatory motif in plant sucrose-phosphate synthases (SPS), centered on a short ATR(N/S/P)xRERS / SPQERN-type sequence upstream of the catalytic glycosyltransferase core. | 6.66 | 2 |
| #1501 | Acidic phosphosite-rich IDRs | Low-complexity intrinsically disordered segments enriched in polar residues with embedded acidic/phospho-regulatory clusters, serving as interaction modules across diverse proteins (nuclear/chromatin/RNA-associated and also secreted) | 6.55 | 11 |
| #13016 | Acidic low-complexity IDRs | Acidic intrinsically disordered low-complexity regions (IDRs)—often terminal tails or linkers—i.e., long D/E-rich acidic tracts commonly found in nuclear/chromatin proteins, viral tegument/IE phosphoproteins, and cytoskeletal regulators | 6.31 | 9 |
| #8237 | Conserved structured domain peaks | A feature that activates within structured partner-recognition/DNA-binding domains and adjacent ordered segments of eukaryotic regulatory proteins, with peaks falling on conserved residues inside these domains rather than on flanking low-complexity stretches. | 6.25 | 3 |
| #13745 | Hydrophobic secondary structure signal | Generic hydrophobic secondary-structure signal: activates on hydrophobic/aromatic side chains that form stable α-helices and β-strands—covering transmembrane helices, coiled-coils, and the hydrophobic cores of globular domains—while avoiding catalytic/metal-binding residues and disordered segments | 6.21 | 3 |
| #14076 | GARIN central LxLTR motif | A conserved central segment in Golgi-associated RAB2 interactor (GARIN) family proteins, characterized by a short "(L/V)xLTR"-like motif (e.g., LELTR, LVLTR) preceded by a basic/polar stretch (often K/R/S-rich). | 6.07 | 16 |
| #14234 | Ordered cores and transmembrane helices | Generic preference for ordered chain regions of folded proteins, with particularly strong responses in hydrophobic α-helices that span membranes; avoids signal peptides and low-confidence N-terminal precursor regions. | 5.93 | 2 |
| #10856 | Short domain-boundary linkers | Short, flexible coil/linker segments at or flanking structured domains—especially domain boundaries, inter‑domain linkers, and proteolytic processing junctions (signal peptides/propeptides)—with occasional peaks in loops within the domains themselves; not broadly marking long intrinsically disordered regions. | 5.84 | 3 |
| #8139 | Structured cores of genome regulators | Core residues of folded domains in eukaryotic genome-function proteins (chromatin-, DNA-, and RNA-associated factors, including transcriptional regulators, RNA-binding proteins, replication and repair factors), emphasizing beta-strand/loop elements of nucleic-acid–associated folds (e.g., B3 DNA-binding and RRM RNA-binding) and the PB1 oligomerization module; the feature marks conserved, well-structured domain cores and avoids long low-complexity/disordered regions. | 5.74 | 3 |
| #6178 | Nuclear basic/polar low-complexity IDRs | Intrinsically disordered, low‑complexity regulatory segments of nuclear proteins—especially transcription factors/cofactors, chromatin regulators, RNA/splicing factors, and nuclear RING‑type E3 ligases—enriched in basic (Lys/Arg) and polar (Ser/Pro/Gln/Glu/Gly/Thr) tracts that serve as activation/repression, partner‑binding, or localization (NLS‑like) regions; typically outside folded DNA‑binding/catalytic domains and often overlapping phosphorylation sites or simple helices within IDRs. | 5.70 | 7 |
| #9000 | Glutamate-biased acidic tract detector | Detector of glutamate identity and glutamate-enriched acidic tracts: strong activation on E residues, especially within acidic, low‑complexity/disordered regions; weaker, sporadic responses on D; largely domain-, function-, and taxonomy-agnostic. | 5.66 | 8 |
| #14895 | Unknown generic feature | Unknown generic feature | 20.17 | 23 |
| #9214 | Unknown generic feature | Unknown generic feature | 17.88 | 23 |
| #1803 | Unknown generic feature | Unknown generic feature | 17.30 | 23 |
| #14534 | Unknown generic feature | Unknown generic feature | 14.60 | 21 |
| #9194 | Unknown generic feature | Unknown generic feature | 13.19 | 21 |
| #9005 | Unknown generic feature | Unknown generic feature | 9.39 | 23 |
Top-K sparse-autoencoder activations on the ESM-C residual stream. Activation is a feature's peak value over its region; Prevalence is the number of residues it fires on; Δ activation is isoform − canonical. Only the top 30 features per category are saved; generic / "unknown" features are hidden until you expand a table. Read more about SAE features →
Clinical variants
Differential region — lost N-terminus (canonical-only)
| Pos (iso) | AA change | Consequence | Source | Clin. sig. | AF (gnomAD) | Impact | AlphaMissense | ESM-C ΔLLR | Link |
|---|---|---|---|---|---|---|---|---|---|
| — | V→V | intronic | gnomAD | — | 2.74e-06 | — | — | 0.00 | chr17-48076939-C-T |
| — | E→E | intronic | gnomAD | — | 2.05e-06 | — | — | 0.00 | chr17-48076945-T-C |
| — | E→E | intronic | gnomAD | — | 1.37e-06 | — | — | 0.00 | chr17-48076951-T-C |
| — | E→G | intronic | gnomAD | — | 6.84e-07 | damaging | likely_pathogenic (0.77) | -9.56 | chr17-48076952-T-C |
| — | E→Q | intronic | gnomAD | — | 1.37e-06 | damaging | likely_pathogenic (0.61) | -9.81 | chr17-48076956-C-G |
| — | E→Q | intronic | gnomAD | — | 6.84e-07 | damaging | likely_pathogenic (0.62) | -9.69 | chr17-48076959-C-G |
| — | E→K | intronic | gnomAD | — | 6.84e-07 | damaging | likely_pathogenic (0.83) | -11.12 | chr17-48076962-C-T |
| — | L→L | intronic | gnomAD | — | 2.05e-06 | — | — | 0.00 | chr17-48076963-T-C |
| — | V→V | intronic | gnomAD | — | 1.37e-06 | — | — | 0.00 | chr17-48076966-C-T |
| — | V→L | intronic | gnomAD | — | 2.05e-06 | damaging | likely_benign (0.30) | -7.78 | chr17-48076968-C-A |
| — | V→M | intronic | gnomAD | — | 2.74e-06 | damaging | likely_benign (0.24) | -8.68 | chr17-48076968-C-T |
| — | V→V | intronic | gnomAD | — | 1.37e-06 | — | — | 0.00 | chr17-48076975-C-T |
| — | V→G | intronic | gnomAD | — | 6.85e-07 | damaging | likely_benign (0.14) | -8.75 | chr17-48076976-A-C |
| — | V→M | intronic | gnomAD | — | 7.53e-06 | damaging | likely_benign (0.13) | -8.37 | chr17-48076977-C-T |
| — | K→N | intronic | gnomAD | — | 6.84e-07 | damaging | likely_pathogenic (0.58) | -9.00 | chr17-48076978-T-A |
| — | — | intronic | gnomAD | — | 1.37e-06 | — | — | — | chr17-48076978-TTTC-T |
| — | K→I | intronic | gnomAD | — | 6.84e-07 | damaging | ambiguous (0.56) | -11.69 | chr17-48076979-T-A |
| — | K→R | intronic | gnomAD | — | 9.58e-06 | damaging | likely_benign (0.12) | -8.94 | chr17-48076979-T-C |
| — | K→* | intronic | gnomAD | — | 6.84e-07 | damaging | — | — | chr17-48076980-T-A |
| — | K→E | intronic | gnomAD | — | 6.84e-07 | damaging | ambiguous (0.42) | -9.69 | chr17-48076980-T-C |
| — | K→N | intronic | gnomAD | — | 6.84e-07 | damaging | likely_pathogenic (0.84) | -9.87 | chr17-48076981-C-A |
| — | K→R | intronic | gnomAD | — | 6.84e-07 | damaging | likely_benign (0.12) | -9.06 | chr17-48076982-T-C |
| — | K→N | intronic | gnomAD | — | 6.84e-07 | damaging | likely_pathogenic (0.73) | -9.94 | chr17-48076984-C-G |
| — | — | intronic | gnomAD | — | 6.84e-07 | — | — | — | chr17-48076984-CTTG-C |
| — | K→Q | intronic | gnomAD | — | 1.37e-06 | damaging | likely_benign (0.24) | -10.25 | chr17-48076986-T-G |
| — | N→K | intronic | gnomAD | — | 1.85e-05 | damaging | likely_pathogenic (0.68) | -8.87 | chr17-48076987-G-C |
| — | Q→Q | intronic | gnomAD | — | 6.84e-07 | — | — | 0.00 | chr17-48076990-T-C |
| — | K→K | intronic | gnomAD | — | 2.74e-06 | — | — | 0.00 | chr17-48076993-T-C |
| — | G→G | intronic | gnomAD | — | 6.85e-07 | — | — | 0.00 | chr17-48076999-C-A |
| — | G→V | intronic | gnomAD | — | 6.85e-07 | damaging | likely_pathogenic (0.88) | -9.25 | chr17-48077000-C-A |
| — | G→W | intronic | gnomAD | — | 6.85e-07 | damaging | likely_pathogenic (0.97) | -12.31 | chr17-48077001-C-A |
| — | M→I | intronic | gnomAD | — | 1.37e-06 | damaging | — | -10.94 | chr17-48077002-C-G |
| — | M→I | intronic | gnomAD | — | 6.85e-07 | damaging | — | -10.94 | chr17-48077002-C-T |
| — | M→T | intronic | gnomAD | — | 6.85e-07 | damaging | — | -11.50 | chr17-48077003-A-G |
| — | M→L | intronic | gnomAD | — | 6.85e-07 | damaging | — | -11.19 | chr17-48077004-T-G |
| — | N→K | intronic | ClinVar | Uncertain significance | — | damaging | likely_pathogenic (0.68) | -8.87 | ClinVar:3827926 |
| — | V→V | intronic | COSMIC | — | — | — | — | 0.00 | COSV99848141 |
| — | E→Q | intronic | COSMIC | — | — | damaging | likely_pathogenic (0.95) | -11.25 | COSV99848313 |
| — | E→K | intronic | COSMIC | — | — | damaging | likely_pathogenic (0.83) | -9.56 | COSV56681509 |
| — | V→L | intronic | COSMIC | — | — | damaging | likely_benign (0.20) | -8.06 | COSV56682280 |
| — | Q→K | intronic | COSMIC | — | — | damaging | ambiguous (0.37) | -10.00 | COSV56682149 |
| — | — | intronic | COSMIC | — | — | damaging | — | — | COSV56681712 |
| — | — | intronic | COSMIC | — | — | damaging | — | — | COSV56681499 |
| — | M→V | intronic | COSMIC | — | — | damaging | — | -11.44 | COSV56682591 |
44 variants in the differential region.
Shared canonical core — sequence common to canonical and isoform; AlphaMissense applies here
| Pos (iso) | AA change | Consequence | Source | Clin. sig. | AF (gnomAD) | Impact | AlphaMissense | ESM-C ΔLLR | Link |
|---|---|---|---|---|---|---|---|---|---|
| 0 | V→V | synonymous_variant | gnomAD | — | 6.84e-07 | — | — | 0.00 | chr17-48076936-C-A |
| 0 | V→V | synonymous_variant | gnomAD | — | 6.84e-07 | — | — | 0.00 | chr17-48076936-C-G |
| 1 | E→V | missense_variant | ClinVar | — | — | damaging | likely_pathogenic (1.00) | -11.31 | ClinVar:4423332 |
| 2 | K→N | missense_variant | gnomAD | — | 6.84e-07 | damaging | likely_pathogenic (0.99) | -9.37 | chr17-48076930-T-A |
| 4 | L→L | synonymous_variant | gnomAD | — | 6.84e-06 | — | — | 0.00 | chr17-48076924-G-A |
| 4 | L→I | missense_variant | COSMIC | — | — | damaging | likely_benign (0.28) | -9.56 | COSV56682110 |
| 5 | D→H | missense_variant | gnomAD | — | 1.37e-06 | damaging | likely_pathogenic (0.94) | -7.89 | chr17-48076923-C-G |
| 5 | D→N | missense_variant | gnomAD | — | 1.85e-05 | damaging | likely_pathogenic (0.56) | -4.01 | chr17-48076923-C-T |
| 6 | R→H | missense_variant | gnomAD | — | 2.74e-06 | damaging | likely_pathogenic (0.79) | -7.03 | chr17-48076919-C-T |
| 6 | R→C | missense_variant | gnomAD | — | 1.23e-05 | damaging | likely_pathogenic (0.87) | -7.12 | chr17-48076920-G-A |
| 7 | R→Q | missense_variant | gnomAD | — | 2.05e-06 | damaging | likely_pathogenic (0.99) | -9.25 | chr17-48076916-C-T |
| 7 | R→R | synonymous_variant | gnomAD | — | 6.84e-07 | — | — | 0.00 | chr17-48076917-G-T |
| 7 | R→R | synonymous_variant | COSMIC | — | — | — | — | 0.00 | COSV56682056 |
| 7 | R→* | stop_gained | COSMIC | — | — | LoF | — | — | COSV56681749 |
| 8 | V→V | synonymous_variant | gnomAD | — | 2.05e-06 | — | — | 0.00 | chr17-48076912-C-T |
| 8 | V→A | missense_variant | gnomAD | — | 1.37e-06 | damaging | likely_pathogenic (0.92) | -7.43 | chr17-48076913-A-G |
| 9 | V→V | synonymous_variant | gnomAD | — | 1.37e-06 | — | — | 0.00 | chr17-48076909-T-C |
| 9 | V→L | missense_variant | gnomAD | — | 6.84e-07 | damaging | likely_pathogenic (0.76) | -6.96 | chr17-48076911-C-A |
| 10 | K→K | synonymous_variant | gnomAD | — | 6.84e-07 | — | — | 0.00 | chr17-48076906-C-T |
| 11 | G→G | synonymous_variant | gnomAD | — | 3.42e-06 | — | — | 0.00 | chr17-48076903-G-A |
| 11 | G→S | missense_variant | gnomAD | — | 6.84e-07 | damaging | likely_pathogenic (0.96) | -7.97 | chr17-48076905-C-T |
| 11 | G→S | missense_variant | COSMIC | — | — | damaging | likely_pathogenic (0.96) | -7.97 | COSV99847936 |
| 12 | K→R | missense_variant | gnomAD | — | 2.05e-06 | — | likely_benign (0.09) | -5.18 | chr17-48076901-T-C |
| 14 | E→D | missense_variant | gnomAD | — | 6.85e-07 | damaging | likely_pathogenic (1.00) | -7.72 | chr17-48076894-C-G |
| 14 | E→K | missense_variant | COSMIC | — | — | damaging | likely_pathogenic (1.00) | -10.19 | COSV56682856 |
| 15 | Y→Y | synonymous_variant | gnomAD | — | 1.71e-05 | — | — | 0.00 | chr17-48076891-G-A |
| 16 | L→L | synonymous_variant | gnomAD | — | 3.43e-06 | — | — | 0.00 | chr17-48076888-G-A |
| 16 | L→R | missense_variant | gnomAD | — | 6.85e-07 | damaging | likely_pathogenic (0.84) | -9.56 | chr17-48076889-A-C |
| 16 | L→F | missense_variant | COSMIC | — | — | — | ambiguous (0.35) | -6.53 | COSV56681814 |
| 17 | L→L | synonymous_variant | gnomAD | — | 6.85e-07 | — | — | 0.00 | chr17-48076885-T-G |
| 18 | K→K | synonymous_variant | gnomAD | — | 6.85e-07 | — | — | 0.00 | chr17-48076882-C-T |
| 18 | K→E | missense_variant | gnomAD | — | 6.85e-07 | damaging | likely_pathogenic (1.00) | -11.19 | chr17-48076884-T-C |
| 21 | G→G | synonymous_variant | gnomAD | — | 4.39e-05 | — | — | 0.00 | chr17-48076873-T-C |
| 22 | F→F | synonymous_variant | gnomAD | — | 6.86e-07 | — | — | 0.00 | chr17-48076870-G-A |
| 22 | F→L | missense_variant | COSMIC | — | — | damaging | likely_pathogenic (1.00) | -10.37 | COSV56682209 |
| 25 | E→K | missense_variant | gnomAD | — | 7.10e-07 | damaging | likely_pathogenic (0.84) | -9.50 | chr17-48076177-C-T |
| 25 | E→Q | missense_variant | ClinVar | — | — | damaging | likely_pathogenic (0.60) | -11.12 | ClinVar:4423331 |
| 25 | E→Q | missense_variant | COSMIC | — | — | damaging | likely_pathogenic (0.60) | -11.12 | COSV56681530 |
| 26 | D→H | missense_variant | gnomAD | — | 7.06e-07 | damaging | likely_pathogenic (0.97) | -12.12 | chr17-48076174-C-G |
| 27 | N→D | missense_variant | gnomAD | — | 1.41e-06 | damaging | likely_pathogenic (0.97) | -10.37 | chr17-48076171-T-C |
| 27 | N→S | missense_variant | COSMIC | — | — | damaging | likely_pathogenic (0.59) | -9.00 | COSV56681968 |
| 30 | E→D | missense_variant | COSMIC | — | — | damaging | likely_pathogenic (1.00) | -10.56 | COSV99848274 |
| 32 | E→D | missense_variant | COSMIC | — | — | damaging | likely_pathogenic (0.96) | -8.69 | COSV99848151 |
| 34 | N→N | synonymous_variant | gnomAD | — | 1.38e-06 | — | — | 0.00 | chr17-48076148-G-A |
| 35 | L→L | synonymous_variant | gnomAD | — | 1.38e-06 | — | — | 0.00 | chr17-48076147-G-A |
| 38 | P→P | synonymous_variant | gnomAD | — | 1.51e-05 | — | — | 0.00 | chr17-48076136-G-A |
| 38 | P→P | synonymous_variant | gnomAD | — | 6.88e-07 | — | — | 0.00 | chr17-48076136-G-C |
| 38 | P→P | synonymous_variant | COSMIC | — | — | — | — | 0.00 | COSV99847948 |
| 39 | D→N | missense_variant | COSMIC | — | — | damaging | likely_pathogenic (0.96) | -8.31 | COSV99847954 |
| 40 | L→L | synonymous_variant | gnomAD | — | 1.23e-05 | — | — | 0.00 | chr17-48076130-G-A |
| 42 | A→A | synonymous_variant | gnomAD | — | 5.48e-06 | — | — | 0.00 | chr17-48076124-A-G |
| 42 | A→A | synonymous_variant | gnomAD | — | 6.85e-07 | — | — | 0.00 | chr17-48076124-A-T |
| 42 | A→G | missense_variant | gnomAD | — | 1.37e-06 | damaging | ambiguous (0.38) | -9.94 | chr17-48076125-G-C |
| 45 | L→L | synonymous_variant | gnomAD | — | 1.37e-06 | — | — | 0.00 | chr17-48076115-C-T |
| 45 | L→L | synonymous_variant | gnomAD | — | 6.85e-07 | — | — | 0.00 | chr17-48076117-G-A |
| 45 | L→L | synonymous_variant | COSMIC | — | — | — | — | 0.00 | COSV105045650 |
| 46 | Q→* | stop_gained | COSMIC | — | — | LoF | — | — | COSV99848158 |
| 47 | S→L | missense_variant | gnomAD | — | 2.74e-06 | damaging | likely_benign (0.27) | -9.37 | chr17-48076110-G-A |
| 47 | S→L | missense_variant | ClinVar | Uncertain significance | — | damaging | likely_benign (0.27) | -9.37 | ClinVar:4648788 |
| 48 | Q→R | missense_variant | gnomAD | — | 6.85e-07 | — | likely_benign (0.33) | -7.47 | chr17-48076107-T-C |
| 48 | Q→* | stop_gained | COSMIC | — | — | LoF | — | — | COSV99848270 |
| 49 | K→R | missense_variant | gnomAD | — | 6.85e-07 | — | likely_benign (0.12) | -6.44 | chr17-48076104-T-C |
| 49 | K→Q | missense_variant | ClinVar | Uncertain significance | — | damaging | likely_benign (0.33) | -8.75 | ClinVar:4531663 |
| 50 | T→A | missense_variant | gnomAD | — | 6.85e-07 | — | likely_benign (0.06) | -6.06 | chr17-48076102-T-C |
| 52 | H→P | missense_variant | gnomAD | — | 4.79e-06 | — | likely_benign (0.11) | -7.21 | chr17-48076095-T-G |
| 54 | T→R | missense_variant | gnomAD | — | 2.74e-06 | damaging | likely_benign (0.23) | -7.90 | chr17-48076089-G-C |
| 55 | D→G | missense_variant | gnomAD | — | 6.84e-07 | — | likely_benign (0.10) | -6.27 | chr17-48076086-T-C |
| 55 | D→N | missense_variant | gnomAD | — | 6.84e-07 | — | likely_benign (0.10) | -6.15 | chr17-48076087-C-T |
| 55 | D→G | missense_variant | ClinVar | Uncertain significance | — | — | likely_benign (0.10) | -6.27 | ClinVar:4220041 |
| 55 | D→N | missense_variant | ClinVar | Uncertain significance | — | — | likely_benign (0.10) | -6.15 | ClinVar:4220044 |
| 56 | K→K | synonymous_variant | gnomAD | — | 4.04e-05 | — | — | 0.00 | chr17-48076082-T-C |
| 56 | K→I | missense_variant | gnomAD | — | 6.84e-07 | damaging | ambiguous (0.43) | -10.75 | chr17-48076083-T-A |
| 56 | K→I | missense_variant | ClinVar | Uncertain significance | — | damaging | ambiguous (0.43) | -10.75 | ClinVar:4220042 |
| 57 | S→S | synonymous_variant | gnomAD | — | 4.79e-06 | — | — | 0.00 | chr17-48076079-T-G |
| 58 | E→K | missense_variant | COSMIC | — | — | damaging | likely_benign (0.26) | -8.68 | COSV99848247 |
| 59 | G→R | missense_variant | gnomAD | — | 6.84e-07 | damaging | likely_pathogenic (0.58) | -8.76 | chr17-48076075-C-G |
| 59 | G→R | missense_variant | COSMIC | — | — | damaging | likely_pathogenic (0.58) | -8.76 | COSV56682528 |
| 62 | R→H | missense_variant | gnomAD | — | 7.53e-06 | damaging | likely_pathogenic (0.89) | -9.31 | chr17-48076065-C-T |
| 62 | R→C | missense_variant | gnomAD | — | 6.16e-06 | damaging | likely_pathogenic (0.96) | -9.31 | chr17-48076066-G-A |
| 62 | R→H | missense_variant | ClinVar | Uncertain significance | — | damaging | likely_pathogenic (0.89) | -9.31 | ClinVar:2290145 |
| 62 | R→H | missense_variant | COSMIC | — | — | damaging | likely_pathogenic (0.89) | -9.31 | COSV56682901 |
| 62 | R→C | missense_variant | COSMIC | — | — | damaging | likely_pathogenic (0.96) | -9.31 | COSV56682920 |
| 63 | K→R | missense_variant | gnomAD | — | 8.90e-06 | damaging | likely_benign (0.11) | -7.87 | chr17-48076062-T-C |
| 64 | A→T | missense_variant | gnomAD | — | 1.37e-06 | — | likely_benign (0.08) | -6.34 | chr17-48076060-C-T |
| 64 | A→T | missense_variant | ClinVar | Uncertain significance | — | — | likely_benign (0.08) | -6.34 | ClinVar:4220043 |
| 65 | D→G | missense_variant | gnomAD | — | 6.85e-07 | damaging | likely_benign (0.15) | -8.36 | chr17-48076056-T-C |
| 66 | S→A | missense_variant | gnomAD | — | 6.85e-07 | damaging | likely_benign (0.06) | -9.87 | chr17-48076054-A-C |
| 66 | S→T | missense_variant | gnomAD | — | 1.37e-06 | — | likely_benign (0.05) | -7.31 | chr17-48076054-A-T |
| 70 | D→V | missense_variant | gnomAD | — | 1.37e-06 | damaging | likely_benign (0.13) | -9.56 | chr17-48076041-T-A |
| 70 | D→V | missense_variant | ClinVar | Uncertain significance | — | damaging | likely_benign (0.13) | -9.56 | ClinVar:2307098 |
| 70 | D→N | missense_variant | COSMIC | — | — | damaging | likely_benign (0.11) | -7.93 | COSV56682563 |
| 71 | K→K | synonymous_variant | gnomAD | — | 1.37e-06 | — | — | 0.00 | chr17-48076037-C-T |
| 71 | K→R | missense_variant | gnomAD | — | 2.05e-06 | — | likely_benign (0.09) | -4.90 | chr17-48076038-T-C |
| 71 | K→E | missense_variant | gnomAD | — | 6.85e-07 | damaging | likely_benign (0.18) | -11.05 | chr17-48076039-T-C |
| 71 | K→K | synonymous_variant | COSMIC | — | — | — | — | 0.00 | COSV56682520 |
| 71 | K→N | missense_variant | COSMIC | — | — | damaging | likely_benign (0.29) | -9.30 | COSV99848170 |
| 72 | G→G | synonymous_variant | gnomAD | — | 2.06e-06 | — | — | 0.00 | chr17-48076034-T-C |
| 72 | — | frameshift_variant | gnomAD | — | 1.37e-06 | LoF | — | — | chr17-48076035-CCCTT-C |
| 72 | G→R | missense_variant | gnomAD | — | 2.74e-06 | damaging | ambiguous (0.54) | -7.55 | chr17-48076036-C-T |
| 73 | E→E | synonymous_variant | gnomAD | — | 6.86e-07 | — | — | 0.00 | chr17-48076031-C-T |
| 73 | E→Q | missense_variant | COSMIC | — | — | damaging | likely_benign (0.23) | -8.62 | COSV99848203 |
| 74 | E→D | missense_variant | gnomAD | — | 6.86e-07 | — | likely_benign (0.05) | -5.84 | chr17-48076028-C-G |
| 75 | S→S | synonymous_variant | gnomAD | — | 4.81e-06 | — | — | 0.00 | chr17-48076025-G-A |
| 75 | S→G | missense_variant | gnomAD | — | 6.86e-07 | — | likely_benign (0.06) | -5.94 | chr17-48076027-T-C |
| 75 | S→S | synonymous_variant | COSMIC | — | — | — | — | 0.00 | COSV99848275 |
| 77 | P→P | synonymous_variant | gnomAD | — | 2.07e-06 | — | — | 0.00 | chr17-48076019-T-C |
| 77 | P→L | missense_variant | gnomAD | — | 1.38e-06 | — | likely_benign (0.09) | -6.30 | chr17-48076020-G-A |
| 79 | K→N | missense_variant | gnomAD | — | 6.90e-07 | damaging | likely_pathogenic (0.93) | -10.37 | chr17-48076013-C-A |
| 79 | K→K | synonymous_variant | gnomAD | — | 1.38e-06 | — | — | 0.00 | chr17-48076013-C-T |
| 80 | K→N | missense_variant | gnomAD | — | 6.91e-07 | damaging | likely_pathogenic (0.91) | -10.25 | chr17-48076010-C-G |
| 80 | K→R | missense_variant | gnomAD | — | 1.38e-06 | — | likely_benign (0.09) | -5.37 | chr17-48076011-T-C |
| 80 | K→N | missense_variant | COSMIC | — | — | damaging | likely_pathogenic (0.91) | -10.25 | COSV56681336 |
| 80 | K→N | missense_variant | COSMIC | — | — | damaging | likely_pathogenic (0.91) | -10.25 | COSV99848297 |
| 81 | — | inframe_deletion | gnomAD | — | 2.77e-06 | — | — | — | chr17-48076007-TTTC-T |
| 81 | K→R | missense_variant | gnomAD | — | 6.92e-07 | — | likely_benign (0.09) | -6.75 | chr17-48076008-T-C |
| 82 | E→G | missense_variant | gnomAD | — | 6.92e-07 | damaging | likely_benign (0.33) | -8.81 | chr17-48076005-T-C |
| 83 | E→V | missense_variant | gnomAD | — | 1.39e-06 | damaging | likely_benign (0.30) | -8.24 | chr17-48076002-T-A |
| 83 | E→Q | missense_variant | gnomAD | — | 6.94e-07 | damaging | ambiguous (0.43) | -8.99 | chr17-48076003-C-G |
| 84 | S→S | synonymous_variant | gnomAD | — | 6.85e-07 | — | — | 0.00 | chr17-48075098-T-A |
| 84 | S→L | missense_variant | gnomAD | — | 6.85e-07 | — | likely_benign (0.08) | -6.09 | chr17-48075099-G-A |
| 85 | E→A | missense_variant | gnomAD | — | 6.85e-07 | damaging | ambiguous (0.50) | -9.43 | chr17-48075096-T-G |
| 87 | P→S | missense_variant | gnomAD | — | 6.84e-07 | — | ambiguous (0.53) | -5.76 | chr17-48075091-G-A |
| 87 | P→Q | missense_variant | COSMIC | — | — | damaging | likely_pathogenic (0.61) | -7.20 | COSV106082017 |
| 87 | P→S | missense_variant | COSMIC | — | — | — | ambiguous (0.53) | -5.76 | COSV56682997 |
| 88 | R→Q | missense_variant | gnomAD | — | 6.84e-07 | damaging | likely_pathogenic (0.92) | -8.69 | chr17-48075087-C-T |
| 88 | R→* | stop_gained | gnomAD | — | 2.05e-06 | LoF | — | — | chr17-48075088-G-A |
| 88 | R→* | stop_gained | COSMIC | — | — | LoF | — | — | COSV56682269 |
| 90 | F→S | missense_variant | gnomAD | — | 6.84e-07 | damaging | likely_pathogenic (1.00) | -10.44 | chr17-48075081-A-G |
| 91 | A→A | synonymous_variant | gnomAD | — | 1.37e-06 | — | — | 0.00 | chr17-48075077-A-C |
| 91 | A→A | synonymous_variant | gnomAD | — | 6.84e-07 | — | — | 0.00 | chr17-48075077-A-G |
| 92 | R→Q | missense_variant | gnomAD | — | 6.84e-07 | damaging | likely_pathogenic (0.99) | -8.12 | chr17-48075075-C-T |
| 92 | R→* | stop_gained | gnomAD | — | 6.84e-07 | LoF | — | — | chr17-48075076-G-A |
| 92 | R→R | synonymous_variant | COSMIC | — | — | — | — | 0.00 | COSV107309107 |
| 92 | R→L | missense_variant | COSMIC | — | — | damaging | likely_pathogenic (1.00) | -10.37 | COSV99847942 |
| 92 | R→Q | missense_variant | COSMIC | — | — | damaging | likely_pathogenic (0.99) | -8.12 | COSV56683114 |
| 92 | R→* | stop_gained | COSMIC | — | — | LoF | — | — | COSV104387334 |
| 93 | G→D | missense_variant | COSMIC | — | — | damaging | likely_pathogenic (0.89) | -9.62 | COSV105045685 |
| 95 | E→K | missense_variant | gnomAD | — | 6.84e-07 | damaging | likely_pathogenic (0.87) | -7.54 | chr17-48075067-C-T |
| 95 | E→E | synonymous_variant | COSMIC | — | — | — | — | 0.00 | COSV56681442 |
| 96 | P→P | synonymous_variant | gnomAD | — | 1.97e-04 | — | — | 0.00 | chr17-48075062-C-T |
| 96 | P→L | missense_variant | gnomAD | — | 7.52e-06 | damaging | likely_pathogenic (1.00) | -10.81 | chr17-48075063-G-A |
| 96 | P→P | synonymous_variant | COSMIC | — | — | — | — | 0.00 | COSV99848187 |
| 96 | P→Q | missense_variant | COSMIC | — | — | damaging | likely_pathogenic (1.00) | -12.12 | COSV56682499 |
| 96 | P→S | missense_variant | COSMIC | — | — | damaging | likely_pathogenic (1.00) | -9.69 | COSV99848079 |
| 97 | E→D | missense_variant | gnomAD | — | 1.37e-06 | damaging | likely_pathogenic (0.92) | -9.06 | chr17-48075059-C-A |
| 97 | E→E | synonymous_variant | gnomAD | — | 1.37e-06 | — | — | 0.00 | chr17-48075059-C-T |
| 97 | E→D | missense_variant | ClinVar | Uncertain significance | — | damaging | likely_pathogenic (0.92) | -9.06 | ClinVar:4648787 |
| 97 | E→K | missense_variant | COSMIC | — | — | damaging | likely_pathogenic (1.00) | -12.12 | COSV56681862 |
| 98 | R→R | synonymous_variant | gnomAD | — | 6.84e-07 | — | — | 0.00 | chr17-48075056-C-T |
| 98 | R→Q | missense_variant | COSMIC | — | — | damaging | likely_pathogenic (0.98) | -7.87 | COSV56682101 |
| 98 | R→L | missense_variant | COSMIC | — | — | damaging | likely_pathogenic (0.99) | -10.12 | COSV99848174 |
| 99 | I→V | missense_variant | gnomAD | — | 6.84e-07 | damaging | likely_pathogenic (0.96) | -7.69 | chr17-48075055-T-C |
| 101 | G→A | missense_variant | COSMIC | — | — | damaging | likely_pathogenic (1.00) | -12.44 | COSV56682437 |
| 103 | T→T | synonymous_variant | gnomAD | — | 2.74e-06 | — | — | 0.00 | chr17-48075041-T-C |
| 105 | S→S | synonymous_variant | gnomAD | — | 6.84e-07 | — | — | 0.00 | chr17-48075035-G-A |
| 106 | S→S | synonymous_variant | COSMIC | — | — | — | — | 0.00 | COSV56681832 |
| 108 | E→Q | missense_variant | gnomAD | — | 6.84e-07 | damaging | likely_pathogenic (0.92) | -10.00 | chr17-48075028-C-G |
| 109 | L→L | synonymous_variant | gnomAD | — | 9.59e-06 | — | — | 0.00 | chr17-48075023-G-A |
| 110 | M→L | missense_variant | gnomAD | — | 6.85e-07 | damaging | likely_pathogenic (0.91) | -10.25 | chr17-48075022-T-A |
| 110 | M→V | missense_variant | gnomAD | — | 6.85e-07 | damaging | likely_pathogenic (0.99) | -11.50 | chr17-48075022-T-C |
| 110 | — | inframe_deletion | COSMIC | — | — | — | — | — | COSV56681204 |
| 111 | F→F | synonymous_variant | gnomAD | — | 6.85e-07 | — | — | 0.00 | chr17-48075017-G-A |
| 112 | L→P | missense_variant | gnomAD | — | 6.85e-07 | damaging | likely_pathogenic (1.00) | -13.06 | chr17-48075015-A-G |
| 112 | L→L | synonymous_variant | COSMIC | — | — | — | — | 0.00 | COSV99848148 |
| 112 | L→L | synonymous_variant | COSMIC | — | — | — | — | 0.00 | COSV56681918 |
| 113 | M→I | missense_variant | gnomAD | — | 6.86e-07 | damaging | likely_pathogenic (0.97) | -8.12 | chr17-48075011-C-A |
| 113 | M→T | missense_variant | gnomAD | — | 6.86e-07 | damaging | likely_pathogenic (1.00) | -11.81 | chr17-48075012-A-G |
| 114 | K→K | synonymous_variant | gnomAD | — | 6.86e-07 | — | — | 0.00 | chr17-48075008-T-C |
| 114 | K→T | missense_variant | gnomAD | — | 6.86e-07 | damaging | likely_pathogenic (1.00) | -11.87 | chr17-48075009-T-G |
| 116 | K→R | missense_variant | gnomAD | — | 8.28e-06 | — | likely_benign (0.14) | -6.94 | chr17-48071577-T-C |
| 117 | N→K | missense_variant | COSMIC | — | — | damaging | likely_pathogenic (0.89) | -8.94 | COSV99848229 |
| 118 | S→F | missense_variant | COSMIC | — | — | damaging | likely_pathogenic (0.99) | -10.37 | COSV56681364 |
| 121 | A→A | synonymous_variant | gnomAD | — | 6.87e-07 | — | — | 0.00 | chr17-48071561-A-G |
| 122 | D→D | synonymous_variant | gnomAD | — | 6.87e-07 | — | — | 0.00 | chr17-48071558-G-A |
| 124 | V→L | missense_variant | gnomAD | — | 6.86e-07 | damaging | likely_pathogenic (1.00) | -11.25 | chr17-48071554-C-G |
| 126 | A→A | synonymous_variant | gnomAD | — | 6.85e-07 | — | — | 0.00 | chr17-48071546-G-A |
| 126 | A→D | missense_variant | gnomAD | — | 6.85e-07 | damaging | likely_pathogenic (1.00) | -12.00 | chr17-48071547-G-T |
| 126 | A→V | missense_variant | COSMIC | — | — | damaging | likely_pathogenic (1.00) | -10.62 | COSV99848117 |
| 128 | E→* | stop_gained | gnomAD | — | 6.85e-07 | LoF | — | — | chr17-48071542-C-A |
| 128 | E→K | missense_variant | COSMIC | — | — | damaging | likely_pathogenic (0.99) | -10.94 | COSV56681808 |
| 130 | N→S | missense_variant | gnomAD | — | 6.84e-07 | damaging | likely_pathogenic (0.91) | -12.44 | chr17-48071535-T-C |
| 131 | V→F | missense_variant | gnomAD | — | 2.74e-06 | damaging | likely_pathogenic (0.57) | -8.73 | chr17-48071533-C-A |
| 131 | V→V | synonymous_variant | COSMIC | — | — | — | — | 0.00 | COSV56681933 |
| 132 | K→K | synonymous_variant | COSMIC | — | — | — | — | 0.00 | COSV56682664 |
| 133 | C→C | synonymous_variant | gnomAD | — | 1.37e-06 | — | — | 0.00 | chr17-48071525-G-A |
| 133 | C→S | missense_variant | gnomAD | — | 6.84e-07 | damaging | likely_pathogenic (1.00) | -5.97 | chr17-48071526-C-G |
| 134 | P→L | missense_variant | gnomAD | — | 6.84e-07 | damaging | likely_pathogenic (1.00) | -10.62 | chr17-48071523-G-A |
| 135 | Q→Q | synonymous_variant | gnomAD | — | 6.84e-07 | — | — | 0.00 | chr17-48071519-C-T |
| 135 | Q→* | stop_gained | gnomAD | — | 6.84e-07 | LoF | — | — | chr17-48071521-G-A |
| 136 | V→F | missense_variant | gnomAD | — | 2.05e-06 | damaging | likely_pathogenic (0.99) | -12.05 | chr17-48071518-C-A |
| 138 | I→I | synonymous_variant | COSMIC | — | — | — | — | 0.00 | COSV56681722 |
| 138 | I→T | missense_variant | COSMIC | — | — | damaging | likely_pathogenic (1.00) | -11.06 | COSV56682067 |
| 139 | S→S | synonymous_variant | gnomAD | — | 2.05e-06 | — | — | 0.00 | chr17-48071507-G-T |
| 140 | F→F | synonymous_variant | gnomAD | — | 6.84e-07 | — | — | 0.00 | chr17-48071504-G-A |
| 141 | Y→Y | synonymous_variant | gnomAD | — | 1.37e-06 | — | — | 0.00 | chr17-48071501-A-G |
| 141 | Y→C | missense_variant | gnomAD | — | 6.84e-07 | damaging | likely_pathogenic (1.00) | -11.94 | chr17-48071502-T-C |
| 144 | R→K | missense_variant | COSMIC | — | — | damaging | likely_pathogenic (0.99) | -11.06 | COSV56682552 |
| 146 | T→T | synonymous_variant | gnomAD | — | 4.11e-06 | — | — | 0.00 | chr17-48071486-C-T |
| 146 | T→K | missense_variant | COSMIC | — | — | damaging | likely_pathogenic (0.99) | -13.94 | COSV56681836 |
| 148 | H→H | synonymous_variant | gnomAD | — | 6.84e-07 | — | — | 0.00 | chr17-48071480-A-G |
| 149 | S→S | synonymous_variant | gnomAD | — | 6.84e-07 | — | — | 0.00 | chr17-48071477-G-A |
| 149 | S→F | missense_variant | gnomAD | — | 6.84e-07 | damaging | likely_pathogenic (0.97) | -11.12 | chr17-48071478-G-A |
| 149 | S→A | missense_variant | gnomAD | — | 6.84e-07 | damaging | likely_benign (0.27) | -9.06 | chr17-48071479-A-C |
| 149 | S→P | missense_variant | gnomAD | — | 6.84e-07 | damaging | likely_pathogenic (0.98) | -11.81 | chr17-48071479-A-G |
| 149 | S→T | missense_variant | gnomAD | — | 6.84e-07 | damaging | ambiguous (0.44) | -10.19 | chr17-48071479-A-T |
| 150 | Y→Y | synonymous_variant | gnomAD | — | 2.94e-05 | — | — | 0.00 | chr17-48071474-G-A |
| 150 | Y→* | stop_gained | gnomAD | — | 6.85e-07 | LoF | — | — | chr17-48071474-G-T |
| 151 | P→S | missense_variant | gnomAD | — | 6.85e-07 | damaging | likely_pathogenic (0.58) | -8.75 | chr17-48071473-G-A |
| 151 | P→S | missense_variant | COSMIC | — | — | damaging | likely_pathogenic (0.58) | -8.75 | COSV108798400 |
| 152 | S→S | synonymous_variant | gnomAD | — | 4.11e-06 | — | — | 0.00 | chr17-48071468-C-T |
| 152 | S→L | missense_variant | gnomAD | — | 2.05e-06 | damaging | likely_benign (0.21) | -8.42 | chr17-48071469-G-A |
| 152 | S→L | missense_variant | ClinVar | Uncertain significance | — | damaging | likely_benign (0.21) | -8.42 | ClinVar:3138046 |
| 152 | S→S | synonymous_variant | COSMIC | — | — | — | — | 0.00 | COSV56681272 |
| 152 | S→* | stop_gained | COSMIC | — | — | LoF | — | — | COSV99848014 |
| 153 | E→K | missense_variant | COSMIC | — | — | damaging | likely_pathogenic (0.75) | -12.00 | COSV104555665 |
| 154 | D→E | missense_variant | gnomAD | — | 4.80e-06 | — | likely_benign (0.09) | -5.24 | chr17-48071462-A-C |
| 154 | D→D | synonymous_variant | gnomAD | — | 6.85e-07 | — | — | 0.00 | chr17-48071462-A-G |
| 154 | D→Y | missense_variant | gnomAD | — | 6.85e-07 | damaging | likely_pathogenic (0.71) | -11.56 | chr17-48071464-C-A |
| 154 | D→N | missense_variant | gnomAD | — | 1.37e-06 | damaging | likely_benign (0.31) | -10.18 | chr17-48071464-C-T |
| 154 | D→E | missense_variant | ClinVar | Uncertain significance | — | — | likely_benign (0.09) | -5.24 | ClinVar:2517719 |
| 156 | — | inframe_deletion | gnomAD | — | 2.06e-06 | — | — | — | chr17-48071456-GTCA-G |
| 156 | D→N | missense_variant | COSMIC | — | — | damaging | likely_benign (0.19) | -8.62 | COSV99848218 |
| 157 | K→N | missense_variant | COSMIC | — | — | damaging | likely_pathogenic (0.72) | -9.62 | COSV99848111 |
| 157 | — | frameshift_variant | COSMIC | — | — | LoF | — | — | COSV56682928 |
| 158 | K→* | stop_gained | gnomAD | — | 6.85e-07 | LoF | — | — | chr17-48071452-T-A |
| 159 | D→V | missense_variant | gnomAD | — | 6.86e-07 | damaging | ambiguous (0.45) | -9.47 | chr17-48071448-T-A |
| 159 | — | frameshift_variant | gnomAD | — | 6.86e-07 | LoF | — | — | chr17-48071449-C-CT |
| 159 | — | frameshift_variant | gnomAD | — | 6.86e-07 | LoF | — | — | chr17-48071449-CTTTT-C |
| 159 | D→H | missense_variant | COSMIC | — | — | damaging | likely_pathogenic (0.73) | -10.72 | COSV56682228 |
| 159 | D→Y | missense_variant | COSMIC | — | — | damaging | likely_pathogenic (0.65) | -9.97 | COSV56681987 |
| 160 | — | frameshift_variant | gnomAD | — | 6.86e-07 | LoF | — | — | chr17-48071445-TC-T |
| 160 | D→Y | missense_variant | gnomAD | — | 1.37e-06 | damaging | likely_pathogenic (0.68) | -10.43 | chr17-48071446-C-A |
| 160 | D→N | missense_variant | gnomAD | — | 6.86e-07 | — | likely_benign (0.27) | -7.43 | chr17-48071446-C-T |
| 160 | — | frameshift_variant | gnomAD | — | 6.86e-07 | LoF | — | — | chr17-48071446-CA-C |
| 162 | N→K | missense_variant | gnomAD | — | 6.87e-07 | damaging | likely_pathogenic (0.68) | -7.21 | chr17-48071438-G-C |
| 162 | — | frameshift_variant | gnomAD | — | 1.37e-06 | LoF | — | — | chr17-48071439-TTCTTG-T |
| 162 | N→K | missense_variant | COSMIC | — | — | damaging | likely_pathogenic (0.68) | -7.21 | COSV56682095 |
237 variants in the shared canonical core.
LoF = frameshift / stop-gain / splice-disrupting — inherently loss-of-function, flagged by consequence (AlphaMissense and ESM-C score only missense/substitutions, so they are blank here by design, not by absence of impact). AlphaMissense is computed in the canonical reading frame, so it scores the shared core but reads N/A across an isoform-unique extension — use ESM-C ΔLLR there.