CBX1
TRUNCATED 143 aa (canonical 185 aa) · UniProt P83916 · CDLMPS
chr17:48076878:-:AAG:ENST00000225603.9
AI summary Truncation removes the N-terminal half of the chromodomain that reads H3K9me3, the core molecular function of HP1-beta.
This truncation deletes the N-terminal 43 residues of canonical CBX1, and domain annotation shows this segment overlaps the chromodomain itself (CATH chromo fold, SMART/Pfam/PROSITE chromo domain, PRINTS chromo domain subgroup) rather than flanking sequence — three independent domain calls diverge in this region. Localization signal (nuclear import, DeepLoc) is unaffected, and whole-protein biophysical character and SAE feature shifts are both below threshold, so this is not a general destabilization but a specific excision of part of the reader module.
CBX1/HP1-beta's defining activity is its chromodomain reading H3K9me3 to organize heterochromatin and recruit SUV39H1, PRC2, and DNA-damage machinery; removing roughly half of that domain would be expected to impair or abolish methyl-lysine reading, which is the molecular event this protein exists to perform. Because nuclear targeting is preserved, any truncated protein made from this isoform would likely still enter the nucleus but arrive without a functional H3K9me3-reading chromodomain, disconnecting it from the heterochromatin-binding and PRC2-recruitment activities described in the literature.
Structural confidence for this region is low (pTM <0.4, high PAE between removed and retained segments), so the fold-based P1/P2 signals could not corroborate the domain-annotation-based loss; the conclusion rests primarily on InterPro domain-boundary overlap rather than confirmed 3D disruption. Germline tolerance and disease-variant density in this region are unremarkable, which tempers confidence that this loss is clinically impactful even though it is conserved and translated.
Folding
Evidence — click any tile for the differential-region detail
LLM reasoning
How well is the isoform-unique region conserved at the amino-acid level across primates (mean percent identity)? High AA identity in close relatives argues the alternative protein is real and translated, not a sequencing or annotation artifact. (Reading-frame intactness is shown as context.)
Comparison
Sequence conservation · primate
mean AA identity over the differential ORF is the score basis; frame intactness across the clade is context
| Conservation metric | Canonical ORF | Differential ORF |
|---|---|---|
| Mean AA identity | 100% | 100% |
| Frame intact (fraction of species) | 84% | 96% |
| Species aligned | 25 | 25 |
| Species frame-intact | 21 | 24 |
| Start codon conserved | 100% | 100% |
| Deepest intact species | Propithecus_coquereli | Microcebus_murinus |
| Phylo depth (MRCA) | 7 | 7 |
The same amino-acid identity test across mammals. Conservation over deeper evolutionary distance is stronger evidence the ORF is under selection to be translated. (Reading-frame intactness is shown as context.)
Comparison
Sequence conservation · mammalian
mean AA identity over the differential ORF is the score basis; frame intactness across the clade is context
| Conservation metric | Canonical ORF | Differential ORF |
|---|---|---|
| Mean AA identity | 98% | 100% |
| Frame intact (fraction of species) | 15% | 100% |
| Species aligned | 20 | 20 |
| Species frame-intact | 3 | 20 |
| Start codon conserved | 100% | 100% |
| Deepest intact species | Callithrix_jacchus | Loxodonta_africana |
| Phylo depth (MRCA) | 6 | 12 |
phyloP / phastCons measure per-base evolutionary constraint. A high absolute phyloP over the differential region means the sequence itself is under strong purifying selection — evidence it is coding and does something. (Enrichment over the shared core is shown as context only.)
Comparison
Per-base conservation · differential region
absolute mean phyloP over the differential region is the score basis (high = strong purifying selection); the shared column and enrichment are context, not the claim
| Track | Differential | Shared | vs shared |
|---|---|---|---|
| phyloP mean | 4.92 | 5.02 | 0.979 |
| phastCons mean | 0.96 | 0.907 | — |
Start site · canonical vs isoform
initiation context + per-base conservation at each start codon
| Property | Canonical | Isoform |
|---|---|---|
| Start codon | ATG | AAG |
| Kozak context (−9..+4) | GCGGGCACTATGG | CTAAAGTGGAAGG |
| phyloP at start codon | 6.43 | 3.9 |
| phastCons at start codon | 1 | 1 |
| phyloP over Kozak window | 6.31 | 5.08 |
| phastCons over Kozak window | 1 | 0.979 |
| Kozak mismatch — full consensus | 3 | 10 |
| Kozak window GC content | 0.692 | 0.462 |
LLM reasoning
Is the alternative start used in more than one cell line? Reproducible initiation across independent samples argues against a one-off ribosome artifact.
Comparison
Per-cell-line usage · canonical vs isoform
Fisher q = 6.37e-14
| Cell line | Canonical | Isoform | p-value |
|---|---|---|---|
| HeLa | — | 21.1 | 6.37e-14 |
| U2OS | — | 1.4 | 0.00116 |
| RPE1 Async | — | 1.84 | 8.99e-05 |
| RPE1 Que | — | 0.316 | 0.00177 |
How efficiently ribosomes initiate at this start (TIS) relative to background, per cell line. Strong, reproducible initiation supports a genuine translation event.
Comparison
Start-Site Usage · canonical vs isoform
ribosome initiation efficiency at the canonical start vs this alternative start, per cell line
| Cell line | Canonical | Isoform |
|---|---|---|
| HeLa | — | 0.577 |
| U2OS | — | 0.0215 |
| RPE1 Async | — | 0.0437 |
| RPE1 Que | — | 0.0155 |
Were tryptic peptides unique to the differential region detected by mass spec? Direct peptide evidence is the strongest proof the isoform protein exists.
Comparison
Peptide Evidence (canonical vs isoform)
isoform-unique peptides are direct evidence the alternative protein exists
| Feature | Canonical | Isoform |
|---|---|---|
| Tryptic peptides (in-silico) | 28 | 2 |
| Validated by mass-spec | 0 | 1 |
| Isoform-unique peptides | — | 1 |
Details
Peptide Evidence (canonical vs isoform)
- validated MGFSDEDNTWEPEENLDCPDLIAEFLQSQK 0–30
LLM reasoning
Do the predicted localization features (DeepLoc prediction, sorting signals, or membrane association) differ between the canonical and the isoform? A protein with changed localization features acts in a different cellular context.
Comparison
Subcellular localization (DeepLoc)
no localization-feature change
| Property | Canonical | Isoform |
|---|---|---|
| Predicted location | Nucleus | Nucleus |
| Sorting signals | Nuclear localization signal | Nuclear localization signal |
| Membrane | Soluble | Soluble |
Do N-terminal targeting signals (SignalP secretion, TargetP mitochondrial/chloroplast) differ between canonical and isoform? N-terminal changes most directly add or remove targeting peptides.
LLM reasoning
Does healthy human germline variation (gnomAD) avoid this region (depletion ratio < 1×), and is it intrinsically constrained (ESM-C constraint delta > 0)? Depletion of population variation plus high sequence constraint mean the region resists change — it is functionally important. Scored only where the unique region is canonical coding sequence, i.e. on truncations. gnomAD is a tolerance catalogue, not a disease one; disease/cancer variants (ClinVar / COSMIC) live in M2.
Comparison
Germline Variants · differential vs shared region
gnomAD (population) variant density per nucleotide — depletion ratio < 1× means healthy human variation avoids the region (constrained), the M1 basis alongside ESM-C constraint
| Variant set | Differential | Shared | Depletion ratio |
|---|---|---|---|
| gnomAD variants | 60 | 132 | 1.5× |
Sequence constraint (ESM-C) · differential vs shared region
lower LLR = more constrained; enrichment = differential / conserved
| Property | Differential | Shared | Enrichment |
|---|---|---|---|
| ESM-C mean LLR | -0.00211 | -0.00855 | — |
| Constrained positions | 0 | 0 | — |
Predicted-damaging variants · germline (gnomAD)
AlphaMissense / ESM-C / LoF calls per region (length-normalized)
| Variant set | Differential | Shared | Enrichment |
|---|---|---|---|
| Scorable variants | 58 | 130 | 1.5× |
| Damaging variants | 36 | 67 | 1.8× |
| — of which loss-of-function | 1 | 12 | 0.28× |
| AlphaMissense-pathogenic | 22 | 36 | 2× |
Predictor scores · germline (gnomAD)
scored: 254 ESM-C · 168 AlphaMissense
| Property | Differential | Shared |
|---|---|---|
| Mean ΔLLR (ESM-C) | -5.75 | -5.59 |
| Min ΔLLR (ESM-C) | -12.3 | -13.1 |
| Mean AlphaMissense | 0.641 | 0.551 |
Are disease (ClinVar / COSMIC) variants enriched per nucleotide in the differential region versus the shared core (disease enrichment ratio ≥ 1×)? Disease variants concentrating in the unique region tie it to phenotype.
Comparison
Clinical-variant burden · differential vs shared region
counts per region; ratio is density-normalized (variants per nucleotide, differential ÷ shared — ≥1× = disease variants concentrate in the differential region, the M2 basis)
| Variant set | Differential | Shared | Enrichment |
|---|---|---|---|
| Disease variants | 16 | 73 | 0.75× |
| Pathogenic | 0 | 0 | — |
Predicted-damaging variants · disease (ClinVar/COSMIC)
AlphaMissense / ESM-C / LoF calls per region (length-normalized)
| Variant set | Differential | Shared | Enrichment |
|---|---|---|---|
| Scorable variants | 16 | 72 | 0.74× |
| Damaging variants | 13 | 53 | 0.82× |
| — of which loss-of-function | 3 | 6 | 1.7× |
| AlphaMissense-pathogenic | 6 | 37 | 0.54× |
Predictor scores · disease (ClinVar/COSMIC)
scored: 254 ESM-C · 168 AlphaMissense
| Property | Differential | Shared |
|---|---|---|
| Mean ΔLLR (ESM-C) | -8.06 | -7.43 |
| Min ΔLLR (ESM-C) | -11.4 | -13.9 |
| Mean AlphaMissense | 0.66 | 0.683 |
LLM reasoning
Does the gained/lost region fold into ordered structure (pLDDT) and shift the protein's biophysical character (charge, hydropathy, disorder)? Structured, biophysically distinct regions are more likely to be functional.
Comparison
Structure (ESMFold2) · canonical vs isoform
TM-score 0.419 · RMSD 2.57 Å
| Metric | Canonical | Isoform |
|---|---|---|
| pLDDT (whole protein) | 0.822 | 0.579 |
Fold confidence · differential vs shared region
mean ESMFold2 pLDDT in each region
| Metric | Differential | Shared | Enrichment |
|---|---|---|---|
| pLDDT | 0.809 | 1.06 | 0.77 |
The shared region is the stretch of protein identical in both the isoform and the canonical (the canonical body for an extension; the post-truncation body for a truncation). Folded in both contexts it normally comes out nearly identical, so its Cα backbone RMSD ≈ 0. A high shared-region RMSD (superposed on the shared residues only, from the ESMFold2 structures) means the extension or truncation reorganizes how the retained region folds — a rare, high-interest functional signal. TM-score is a length-normalized companion; the check only fires when both structures are confidently folded, and uORF/altORF isoforms (no shared region) are not evaluable.
Comparison
Shared-region structural change
Cα RMSD superposed on the shared residues only; TM-score is length-normalized · shared-region Cα RMSD 30.1 Å · shared TM-score 0.472 · shared region 142 aa · min shared pLDDT 0.58 · global TM-score 0.419 · global RMSD 2.57 Å
| Property | Canonical | Isoform |
|---|---|---|
| Shared-region pLDDT | 0.827 | 0.58 |
Does the differential region contain actual secondary structure — a helix or strand — rather than coil? ESMFold2 predicts coordinates but no secondary structure, so elements are assigned from those coordinates (P-SEA, Cα geometry). Because they are derived from a prediction, each element carries its own mean pLDDT: a geometrically clean helix running through a disordered stretch is geometry fitted to a guess, so BOTH length and confidence are required to score. The direction differs by ORF type — an extension GAINS the element (a candidate functional addition), a truncation LOSES one, and there the element is read off the canonical structure because the removed segment exists only in it. This says nothing about whether the element is integrated with the rest of the fold; the contact and PAE evidence in this same category answers that.
Comparison
Secondary structure · differential vs shared region
helices and strands assigned from the predicted coordinates (P-SEA)
| Metric | Differential | Shared |
|---|---|---|
| Alpha helices | 0 | 3 |
| Beta strands | 0 | 2 |
| Longest element (aa) | 0 | 15 |
| Mean pLDDT | — | 0.88 |
Elements and coordinates
0 in the differential region, 5 in the shared core — residue numbering is 1-based on the protein holding the region
| Removed (canonical) | Shared core |
|---|---|
| — | alpha helix 61–75 15 aa · pLDDT 0.82 |
| — | beta strand 86–92 7 aa · pLDDT 0.79 |
| — | beta strand 143–148 6 aa · pLDDT 0.90 |
| — | alpha helix 149–154 6 aa · pLDDT 0.93 |
| — | alpha helix 157–166 10 aa · pLDDT 0.94 |
Below threshold
1 in the differential region, 3 in the shared core — shorter than 6 aa or below pLDDT 0.70, so not counted above
| Removed (canonical) | Shared core |
|---|---|
| beta strand 35–39 5 aa · pLDDT 0.92 | beta strand 50–53 4 aa · pLDDT 0.92 |
| — | beta strand 96–100 5 aa · pLDDT 0.74 |
| — | beta strand 168–172 5 aa · pLDDT 0.83 |
LLM reasoning
Does the isoform gain or lose a real InterPro functional domain in the differential region (disorder/structural-only signatures excluded)? Gaining or losing a domain changes function directly.
Comparison
Domains & motifs (canonical vs isoform)
gained = only in the isoform; lost = only in the canonical
| Feature | Canonical | Isoform |
|---|---|---|
| InterPro domains | 15 | 11 |
| Short linear motifs | 1 | 1 |
Details
Domains & motifs (canonical vs isoform)
- lost Chromobox protein homolog 1 domain
- lost PR00504 domain
- lost Chromo domain signature domain
- lost chromodomain of heterochromatin protein 1 homolog beta domain
Biophysical character of the isoform-differential region versus the shared canonical core — pI, hydropathy, charge, disorder and related properties. Descriptive; the folding (P1) score keys off the GRAVY / charge / disorder deltas.
Comparison
Biophysics · differential vs shared region
highlighted rows are enriched in the differential region
| Property | Differential | Shared | Enrichment |
|---|---|---|---|
| Isoelectric point (pI) | 9.32 | 4.32 | 2.16 |
| Hydropathy (GRAVY) | -1.1 | -1.31 | 0.835 |
| Fraction charged | 0.535 | 0.434 | 1.23 |
| Disorder fraction | 0.229 | 0.228 | 1.01 |
| Disorder-promoting | 0.605 | 0.699 | 0.865 |
| Low-complexity fraction | 0.512 | 0.133 | 3.85 |
| Prion-like fraction | 0.14 | 0.217 | 0.643 |
| LLPS score | 0.202 | 0.187 | 1.08 |
| π–π propensity | 0.163 | 0.175 | 0.931 |
| Aromaticity | 0.0698 | 0.0699 | 0.999 |
| Instability index | 63.3 | 50.2 | 1.26 |
| Shannon entropy | 3 | 3.96 | 0.758 |
| Normalized complexity | 0.694 | 0.915 | 0.758 |
Sparse-autoencoder (SAE) interpretability features on the ESM-C residual stream, comparing the isoform against the canonical protein. This is the S3 criterion — a presence check that fires when any interpretable feature is gained or lost. Only the top 30 features per category (ranked by activation, prevalence ≥ 2) are saved; each table shows the top 10 interpretable features, and generic / "unknown" features are hidden unless you expand it. What are SAE features? →
Part 1 · Global feature comparison
Which SAE features fire in the isoform vs the canonical protein, across the whole sequence.
Isoform-only features — 34 total (prevalence ≥ 2)
| Feature | Label | Description | Activation | Prevalence |
|---|---|---|---|---|
| #787 | Basic disordered microprotein regions | The feature activates within small proteins, viral accessory/regulatory proteins, and short ORFs, often on disordered or low-complexity segments. Activation is frequently associated with basic-residue–containing patches (Lys/Arg) but is not restricted to a particular position in the protein. | 2.26 | 2 |
| #1745 | Sugar nucleotide pocket hydrophobe | A conserved hydrophobic position in a structured secondary element of the catalytic/binding core that shapes the sugar(‑nucleotide) binding/transfer pocket of carbohydrate‑active proteins—predominantly GT‑A/GT‑B glycosyltransferases and sugar‑1‑phosphate nucleotidyltransferases—plus analogous sites in solute/carbohydrate‑binding proteins and occasional nucleotide‑handling regulators (e.g., eIF2B subunits). In GT‑C membrane enzymes, the position lies within a catalytic transmembrane helix lining the active cavity. | 2.17 | 2 |
| #3508 | Alpha-helix boundary/capping residues | Alpha-helix boundary/capping residues: the feature marks residues at α-helix termini (especially C-caps) and immediately adjacent coil linkers, with occasional spillover into short stretches within the helix; it is a broadly structural signal independent of protein family and often falls near, but not on, catalytic sites | 2.16 | 2 |
| #9154 | Extracellular beta-strand/loop repeats | Short beta-strand segments and strand–loop junctions that repeat across beta-rich extracellular domains—especially beta‑propeller blades and related repeat scaffolds—common to secreted carbohydrate-active enzymes and adhesion/ECM receptors; the same structural element is also detected in analogous beta-strand→coil sites of soluble enzymes | 2.13 | 2 |
| #10487 | Nuclear regulatory low-complexity IDRs | Intrinsically disordered, low‑complexity regulatory regions of large metazoan regulatory and scaffolding proteins—enriched in nuclear gene‑regulatory factors—especially transactivation domains, flexible linkers, and long regulatory tails that are S/P/Q/G/T- and acidic‑rich, harbor SP/TP phospho‑motifs and short homorepeats; the feature preferentially avoids compact folded DNA-/ligand-/zinc‑binding domains. | 2.11 | 3 |
| #13758 | Disordered K/R-rich polyanion-binding modules | Basic K/R-rich, polyanion-interacting modules in nucleoproteins and related proteins: the feature highlights flexible, low-complexity/disordered basic segments adjacent to canonical DNA-binding elements (AP2/ERF and HMG-box) that contact chromatin/DNA (e.g., protamine-like regions and the N-terminal DNA-binding arm of tyrosine recombinases), and also basic N-terminal targeting peptides (signal peptide n-regions, chloroplast transit peptides) or basic loops in carbohydrate-binding modules. | 2.09 | 7 |
| #10894 | Helical hydrophobic packing signal | Generic alpha-helical hydrophobic packing signal: prefers small/hydrophobic residues within alpha-helices—spanning transmembrane helices, coiled-coils, and tandem helical repeat solenoids—rather than catalytic motifs | 2.01 | 2 |
| #10898 | Coil-to-beta edge loops | Coil-to-β-strand transition motifs (β-strand starts/edges) and the short preceding loops that frame solvent‑exposed sheet edges and pocket rims, often adjacent to substrate/ligand-binding sites across diverse proteins | 1.95 | 2 |
| #10150 | Membrane-interface helix-edge hydrophobic motifs | Short hydrophobic "helix-edge" motifs at membrane interfaces and signal-peptide entry regions: the N-terminal boundary/capping region of select transmembrane helices and adjacent cytosolic juxtamembrane amphipathic helices in multi-pass membrane proteins; occasionally short patches within a transmembrane helix (e.g., GPCR TM segments); rare analogous short hydrophobic micro-motifs also occur in soluble enzymes. | 1.94 | 2 |
| #4673 | NA-binding SLiMs and cystine-knots | Short conserved interaction motifs and broader disordered/low-complexity binding segments in nucleic-acid/chromatin-associated proteins, with peaks in disordered acidic/basic and Pro/polar-rich tails of retroelement Pol proteins, key DNA/RNA-contacting loops/β-strands (OB-folds, polymerase loops), and disordered acidic/trafficking SLiMs (e.g., AHA, NES/NLS) that mediate regulatory binding; also includes disulfide-bonded cystine-knot cores in secreted TGF-beta/BMP/GDF growth factors. | 1.93 | 3 |
| #4365 | Charged low-complexity IDRs | Charge-dense, low-complexity intrinsically disordered regions (IDRs) in eukaryotic proteins—long E/D- or K/R-rich tracts (often interspersed with Pro/Gly) that act as scaffold/interaction segments, including RS- and RGG-type regions and coiled-coil–prone linkers, common in chromatin, RNA/RNP, p97/Cdc48 adaptors, and cytoskeletal/ciliary regulators | 1.92 | 3 |
| #10290 | Disordered phospho-rich regulatory tails | Eukaryotic intrinsically disordered, low‑complexity regulatory linkers and tails — often enriched in polar and acidic residues, with embedded phospho‑Ser/Thr/Tyr clusters — in scaffold/adaptor and signaling proteins; well‑folded domains such as PTB are de‑emphasized. | 1.87 | 7 |
| #7170 | Beta-sandwich domain junctions | Terminal/edge beta-strands and adjacent linker segments of modular beta-sandwich domains, including fibronectin type III, cohesin, and Calx-beta/Ig-like modules, repeated across multi-domain bacterial cellulosomal scaffoldins, secreted modular glycosidases, and animal adhesion receptors. | 1.85 | 3 |
| #12901 | DG-centered short structured motif | A motif-based feature firing on a short structured element within globular domains, characterized by a recurring small-residue–D-G core (e.g., "xDGxxx") embedded in a β-strand/α-helix and often followed downstream by a W–H–T-like motif; prominent in carbohydrate-binding F5/8 type C domains of glycosidases but also seen in other ordered domains (DOC, BTB/POZ-containing proteins, centrosomal proteins). | 1.83 | 2 |
| #2588 | Exposed beta-loop interfaces | A cross-kingdom feature marking solvent-exposed beta-strand/loop segments within repeated, beta-rich binding/scaffold domains and processivity modules, recurring in nucleic-acid–binding S1/OB-fold proteins (rRNA biogenesis), DNA-binding plant factors (AP2/ERF, B3), replication clamps, extracellular adhesion modules (Ig/FN3/Laminin-G) and Ly6/uPAR three-finger proteins, as well as cupredoxin/cupin enzyme scaffolds; also present in nuclear transcriptional regulators. | 1.82 | 2 |
| #9873 | Pro-aromatic C-terminal FAD motif | A conserved C-terminal active-site/cofactor-binding motif in FAD-dependent pyridine nucleotide–disulfide oxidoreductases, captured as a compact Pro/aromatic-rich window centered on the FAD-contacting aromatic residue (Tyr/Phe) and adjacent substrate-binding lysine, with extensions into nearby helices and turns. | 1.81 | 2 |
| #6038 | β-sheet globular domain cores | Compact β‑sheet–dominated globular domain cores, most often from secretory‑pathway/extracellular proteins and frequently disulfide‑stabilized (cystatin, cystine‑knot and MRH lectin folds, RNase A–like), with spillover to intracellular β‑sheet–rich globular domains (e.g., NTF2‑like, WD40/β‑propeller, OB-/ribosomal cores) in cytosolic/mitochondrial proteins. | 1.80 | 2 |
| #6335 | Secondary-structure caps and TM boundaries | Secondary-structure transition/capping residues: helix/strand caps and loop/turn anchors, including residues flanking the starts/ends of transmembrane helices (generic hinge/edge positions rather than a specific sequence motif). | 1.80 | 2 |
| #10537 | WD40 blade boundary loops | Short coil/turn linkers that define WD40 β-propeller blade boundaries—i.e., the strand-connecting loops within a WD repeat and the inter-blade (repeat-to-repeat) junctions—especially at the C-terminal end of repeats but also at repeat starts; consistently marking the flexible boundary elements of WD40 domains (often involving small/polar residues such as glycine). Additionally, in WD-associated partner proteins that lack WD repeats, the feature can mark short coil/turn segments at propeller-binding interfaces. | 1.79 | 2 |
| #6490 | Folded globular domain detector | General detector of folded, globular domains (alpha/beta cores) that sharply distinguishes ordered functional domains from low-complexity/disordered termini and linkers, independent of specific function. | 1.76 | 2 |
| #15327 | Ectodomain surface loop patches | Surface-exposed loop/edge segments in extracytoplasmic proteins (secreted, lumenal/periplasmic) — most prominently within modular β‑rich recognition domains (Ig‑like, fibronectin type‑III, complement/macroglobulin), but also in exposed loops of diverse secreted enzymes and secretion/accessory lipoproteins; these patches often include N‑linked glycosylation sequons in eukaryotes | 1.76 | 2 |
| #13186 | CapZ alpha C-terminal motif | Conserved C-terminal alpha-helix/turn segment of the F-actin-capping protein alpha subunit, marking a structured region near residues ~215–255 that contains a characteristic SENY(Q/A)TM-SDTTFKAL-like motif. | 1.73 | 4 |
| #437 | Short linear helical motifs | Short alpha-helical linear motifs (∼8-20 aa) used for assembly/targeting and interface formation—often amphipathic, mixed charged/hydrophobic (E/D/K/R with hydrophobics) and frequently proline-flanked in flexible regions, but also including short hydrophobic helices such as brief transmembrane segments. | 1.72 | 2 |
| #15352 | Acidic IDR linear motifs | Short low-complexity intrinsically disordered linear motifs used for protein–protein interactions—including acidic D/E-rich tracts as well as other short linear motifs (e.g., PIP boxes, viral late-domain motifs, and the conserved acidic C-terminal tail of ribosomal P-stalk proteins, …GFGLFD-like) | 1.69 | 2 |
| #1591 | Phospho-rich coiled-coil scaffolds | Heptad-repeat coiled-coil oligomerization helices embedded within low-complexity, phosphorylation-prone intracellular (non-membrane) regions, often flanked by Ser/Thr/Glu/Asp/Gln/Asn-rich IDRs, used as scaffolds in nuclear/chromatin, spindle/centrosome, viral and RNA-processing assemblies. | 1.68 | 2 |
| #9537 | N-terminal low-hydrophobic presequence detector | N-terminal low-hydrophobic presequence detector: preferentially marks low-complexity, small/polar-rich leaders (e.g., chloroplast transit peptides and some mitochondrial transit peptides), but generally not classical hydrophobic signal peptides or strongly Arg-rich amphipathic leaders | 1.65 | 2 |
| #6961 | PDZ-flanking low-complexity IDRs | Intrinsically disordered linker and tail regions that flank or connect folded domains—most prominently between and after PDZ domains—in large metazoan scaffold/adaptor proteins (synaptic/junctional PDZ scaffolds, MAGUK-related adaptors, exocytosis regulators); the feature largely avoids the structured PDZ cores and highlights compositionally biased regulatory segments rich in polar, basic/acidic, and proline residues. | 1.65 | 2 |
| #1997 | Broad activation across nuclear coactivators | Broad regional activation across large eukaryotic nuclear proteins—commonly Mediator, SAGA, chromatin-remodeling, RNA Pol I, and other transcription/coactivator complex subunits—covering long continuous stretches that include both low-complexity acidic/Q/S/T/P-rich segments and adjacent structured (helix/strand) elements | 1.64 | 2 |
| #3653 | C-terminal repeat-boundary signature | Repeat-associated segments at repeat-unit boundaries in modular surface and secreted proteins, typically highlighting the C-terminal portions of repeat modules and adjacent linkers. In bacterial surface proteins the feature emphasizes terminal helix/strand regions of repeat domains (e.g., helical immunoglobulin-binding B/D/E domains and beta-strand caps of IgG-binding repeats); the same signature also appears in cytosolic repeat assemblies (e.g., RbcS-like repeats in carboxysome proteins) and in mucin-like Ser/Thr/Pro-rich tandem repeats of secreted adhesins. | 1.62 | 6 |
| #3342 | Unknown generic feature | Unknown generic feature | 1.79 | 2 |
Canonical-only features — 251 total (prevalence ≥ 2)
| Feature | Label | Description | Activation | Prevalence |
|---|---|---|---|---|
| #6670 | β-strand edge residue detector | Residue-level recognition of β-strand microenvironments—typically hydrophobic (Val/Ile/Leu) or small residues embedded in or immediately adjacent to β-strands, often at strand termini/edge strands—spanning many β-rich domains across diverse proteins rather than a function-specific signal. | 14.21 | 3 |
| #9666 | Alpha-helix caps and anchors | Residue-level recognition of alpha-helix termini/interfacial anchor residues: polar/aromatic sites (e.g., Asn/His/Tyr/Gly and nearby Y/F/W often following Lys/Arg) that cap or stabilize helices, including membrane–interface aromatic anchors and soluble helix caps, occurring broadly across proteins | 10.72 | 2 |
| #15815 | Basic N-terminal targeting patch | Short, basic/polar N‑terminal leader/transit segment immediately after the initiator methionine (positions ~3–6) — i.e., the positively charged/polar start of signal peptides and organelle transit peptides, and more generally low‑acidic, disordered N‑terminal patches (including occasional cysteine‑rich starts) used for targeting, export, rRNA binding, or membrane topology cues | 9.70 | 4 |
| #13282 | Conserved beta-strand core hydrophobics | Conserved hydrophobic positions within beta-strands (beta-sheet core/packing residues) in well-folded domains across diverse proteins | 9.37 | 2 |
| #10704 | N-terminal PEST-like IDRs | Low-complexity, often acidic/proline-rich intrinsically disordered N-terminal tails of intracellular eukaryotic proteins, frequently overlapping annotated PEST-like segments. | 8.87 | 10 |
| #7996 | Generic extreme N-terminus detector | Generic extreme N-terminus detector: marks residues around position 5 (occasionally 6–7) at the very start of the polypeptide, most often within the N-region of signal/transit peptides of secreted or organelle-targeted precursors, but also in generally disordered N-termini of cytosolic/nuclear proteins; not motif- or residue-specific | 8.62 | 2 |
| #4944 | PotA C-terminal YKG motif | A feature that fires predominantly within the C-terminal region of bacterial PotA-family ABC transporter ATP-binding proteins (spermidine/putrescine import), downstream of the nucleotide-binding domain. It also fires sparsely in a few other small bacterial proteins. | 8.59 | 3 |
| #7156 | Beta-strand connecting loops | Beta-strand–connecting loop/turn residues (beta-hairpin/strand–strand linkers) in beta-sheet–rich domains, often enriched for Gly and acidic residues; a generic structural motif rather than a function-specific signal | 8.00 | 2 |
| #8542 | Beta-strand conservation peaks | Conserved positions in beta-strands of structured domains (including OB/S1/KOW small beta-barrels, beta-sandwiches, and mixed alpha/beta folds), irrespective of function, detected as single-residue peaks across diverse translation-, transcription-, nucleic acid–binding, membrane adaptor, chaperonin, and ATPase proteins | 7.30 | 4 |
| #14618 | Beta-strand loop capping positions | Structural motif: β-strand-associated positions near strand→coil junctions (β-strand C-terminal "capping" positions and the first loop residues), typically at sheet edges and hairpin turns, often followed by glycine-rich loop starts; a general feature recurring across β-rich folds (β-barrels, β-propellers, β-sandwiches, and α/β enzyme cores) | 7.20 | 2 |
| #10977 | Early N-terminal disordered patch | Short N-terminal patch at the extreme N-terminus (around residues 5–10), capturing the first distinctive segment of preproteins/precursors and N-terminal disordered tails; residue composition is heterogeneous but frequently polar (Ser/Thr) and/or charged. | 6.83 | 3 |
| #5941 | Polar/acidic junction motifs | Short polar/acidic secondary-structure junctions—β-turns, β‑strand edges and α‑helix N‑caps—enriched in G/S/T/D/E (often with P), typically forming compact loop/turn segments that frame active/ligand‑binding surfaces or mark the starts of domains across diverse folds. | 6.80 | 3 |
| #6441 | Beta-strand core recognizer | Residue-level recognition of well-ordered beta-strand positions that form the cores of beta-sheet–rich folds across diverse proteins (Ig-like beta-sandwiches, beta-propellers, arrestin-like, F-box associated/FBA, ML/MD-2, and beta-strand enzymes); in secreted/Ig-like contexts the peaks often lie next to disulfide-bonded cysteines and frequently coincide with or flank N-glycosylation sequons. | 6.74 | 2 |
| #3904 | Beta-strand residue signal | Beta-strand secondary-structure signal: the feature marks residues within ordered β-sheets (DSSP class E), often in extracellular/luminal domains but broadly across taxa and functions; enrichment for polar/charged side chains within strands rather than specific catalytic motifs | 6.71 | 2 |
| #1635 | Glycine-rich flexible active-site loops | Glycine-centered flexibility sites: glycine (and occasionally other small residues) in loops/turns and at helix/strand junctions that form or flank catalytic and ligand/metal/nucleotide-binding motifs (glycine-rich loops such as P-loop/DFG, glycine clusters in cofactor-binding loops, and glycine positions adjacent to conserved catalytic residues) across diverse proteins. | 6.66 | 2 |
| #9192 | Single residue secondary structure detector | Generic detector of isolated residues embedded in canonical secondary-structure elements (alpha-helices and beta-strands), including transmembrane and signal-peptide helices and coiled-coils; reflects local backbone order/packing rather than any specific functional motif. | 6.57 | 3 |
| #1911 | Solvent-exposed beta-sheet edges | Solvent‑exposed residues in well‑ordered β‑strands and their adjoining turns, especially edge/terminal strands of β‑sheets in β‑sandwich or α/β enzyme cores (often in extracellular/periplasmic/ER‑lumenal domains and near β‑sheet→TM junctions), enriched for acidic/hydroxyl residues with frequent Gly/Pro and occasional aromatics. | 6.47 | 2 |
| #8021 | Extracellular exposed-loop activations | Extracytoplasmic/secreted proteins and extracellular or luminal domains (secretory pathway, periplasm, outer membrane, virion surface), with enrichment for carbohydrate-active enzymes and Ca2+-dependent recognition modules; the feature highlights solvent-exposed loop/turn residues across these regions rather than a specific catalytic motif. A minority of cytosolic glycosidases with similar loop features can also activate. | 6.22 | 2 |
| #1298 | Conserved catalytic beta-strand patches | Short local sequence patches within folded catalytic/structured regions of large enzymes, frequently landing on β-strand elements and conserved motifs inside catalytic domains; the feature recurs across diverse enzyme families with no strong preference for a single residue identity (peaks fall on a mix of hydrophobic, polar, and charged residues in conserved structural contexts). | 6.02 | 3 |
| #2586 | General beta-strand residue detector | General beta-strand recognition: the feature marks individual residues embedded in well-ordered beta-sheets within structured, beta-rich domains across diverse folds (e.g., alpha-crystallin/sHSP ACD, sliding clamp, glycoside hydrolases, CS/p23, FERM, Rieske). It prefers strand positions with typical side chains (hydrophobic/aromatic and Ser/Thr, with some Lys/Arg), often adjacent to functional sites but not the catalytic/liganding residues themselves, and avoids disordered regions and helices. | 5.71 | 4 |
| #5087 | Metal-assisted anionic catalysis | Anionic group-transfer/hydrolysis microenvironments: residues in active-site loops/β–α junctions that either coordinate divalent metals for phosphate chemistry (Mg2+/Mn2+/Zn2+) or stabilize negatively charged oxyanion/sulfane intermediates; most common in phosphodiester hydrolysis and phosphoryl/nucleotidyl transfer, but also present in persulfide sulfur transfer (rhodanese fold) | 5.65 | 2 |
| #15584 | Surface loops near active sites | Non-catalytic surface loops within mature trypsin-like serine protease (Peptidase S1) domains—especially the loop immediately C-terminal to the catalytic Asp of the charge-relay system—with broader cross-activation on analogous surface loop/β elements flanking active or binding pockets in other enzyme cores. | 5.59 | 2 |
| #6781 | M13 domain conserved non-catalytic peaks | M13 family metalloendopeptidase feature; activates at a small set of conserved, non-catalytic positions within the lumenal/extracellular Peptidase M13 domain across neprilysin, endothelin-converting enzyme, PHEX, and Kell-type proteins from diverse metazoa. | 5.36 | 3 |
| #8063 | Disordered N-terminal activation window | Short N-terminal disordered segment found in a variety of proteins, including plant chloroplast transit peptides, bacterial prokaryotic ubiquitin-like protein Pup N-termini, anti-sigma factor N-tails, and some small RiPP precursor leaders. The activating region is Ser/Thr/Ala/Pro/Gly/Gln-rich, low in acidic residues, and typically lies in a disordered N-terminal stretch; the feature does not generally mark all transit peptides or all RiPP leaders. | 5.34 | 7 |
| #379 | Beta-strand residue detector | Residue-level detector for well-ordered beta-strand positions (DSSP class E) in beta-sheets, independent of specific function; commonly realized in beta-sandwich and beta-propeller contexts (e.g., Rieske ferredoxin domains, PH domains, alpha-crystallin/sHSPs, glycoside hydrolases) and other alpha-beta proteins. | 5.31 | 2 |
| #10942 | Protein kinase catalytic motifs | Conserved catalytic motifs of the protein kinase core domain across Ser/Thr, Tyr, dual‑specificity, and kinase‑like (including pseudokinase and BY‑kinase) families | 5.30 | 2 |
| #7405 | Generic hydrophobic packing signal | A generic structural signal for nonpolar/compact residues at packed positions—primarily aliphatic (Leu/Ile/Val/Ala) and some aromatics (Phe/Tyr) in well-ordered α-helices and β-cores—with occasional glycine at tight turns or helix caps; function-agnostic and pervasive across folds. | 5.08 | 2 |
| #16214 | Nontransmembrane beta strands/turns | A general secondary-structure signal for short beta-strands and their flanking turns/coil in non-transmembrane regions (extracellular or cytosolic), with a preference for small/flexible or polar/charged residues; strongly depleted in alpha-helices, especially membrane-spanning helices. | 5.04 | 2 |
| #14744 | Extracellular beta-strand hydrophobics | Hydrophobic residues positioned within well‑ordered β‑strands of β‑sheet architectures (e.g., β‑propellers, β‑sandwich/Ig‑like, and other β‑rich domains), especially in extracellular/periplasmic or lumenal portions of secreted and membrane‑associated proteins; the signal avoids helices, transmembrane segments, glycosylation/metal‑binding sites, and most disulfide‑anchored positions. | 5.04 | 2 |
| #5506 | Cysteine-rich extracellular repeat modules | Cysteine-rich, disulfide‑stabilized extracellular repeat modules—principally EGF‑like repeats and closely related small domains (CCP/Sushi, WAP/TIL, disintegrin, LDLR class‑A, and von Willebrand–type modules)—found across secreted proteins, extracellular matrix components, proteases/inhibitors, and receptor ectodomains. | 5.00 | 2 |
Shared features by |Δ| activation — 704 total (prevalence ≥ 2)
| Feature | Label | Description | Δ activation | Iso act. | Canon act. | Iso prev. | Canon prev. |
|---|---|---|---|---|---|---|---|
| #512 | Short functional domain segments | Short sequence segments distributed across diverse proteins, with peaks that often fall within structured catalytic, transporter, or fold-defining domains as well as occasional acidic/charged stretches and flexible linkers. | +4.35 | 6.17 | 1.82 | 23 | 5 |
| #1068 | Ala-centered OB-fold β-strand | Conserved small/aliphatic (Ala/Val/Ile)–containing β-strand positions within S1-like/OB-fold RNA-binding domains, where alanine flanked by aliphatic residues sits in tight-packing β-elements that form RNA-contacting surfaces; activation also appears on cytoplasmic regions of some bacterial membrane-associated proteins. | -3.58 | 7.98 | 11.55 | 5 | 9 |
| #11820 | Generic beta-strand core signal | Residues located in well-ordered beta-strands, typically the central/core positions of beta-sheets across diverse folds; prominently seen in small beta-barrel RNA-binding domains (OB/KOW/S1) and in beta-strands of Rossmann-like oxidoreductases, as well as GroES/Hsp10 chaperonins—reflecting a generic beta-strand structural signal rather than a sequence- or function-specific site. | -3.44 | 6.01 | 9.45 | 2 | 5 |
| #5780 | Conserved N-terminal beta-strand cores | Conserved β‑strand residues within β‑sheet cores—especially in OB/Cold‑shock–like β‑barrels and analogous β‑rich domains—most often at the first β‑strands near domain N‑termini; captures the structured strand face used widely in nucleic‑acid–binding modules and in diverse β‑sheet proteins | -3.34 | 7.96 | 11.30 | 7 | 12 |
| #2600 | Eukaryotic regulatory helical modules | Ordered structured modules in eukaryotic regulatory and signaling proteins, with strongest activation on small helix-containing folds — including SANT/Myb-like domains in chromatin remodelers and EF-hand calcium-binding domains — and on associated helical/scaffolding elements within these proteins. The feature emphasizes residues within ordered modules rather than catalytic ATPase or DNA-contact sites, and appears across diverse eukaryotes. | -2.77 | 5.22 | 7.99 | 19 | 28 |
| #4567 | Acidic glycine beta-strand edge motif | A residue-level detector for small/hydrophobic residues embedded in short beta-strand motifs of the form V-[AT]-V-G/S that are flanked by acidic and glycine residues, typical of compact beta-sandwich/beta-barrel folds. The feature captures edge beta-strand elements of small beta-rich domains rather than a single biochemical function. | -2.63 | 7.99 | 10.62 | 2 | 5 |
| #12786 | Aromatic helical binding face | An aromatic-rich alpha-helical recognition segment common to small helix-rich domains—specifically the core helices of DnaJ/Hsp40 J domains and the DNA-recognition helices of HMG-box and AP2/ERF domains—marking the hydrophobic/aromatic face (Tyr/Trp/Phe ± His) used for partner binding, rather than family-specific motifs (e.g., not the J-domain HPD itself). | -2.61 | 4.29 | 6.89 | 3 | 5 |
| #2454 | Aromatic hydrophobic hotspot motifs | A short, local hydrophobic–aromatic micro‑motif (often including Tyr together with a charged residue) that marks interaction/catalytic hotspots across many folds—commonly the signature segments of active sites or nucleic‑acid/protein interfaces (e.g., kinase DFG/HRD/APE segments, RRM RNP1/2 aromatic motifs, SDR YxxxK)—rather than a family‑specific pattern | -2.55 | 4.89 | 7.43 | 2 | 4 |
| #1135 | Homopolymeric low-complexity tracts | Compositionally biased, low-complexity sequence segments characterized by homopolymeric residue runs; the feature detects local stretches of repeated single amino acids regardless of chemistry (polar, basic, or hydrophobic), most often within disordered or unstructured contexts but also occasionally within folded domains where such runs occur. | -2.54 | 2.42 | 4.96 | 9 | 18 |
| #6384 | Widespread activation in folded proteins | Proteins showing broad, diffuse activation across large fractions of their sequence, with peaks frequently occurring in folded regions of single-domain or multi-domain enzymes and structural proteins. | +2.46 | 5.03 | 2.57 | 14 | 23 |
| #6360 | N-terminal low-complexity IDRs | Eukaryotic intrinsically disordered, low‑complexity N‑terminal tails enriched in Lys/Arg/Ser/Asp/Glu that precede structured domains of RNA biogenesis/translation and chromatin‑remodeling factors (nucleolar/ssu‑processome, spliceosome/DEAD‑box helicases, H2A.Z chaperones); these flexible segments are typical interaction/assembly regions. | -2.45 | 2.32 | 4.77 | 5 | 23 |
| #4977 | Strand-edge turn motifs | Short coil-to-β-strand transition motifs—beta-turns/strand-edge loops (often the N-termini of β-strands)—recurring across diverse folds and functions. | -2.33 | 6.11 | 8.45 | 10 | 19 |
| #7726 | Amphipathic helical interaction motifs | Amphipathic alpha‑helical interaction motifs in eukaryotic regulatory proteins—most prominently the helices of the helix‑loop‑helix/basic HLH region, but also short helical binding elements such as PAS-domain helices, BH3 motifs, IQ/calmodulin‑binding helices, and helices within helical catalytic folds (e.g., Sec7)—enriched in basic and hydrophobic residues and used for DNA binding, dimerization, and protein–protein recognition. | -2.20 | 4.20 | 6.40 | 9 | 16 |
| #4900 | Arginine-rich segments and residues | Arginine residue identity/basic-tract feature: primarily marks arginine (R) residues—and dense arginine-rich, low-complexity segments—independent of protein family, fold, or function; shows a strong bias for R over other residues (occasional weak lysine signal), appearing both in disordered/repetitive regions and on individual R within structured domains or near transmembrane boundaries. | -2.15 | 4.45 | 6.60 | 5 | 7 |
| #13334 | N-terminal IDR boundary | Recognition of short, low-complexity, intrinsically disordered N-terminal tails immediately upstream of the first folded domain; these segments are often Ser/Thr/Pro/Gly/Ala-biased and sometimes extend a few residues into the initial helix or the very beginning of the first domain (domain boundary/disorder-to-order transition), especially in transcription regulators and co-regulators. | -2.14 | 2.85 | 5.00 | 6 | 25 |
| #6579 | Tyrosine detector with zinc-finger bias | Sequence-level detector for tyrosine residue identity (aromatic Tyr recognition), largely context-independent but with frequent emphasis on tyrosines positioned near Cys/His in C2H2-type zinc-finger motifs; weak cross-reactivity to other aromatics (Phe/Trp) and occasional neighboring acidic/amide residues. | -2.11 | 4.92 | 7.03 | 2 | 4 |
| #11990 | Paired hydrophobic anchor motif | A broad, structural micro‑motif: a pair of hydrophobic “anchor” residues that sit on short β‑strand or adjacent core positions stabilizing compact helix‑based modules (winged‑helix/HTH, histone‑fold handshakes, EF‑hand pairs, HhH/base‑flipping elements). These anchors lie next to, but not at, the primary functional sites (DNA‑ or Ca2+‑binding loops), and recur early in proteins across diverse taxa. | -2.08 | 2.72 | 4.80 | 2 | 2 |
| #10307 | Amphipathic interaction helices and caps | Short amphipathic α-helices and their capping/turn residues within compact interaction modules (protein–protein or nucleic‑acid binding), e.g., recognition helices of C2H2 zinc fingers, La/SSB HTH helices, BH3 helices, TIR/RA helical elements, and kinase helical caps; not generic helical content or long IDRs. | -1.96 | 4.71 | 6.67 | 18 | 21 |
| #4506 | Structured domain functional hotspots | Residue-level marker of positions within ordered structural domains that often coincide with functionally constrained or ligand/cofactor-interacting sites (notably heme‑ligating histidines), with no strict secondary-structure preference and generally reduced activation in extended intrinsically disordered regions. | -1.96 | 3.11 | 5.07 | 2 | 4 |
| #13122 | Charged polar low-hydrophobicity segments | Compositionally biased, low-hydrophobicity segments enriched in charged and small polar residues (E, D, K, R, S, T; often with P/G), most often found as intrinsically disordered N‑terminal stretches or propeptides but also occurring as surface-exposed polar helices/β-strands within folded domains | -1.89 | 1.79 | 3.68 | 2 | 19 |
| #10761 | SH3 tryptophan ligand groove | Conserved tryptophan‑centered aromatic signature that marks the ligand‑binding surface of SH3 and related Trp‑rich peptide/epigenetic reader modules (e.g., WW/PWWP), typically seen as GW/DW/WW or PWWP motifs within β‑strands forming the core binding groove; prevalent across eukaryotic scaffold/adaptor proteins and chromatin readers. | -1.88 | 6.50 | 8.38 | 53 | 72 |
| #8293 | Polar low-complexity IDRs | Intrinsically disordered, low-complexity segments enriched in small/polar residues (notably S, P, G, Q), corresponding to activation/modulatory regions and flexible linkers; common in transcription factors but present across many protein classes | -1.85 | 2.46 | 4.30 | 5 | 7 |
| #1501 | Acidic phosphosite-rich IDRs | Low-complexity intrinsically disordered segments enriched in polar residues with embedded acidic/phospho-regulatory clusters, serving as interaction modules across diverse proteins (nuclear/chromatin/RNA-associated and also secreted) | -1.75 | 4.81 | 6.55 | 10 | 21 |
| #8048 | Extracellular beta-strand glycan-binding | Extracellular/lumenal ectodomain signal focusing on beta‑strand–rich modules and carbohydrate-recognition contexts in secretory proteins—highlighting residues (often Asn, Gly, acidic/polar) within lectin/β‑propeller/FN3/cadherin-type domains, frequently at or adjacent to Ca(2+)/sugar-binding sites, rather than catalytic cores or transmembranes | -1.73 | 4.45 | 6.18 | 2 | 3 |
| #11491 | Loop-to-beta transition signal | A generic structural signal for beta-strand entry/edge sites: the feature activates on residues at coil-to-beta transitions and early positions within beta-strands in well-folded domains, often with acidic and/or glycine residues immediately preceding the strand and aromatic/hydrophobic residues (e.g., Phe/Tyr) at the strand start. It is not tied to specific active sites but recurs across many enzyme and non-enzyme folds. | -1.73 | 6.50 | 8.23 | 2 | 5 |
| #5022 | Disordered zinc-finger linkers | Short, intrinsically disordered linker segments that flank or connect zinc‑binding domains—especially tandem zinc‑finger repeats (C2H2, RanBP2/CCCH, matrin‑type)—with selectivity for the immediate N‑terminal side or inter‑finger spacers, rather than the folded zinc‑binding cores or other structured domains; most common in eukaryotic chromatin/nuclear regulatory proteins. | -1.72 | 3.13 | 4.85 | 6 | 16 |
| #4863 | Diffuse nonspecific background activation | An almost-null, non-specific background feature that weakly reflects generic protein context rather than any particular function or motif; when present, it shows slight, diffuse activation across ordered secondary structures in diverse proteins but is not enriched at catalytic, metal-binding, or ligand-interaction sites. | -1.71 | 8.91 | 10.62 | 29 | 40 |
| #1803 | Unknown generic feature | Unknown generic feature | -4.33 | 12.97 | 17.30 | 141 | 182 |
| #14534 | Unknown generic feature | Unknown generic feature | -4.16 | 10.44 | 14.60 | 141 | 181 |
| #14895 | Unknown generic feature | Unknown generic feature | -2.33 | 23.03 | 25.36 | 143 | 185 |
Part 2 · Differential coordinates
Features firing on just the canonical-lost (truncated) residues.
Unique-region features — 378 distinct (prevalence ≥ 2)
| Feature | Label | Description | Activation | Prevalence |
|---|---|---|---|---|
| #2114 | Chromodomain and integrase-proximal activation | A feature that activates broadly across chromodomain-containing proteins and other chromatin-associated factors, as well as on retrotransposon Gag-Pol polyproteins (in integrase-proximal regions). Activation is distributed widely along the chain with peaks tending to fall within or adjacent to chromodomain folds rather than on catalytic enzyme cores. | 13.35 | 41 |
| #1068 | Ala-centered OB-fold β-strand | Conserved small/aliphatic (Ala/Val/Ile)–containing β-strand positions within S1-like/OB-fold RNA-binding domains, where alanine flanked by aliphatic residues sits in tight-packing β-elements that form RNA-contacting surfaces; activation also appears on cytoplasmic regions of some bacterial membrane-associated proteins. | 11.55 | 4 |
| #5780 | Conserved N-terminal beta-strand cores | Conserved β‑strand residues within β‑sheet cores—especially in OB/Cold‑shock–like β‑barrels and analogous β‑rich domains—most often at the first β‑strands near domain N‑termini; captures the structured strand face used widely in nucleic‑acid–binding modules and in diverse β‑sheet proteins | 11.30 | 5 |
| #4567 | Acidic glycine beta-strand edge motif | A residue-level detector for small/hydrophobic residues embedded in short beta-strand motifs of the form V-[AT]-V-G/S that are flanked by acidic and glycine residues, typical of compact beta-sandwich/beta-barrel folds. The feature captures edge beta-strand elements of small beta-rich domains rather than a single biochemical function. | 10.62 | 3 |
| #4863 | Diffuse nonspecific background activation | An almost-null, non-specific background feature that weakly reflects generic protein context rather than any particular function or motif; when present, it shows slight, diffuse activation across ordered secondary structures in diverse proteins but is not enriched at catalytic, metal-binding, or ligand-interaction sites. | 10.62 | 9 |
| #13702 | Regulatory low-complexity IDRs | Eukaryotic intrinsically disordered, low‑complexity regulatory segments (terminal tails and inter‑domain linkers) enriched in Pro/Ser/Thr/acidic composition and short linear motifs for modular protein–protein interactions—frequently including WW‑domain–binding PY/PPxY segments and PTM hotspots—with reduced but not absent activation within folded recognition/catalytic domains (WW, PTB/PID, chromo/chromoshadow, SET). | 10.05 | 35 |
| #15815 | Basic N-terminal targeting patch | Short, basic/polar N‑terminal leader/transit segment immediately after the initiator methionine (positions ~3–6) — i.e., the positively charged/polar start of signal peptides and organelle transit peptides, and more generally low‑acidic, disordered N‑terminal patches (including occasional cysteine‑rich starts) used for targeting, export, rRNA binding, or membrane topology cues | 9.70 | 4 |
| #11820 | Generic beta-strand core signal | Residues located in well-ordered beta-strands, typically the central/core positions of beta-sheets across diverse folds; prominently seen in small beta-barrel RNA-binding domains (OB/KOW/S1) and in beta-strands of Rossmann-like oxidoreductases, as well as GroES/Hsp10 chaperonins—reflecting a generic beta-strand structural signal rather than a sequence- or function-specific site. | 9.45 | 3 |
| #8139 | Structured cores of genome regulators | Core residues of folded domains in eukaryotic genome-function proteins (chromatin-, DNA-, and RNA-associated factors, including transcriptional regulators, RNA-binding proteins, replication and repair factors), emphasizing beta-strand/loop elements of nucleic-acid–associated folds (e.g., B3 DNA-binding and RRM RNA-binding) and the PB1 oligomerization module; the feature marks conserved, well-structured domain cores and avoids long low-complexity/disordered regions. | 9.06 | 9 |
| #15240 | Beta-strand recognition modules | Structured β‑strand/turn binding interfaces of compact recognition modules in eukaryotic regulators—predominantly β‑rich folds that present aromatic/charged residues (and often Zn‑coordinating Cys/His) to recognize short peptides/modified histone tails, DNA, or lipids. The feature highlights reader/adaptor domains (Royal‑family histone readers: chromodomain/MBT/Tudor/PWWP; BAH; PHD/CW), peptide‑binding scaffolds (PDZ, CAP‑Gly), lipid/calcium binders (C2), and β‑rich DNA‑binding domains (B3, Rel homology), and also structured interaction surfaces of PB1—while remaining silent on long disordered linkers. | 8.88 | 7 |
| #10704 | N-terminal PEST-like IDRs | Low-complexity, often acidic/proline-rich intrinsically disordered N-terminal tails of intracellular eukaryotic proteins, frequently overlapping annotated PEST-like segments. | 8.87 | 10 |
| #4944 | PotA C-terminal YKG motif | A feature that fires predominantly within the C-terminal region of bacterial PotA-family ABC transporter ATP-binding proteins (spermidine/putrescine import), downstream of the nucleotide-binding domain. It also fires sparsely in a few other small bacterial proteins. | 8.59 | 2 |
| #4977 | Strand-edge turn motifs | Short coil-to-β-strand transition motifs—beta-turns/strand-edge loops (often the N-termini of β-strands)—recurring across diverse folds and functions. | 8.45 | 9 |
| #10761 | SH3 tryptophan ligand groove | Conserved tryptophan‑centered aromatic signature that marks the ligand‑binding surface of SH3 and related Trp‑rich peptide/epigenetic reader modules (e.g., WW/PWWP), typically seen as GW/DW/WW or PWWP motifs within β‑strands forming the core binding groove; prevalent across eukaryotic scaffold/adaptor proteins and chromatin readers. | 8.38 | 22 |
| #11491 | Loop-to-beta transition signal | A generic structural signal for beta-strand entry/edge sites: the feature activates on residues at coil-to-beta transitions and early positions within beta-strands in well-folded domains, often with acidic and/or glycine residues immediately preceding the strand and aromatic/hydrophobic residues (e.g., Phe/Tyr) at the strand start. It is not tied to specific active sites but recurs across many enzyme and non-enzyme folds. | 8.23 | 3 |
| #10571 | Beta-sandwich anion-binding interface | β-strand–rich binding-surface signature of β-sandwich/β-barrel folds used to engage anionic ligands, most prominently phosphoinositides at membranes and DNA/RNA in nuclei. The feature marks conserved β-strands and adjacent loops of PH/C2-like lipid-binding domains and Ig-like Rel-homology–type DNA-binding domains, with similar activation on related β-rich reader domains (e.g., PWWP/MBD) and β/loop interfaces in diverse eukaryotic trafficking, cytoskeletal, and chromatin proteins. | 8.15 | 9 |
| #13745 | Hydrophobic secondary structure signal | Generic hydrophobic secondary-structure signal: activates on hydrophobic/aromatic side chains that form stable α-helices and β-strands—covering transmembrane helices, coiled-coils, and the hydrophobic cores of globular domains—while avoiding catalytic/metal-binding residues and disordered segments | 8.06 | 10 |
| #10583 | IAPP motif and signal-peptide detector | Detector of a conserved motif within the mature region of amylin/IAPP-family secreted peptide hormones, with secondary activations in the upstream signal peptide of the same precursors. | 8.00 | 6 |
| #2600 | Eukaryotic regulatory helical modules | Ordered structured modules in eukaryotic regulatory and signaling proteins, with strongest activation on small helix-containing folds — including SANT/Myb-like domains in chromatin remodelers and EF-hand calcium-binding domains — and on associated helical/scaffolding elements within these proteins. The feature emphasizes residues within ordered modules rather than catalytic ATPase or DNA-contact sites, and appears across diverse eukaryotes. | 7.99 | 9 |
| #5895 | PDZ and chromatin readers | Eukaryote-biased feature marking scaffold/signaling PDZ-domain proteins and nuclear chromatin regulators (readers/writers/remodelers), with additional coverage of ubiquitin/SUMO-pathway E3s, polarity/Rho-pathway components, PP1 glycogen-targeting CBM21 proteins, and ATAT1 acetyltransferases; broad functional/subcellular coverage and low specificity in SwissProt. | 7.82 | 7 |
| #14795 | MutS2 Smr upstream hydrophobic motif | Recognition of a conserved hydrophobic motif in bacterial MutS2 endonucleases (Ribosome-associated protein quality control upstream factor), located immediately N-terminal to the C-terminal Smr domain. | 7.72 | 2 |
| #8668 | Aromatic residues in helices | Aromatic side chains (Tyr/Phe/Trp) embedded in ordered alpha‑helical contexts—most prominently transmembrane helices of multi‑pass membrane proteins, but also helices in DNA‑binding modules—rather than a family‑specific motif. | 7.68 | 2 |
| #3343 | Interaction-domain beta-sheet cores | β-strand cores of modular interaction/reader domains in eukaryotic scaffold and signaling proteins, especially PDZ and PH-like (PH/GRAM/FYVE) folds and chromatin-reader barrels (PWWP/PHD), with strong activation in the Smad MH2 β-sheet; the feature marks folded domain interiors/binding-sheet surfaces and avoids disordered linkers | 7.66 | 7 |
| #4066 | Generic functional region marker | Generic marker of “main functional regions” within proteins: activates inside annotated functional modules (folded catalytic/ligand‑binding domains, extracellular globular domains, mature secreted peptide segments, and disordered protein–protein interaction modules), with occasional emphasis near catalytic/ligand-binding or interaction hotspots and a tendency to spare translocation/membrane-embedding segments | 7.62 | 7 |
| #14895 | Unknown generic feature | Unknown generic feature | 21.61 | 43 |
| #9214 | Unknown generic feature | Unknown generic feature | 17.88 | 43 |
| #1803 | Unknown generic feature | Unknown generic feature | 17.30 | 43 |
| #14534 | Unknown generic feature | Unknown generic feature | 14.60 | 41 |
| #9194 | Unknown generic feature | Unknown generic feature | 14.05 | 41 |
| #9005 | Unknown generic feature | Unknown generic feature | 10.55 | 41 |
Top-K sparse-autoencoder activations on the ESM-C residual stream. Activation is a feature's peak value over its region; Prevalence is the number of residues it fires on; Δ activation is isoform − canonical. Only the top 30 features per category are saved; generic / "unknown" features are hidden until you expand a table. Read more about SAE features →
Clinical variants
Differential region — lost N-terminus (canonical-only)
| Pos (iso) | AA change | Consequence | Source | Clin. sig. | AF (gnomAD) | Impact | AlphaMissense | ESM-C ΔLLR | Link |
|---|---|---|---|---|---|---|---|---|---|
| — | K→K | intronic | gnomAD | — | 6.85e-07 | — | — | 0.00 | chr17-48076882-C-T |
| — | K→E | intronic | gnomAD | — | 6.85e-07 | damaging | likely_pathogenic (1.00) | -11.19 | chr17-48076884-T-C |
| — | L→L | intronic | gnomAD | — | 6.85e-07 | — | — | 0.00 | chr17-48076885-T-G |
| — | L→L | intronic | gnomAD | — | 3.43e-06 | — | — | 0.00 | chr17-48076888-G-A |
| — | L→R | intronic | gnomAD | — | 6.85e-07 | damaging | likely_pathogenic (0.84) | -9.56 | chr17-48076889-A-C |
| — | Y→Y | intronic | gnomAD | — | 1.71e-05 | — | — | 0.00 | chr17-48076891-G-A |
| — | E→D | intronic | gnomAD | — | 6.85e-07 | damaging | likely_pathogenic (1.00) | -7.72 | chr17-48076894-C-G |
| — | K→R | intronic | gnomAD | — | 2.05e-06 | — | likely_benign (0.09) | -5.18 | chr17-48076901-T-C |
| — | G→G | intronic | gnomAD | — | 3.42e-06 | — | — | 0.00 | chr17-48076903-G-A |
| — | G→S | intronic | gnomAD | — | 6.84e-07 | damaging | likely_pathogenic (0.96) | -7.97 | chr17-48076905-C-T |
| — | K→K | intronic | gnomAD | — | 6.84e-07 | — | — | 0.00 | chr17-48076906-C-T |
| — | V→V | intronic | gnomAD | — | 1.37e-06 | — | — | 0.00 | chr17-48076909-T-C |
| — | V→L | intronic | gnomAD | — | 6.84e-07 | damaging | likely_pathogenic (0.76) | -6.96 | chr17-48076911-C-A |
| — | V→V | intronic | gnomAD | — | 2.05e-06 | — | — | 0.00 | chr17-48076912-C-T |
| — | V→A | intronic | gnomAD | — | 1.37e-06 | damaging | likely_pathogenic (0.92) | -7.43 | chr17-48076913-A-G |
| — | R→Q | intronic | gnomAD | — | 2.05e-06 | damaging | likely_pathogenic (0.99) | -9.25 | chr17-48076916-C-T |
| — | R→R | intronic | gnomAD | — | 6.84e-07 | — | — | 0.00 | chr17-48076917-G-T |
| — | R→H | intronic | gnomAD | — | 2.74e-06 | damaging | likely_pathogenic (0.79) | -7.03 | chr17-48076919-C-T |
| — | R→C | intronic | gnomAD | — | 1.23e-05 | damaging | likely_pathogenic (0.87) | -7.12 | chr17-48076920-G-A |
| — | D→H | intronic | gnomAD | — | 1.37e-06 | damaging | likely_pathogenic (0.94) | -7.89 | chr17-48076923-C-G |
| — | D→N | intronic | gnomAD | — | 1.85e-05 | damaging | likely_pathogenic (0.56) | -4.01 | chr17-48076923-C-T |
| — | L→L | intronic | gnomAD | — | 6.84e-06 | — | — | 0.00 | chr17-48076924-G-A |
| — | K→N | intronic | gnomAD | — | 6.84e-07 | damaging | likely_pathogenic (0.99) | -9.37 | chr17-48076930-T-A |
| — | V→V | intronic | gnomAD | — | 6.84e-07 | — | — | 0.00 | chr17-48076936-C-A |
| — | V→V | intronic | gnomAD | — | 6.84e-07 | — | — | 0.00 | chr17-48076936-C-G |
| — | V→V | intronic | gnomAD | — | 2.74e-06 | — | — | 0.00 | chr17-48076939-C-T |
| — | E→E | intronic | gnomAD | — | 2.05e-06 | — | — | 0.00 | chr17-48076945-T-C |
| — | E→E | intronic | gnomAD | — | 1.37e-06 | — | — | 0.00 | chr17-48076951-T-C |
| — | E→G | intronic | gnomAD | — | 6.84e-07 | damaging | likely_pathogenic (0.77) | -9.56 | chr17-48076952-T-C |
| — | E→Q | intronic | gnomAD | — | 1.37e-06 | damaging | likely_pathogenic (0.61) | -9.81 | chr17-48076956-C-G |
| — | E→Q | intronic | gnomAD | — | 6.84e-07 | damaging | likely_pathogenic (0.62) | -9.69 | chr17-48076959-C-G |
| — | E→K | intronic | gnomAD | — | 6.84e-07 | damaging | likely_pathogenic (0.83) | -11.12 | chr17-48076962-C-T |
| — | L→L | intronic | gnomAD | — | 2.05e-06 | — | — | 0.00 | chr17-48076963-T-C |
| — | V→V | intronic | gnomAD | — | 1.37e-06 | — | — | 0.00 | chr17-48076966-C-T |
| — | V→L | intronic | gnomAD | — | 2.05e-06 | damaging | likely_benign (0.30) | -7.78 | chr17-48076968-C-A |
| — | V→M | intronic | gnomAD | — | 2.74e-06 | damaging | likely_benign (0.24) | -8.68 | chr17-48076968-C-T |
| — | V→V | intronic | gnomAD | — | 1.37e-06 | — | — | 0.00 | chr17-48076975-C-T |
| — | V→G | intronic | gnomAD | — | 6.85e-07 | damaging | likely_benign (0.14) | -8.75 | chr17-48076976-A-C |
| — | V→M | intronic | gnomAD | — | 7.53e-06 | damaging | likely_benign (0.13) | -8.37 | chr17-48076977-C-T |
| — | K→N | intronic | gnomAD | — | 6.84e-07 | damaging | likely_pathogenic (0.58) | -9.00 | chr17-48076978-T-A |
| — | — | intronic | gnomAD | — | 1.37e-06 | — | — | — | chr17-48076978-TTTC-T |
| — | K→I | intronic | gnomAD | — | 6.84e-07 | damaging | ambiguous (0.56) | -11.69 | chr17-48076979-T-A |
| — | K→R | intronic | gnomAD | — | 9.58e-06 | damaging | likely_benign (0.12) | -8.94 | chr17-48076979-T-C |
| — | K→* | intronic | gnomAD | — | 6.84e-07 | damaging | — | — | chr17-48076980-T-A |
| — | K→E | intronic | gnomAD | — | 6.84e-07 | damaging | ambiguous (0.42) | -9.69 | chr17-48076980-T-C |
| — | K→N | intronic | gnomAD | — | 6.84e-07 | damaging | likely_pathogenic (0.84) | -9.87 | chr17-48076981-C-A |
| — | K→R | intronic | gnomAD | — | 6.84e-07 | damaging | likely_benign (0.12) | -9.06 | chr17-48076982-T-C |
| — | K→N | intronic | gnomAD | — | 6.84e-07 | damaging | likely_pathogenic (0.73) | -9.94 | chr17-48076984-C-G |
| — | — | intronic | gnomAD | — | 6.84e-07 | — | — | — | chr17-48076984-CTTG-C |
| — | K→Q | intronic | gnomAD | — | 1.37e-06 | damaging | likely_benign (0.24) | -10.25 | chr17-48076986-T-G |
| — | N→K | intronic | gnomAD | — | 1.85e-05 | damaging | likely_pathogenic (0.68) | -8.87 | chr17-48076987-G-C |
| — | Q→Q | intronic | gnomAD | — | 6.84e-07 | — | — | 0.00 | chr17-48076990-T-C |
| — | K→K | intronic | gnomAD | — | 2.74e-06 | — | — | 0.00 | chr17-48076993-T-C |
| — | G→G | intronic | gnomAD | — | 6.85e-07 | — | — | 0.00 | chr17-48076999-C-A |
| — | G→V | intronic | gnomAD | — | 6.85e-07 | damaging | likely_pathogenic (0.88) | -9.25 | chr17-48077000-C-A |
| — | G→W | intronic | gnomAD | — | 6.85e-07 | damaging | likely_pathogenic (0.97) | -12.31 | chr17-48077001-C-A |
| — | M→I | intronic | gnomAD | — | 1.37e-06 | damaging | — | -10.94 | chr17-48077002-C-G |
| — | M→I | intronic | gnomAD | — | 6.85e-07 | damaging | — | -10.94 | chr17-48077002-C-T |
| — | M→T | intronic | gnomAD | — | 6.85e-07 | damaging | — | -11.50 | chr17-48077003-A-G |
| — | M→L | intronic | gnomAD | — | 6.85e-07 | damaging | — | -11.19 | chr17-48077004-T-G |
| — | N→K | intronic | ClinVar | Uncertain significance | — | damaging | likely_pathogenic (0.68) | -8.87 | ClinVar:3827926 |
| — | E→V | intronic | ClinVar | — | — | damaging | likely_pathogenic (1.00) | -11.31 | ClinVar:4423332 |
| — | L→F | intronic | COSMIC | — | — | — | ambiguous (0.35) | -6.53 | COSV56681814 |
| — | E→K | intronic | COSMIC | — | — | damaging | likely_pathogenic (1.00) | -10.19 | COSV56682856 |
| — | G→S | intronic | COSMIC | — | — | damaging | likely_pathogenic (0.96) | -7.97 | COSV99847936 |
| — | R→R | intronic | COSMIC | — | — | — | — | 0.00 | COSV56682056 |
| — | R→* | intronic | COSMIC | — | — | damaging | — | — | COSV56681749 |
| — | L→I | intronic | COSMIC | — | — | damaging | likely_benign (0.28) | -9.56 | COSV56682110 |
| — | V→V | intronic | COSMIC | — | — | — | — | 0.00 | COSV99848141 |
| — | E→Q | intronic | COSMIC | — | — | damaging | likely_pathogenic (0.95) | -11.25 | COSV99848313 |
| — | E→K | intronic | COSMIC | — | — | damaging | likely_pathogenic (0.83) | -9.56 | COSV56681509 |
| — | V→L | intronic | COSMIC | — | — | damaging | likely_benign (0.20) | -8.06 | COSV56682280 |
| — | Q→K | intronic | COSMIC | — | — | damaging | ambiguous (0.37) | -10.00 | COSV56682149 |
| — | — | intronic | COSMIC | — | — | damaging | — | — | COSV56681712 |
| — | — | intronic | COSMIC | — | — | damaging | — | — | COSV56681499 |
| — | M→V | intronic | COSMIC | — | — | damaging | — | -11.44 | COSV56682591 |
76 variants in the differential region.
Shared canonical core — sequence common to canonical and isoform; AlphaMissense applies here
| Pos (iso) | AA change | Consequence | Source | Clin. sig. | AF (gnomAD) | Impact | AlphaMissense | ESM-C ΔLLR | Link |
|---|---|---|---|---|---|---|---|---|---|
| 1 | G→G | synonymous_variant | gnomAD | — | 4.39e-05 | — | — | 0.00 | chr17-48076873-T-C |
| 2 | F→F | synonymous_variant | gnomAD | — | 6.86e-07 | — | — | 0.00 | chr17-48076870-G-A |
| 2 | F→L | missense_variant | COSMIC | — | — | damaging | likely_pathogenic (1.00) | -10.37 | COSV56682209 |
| 5 | E→K | missense_variant | gnomAD | — | 7.10e-07 | damaging | likely_pathogenic (0.84) | -9.50 | chr17-48076177-C-T |
| 5 | E→Q | missense_variant | ClinVar | — | — | damaging | likely_pathogenic (0.60) | -11.12 | ClinVar:4423331 |
| 5 | E→Q | missense_variant | COSMIC | — | — | damaging | likely_pathogenic (0.60) | -11.12 | COSV56681530 |
| 6 | D→H | missense_variant | gnomAD | — | 7.06e-07 | damaging | likely_pathogenic (0.97) | -12.12 | chr17-48076174-C-G |
| 7 | N→D | missense_variant | gnomAD | — | 1.41e-06 | damaging | likely_pathogenic (0.97) | -10.37 | chr17-48076171-T-C |
| 7 | N→S | missense_variant | COSMIC | — | — | damaging | likely_pathogenic (0.59) | -9.00 | COSV56681968 |
| 10 | E→D | missense_variant | COSMIC | — | — | damaging | likely_pathogenic (1.00) | -10.56 | COSV99848274 |
| 12 | E→D | missense_variant | COSMIC | — | — | damaging | likely_pathogenic (0.96) | -8.69 | COSV99848151 |
| 14 | N→N | synonymous_variant | gnomAD | — | 1.38e-06 | — | — | 0.00 | chr17-48076148-G-A |
| 15 | L→L | synonymous_variant | gnomAD | — | 1.38e-06 | — | — | 0.00 | chr17-48076147-G-A |
| 18 | P→P | synonymous_variant | gnomAD | — | 1.51e-05 | — | — | 0.00 | chr17-48076136-G-A |
| 18 | P→P | synonymous_variant | gnomAD | — | 6.88e-07 | — | — | 0.00 | chr17-48076136-G-C |
| 18 | P→P | synonymous_variant | COSMIC | — | — | — | — | 0.00 | COSV99847948 |
| 19 | D→N | missense_variant | COSMIC | — | — | damaging | likely_pathogenic (0.96) | -8.31 | COSV99847954 |
| 20 | L→L | synonymous_variant | gnomAD | — | 1.23e-05 | — | — | 0.00 | chr17-48076130-G-A |
| 22 | A→A | synonymous_variant | gnomAD | — | 5.48e-06 | — | — | 0.00 | chr17-48076124-A-G |
| 22 | A→A | synonymous_variant | gnomAD | — | 6.85e-07 | — | — | 0.00 | chr17-48076124-A-T |
| 22 | A→G | missense_variant | gnomAD | — | 1.37e-06 | damaging | ambiguous (0.38) | -9.94 | chr17-48076125-G-C |
| 25 | L→L | synonymous_variant | gnomAD | — | 1.37e-06 | — | — | 0.00 | chr17-48076115-C-T |
| 25 | L→L | synonymous_variant | gnomAD | — | 6.85e-07 | — | — | 0.00 | chr17-48076117-G-A |
| 25 | L→L | synonymous_variant | COSMIC | — | — | — | — | 0.00 | COSV105045650 |
| 26 | Q→* | stop_gained | COSMIC | — | — | LoF | — | — | COSV99848158 |
| 27 | S→L | missense_variant | gnomAD | — | 2.74e-06 | damaging | likely_benign (0.27) | -9.37 | chr17-48076110-G-A |
| 27 | S→L | missense_variant | ClinVar | Uncertain significance | — | damaging | likely_benign (0.27) | -9.37 | ClinVar:4648788 |
| 28 | Q→R | missense_variant | gnomAD | — | 6.85e-07 | — | likely_benign (0.33) | -7.47 | chr17-48076107-T-C |
| 28 | Q→* | stop_gained | COSMIC | — | — | LoF | — | — | COSV99848270 |
| 29 | K→R | missense_variant | gnomAD | — | 6.85e-07 | — | likely_benign (0.12) | -6.44 | chr17-48076104-T-C |
| 29 | K→Q | missense_variant | ClinVar | Uncertain significance | — | damaging | likely_benign (0.33) | -8.75 | ClinVar:4531663 |
| 30 | T→A | missense_variant | gnomAD | — | 6.85e-07 | — | likely_benign (0.06) | -6.06 | chr17-48076102-T-C |
| 32 | H→P | missense_variant | gnomAD | — | 4.79e-06 | — | likely_benign (0.11) | -7.21 | chr17-48076095-T-G |
| 34 | T→R | missense_variant | gnomAD | — | 2.74e-06 | damaging | likely_benign (0.23) | -7.90 | chr17-48076089-G-C |
| 35 | D→G | missense_variant | gnomAD | — | 6.84e-07 | — | likely_benign (0.10) | -6.27 | chr17-48076086-T-C |
| 35 | D→N | missense_variant | gnomAD | — | 6.84e-07 | — | likely_benign (0.10) | -6.15 | chr17-48076087-C-T |
| 35 | D→G | missense_variant | ClinVar | Uncertain significance | — | — | likely_benign (0.10) | -6.27 | ClinVar:4220041 |
| 35 | D→N | missense_variant | ClinVar | Uncertain significance | — | — | likely_benign (0.10) | -6.15 | ClinVar:4220044 |
| 36 | K→K | synonymous_variant | gnomAD | — | 4.04e-05 | — | — | 0.00 | chr17-48076082-T-C |
| 36 | K→I | missense_variant | gnomAD | — | 6.84e-07 | damaging | ambiguous (0.43) | -10.75 | chr17-48076083-T-A |
| 36 | K→I | missense_variant | ClinVar | Uncertain significance | — | damaging | ambiguous (0.43) | -10.75 | ClinVar:4220042 |
| 37 | S→S | synonymous_variant | gnomAD | — | 4.79e-06 | — | — | 0.00 | chr17-48076079-T-G |
| 38 | E→K | missense_variant | COSMIC | — | — | damaging | likely_benign (0.26) | -8.68 | COSV99848247 |
| 39 | G→R | missense_variant | gnomAD | — | 6.84e-07 | damaging | likely_pathogenic (0.58) | -8.76 | chr17-48076075-C-G |
| 39 | G→R | missense_variant | COSMIC | — | — | damaging | likely_pathogenic (0.58) | -8.76 | COSV56682528 |
| 42 | R→H | missense_variant | gnomAD | — | 7.53e-06 | damaging | likely_pathogenic (0.89) | -9.31 | chr17-48076065-C-T |
| 42 | R→C | missense_variant | gnomAD | — | 6.16e-06 | damaging | likely_pathogenic (0.96) | -9.31 | chr17-48076066-G-A |
| 42 | R→H | missense_variant | ClinVar | Uncertain significance | — | damaging | likely_pathogenic (0.89) | -9.31 | ClinVar:2290145 |
| 42 | R→H | missense_variant | COSMIC | — | — | damaging | likely_pathogenic (0.89) | -9.31 | COSV56682901 |
| 42 | R→C | missense_variant | COSMIC | — | — | damaging | likely_pathogenic (0.96) | -9.31 | COSV56682920 |
| 43 | K→R | missense_variant | gnomAD | — | 8.90e-06 | damaging | likely_benign (0.11) | -7.87 | chr17-48076062-T-C |
| 44 | A→T | missense_variant | gnomAD | — | 1.37e-06 | — | likely_benign (0.08) | -6.34 | chr17-48076060-C-T |
| 44 | A→T | missense_variant | ClinVar | Uncertain significance | — | — | likely_benign (0.08) | -6.34 | ClinVar:4220043 |
| 45 | D→G | missense_variant | gnomAD | — | 6.85e-07 | damaging | likely_benign (0.15) | -8.36 | chr17-48076056-T-C |
| 46 | S→A | missense_variant | gnomAD | — | 6.85e-07 | damaging | likely_benign (0.06) | -9.87 | chr17-48076054-A-C |
| 46 | S→T | missense_variant | gnomAD | — | 1.37e-06 | — | likely_benign (0.05) | -7.31 | chr17-48076054-A-T |
| 50 | D→V | missense_variant | gnomAD | — | 1.37e-06 | damaging | likely_benign (0.13) | -9.56 | chr17-48076041-T-A |
| 50 | D→V | missense_variant | ClinVar | Uncertain significance | — | damaging | likely_benign (0.13) | -9.56 | ClinVar:2307098 |
| 50 | D→N | missense_variant | COSMIC | — | — | damaging | likely_benign (0.11) | -7.93 | COSV56682563 |
| 51 | K→K | synonymous_variant | gnomAD | — | 1.37e-06 | — | — | 0.00 | chr17-48076037-C-T |
| 51 | K→R | missense_variant | gnomAD | — | 2.05e-06 | — | likely_benign (0.09) | -4.90 | chr17-48076038-T-C |
| 51 | K→E | missense_variant | gnomAD | — | 6.85e-07 | damaging | likely_benign (0.18) | -11.05 | chr17-48076039-T-C |
| 51 | K→K | synonymous_variant | COSMIC | — | — | — | — | 0.00 | COSV56682520 |
| 51 | K→N | missense_variant | COSMIC | — | — | damaging | likely_benign (0.29) | -9.30 | COSV99848170 |
| 52 | G→G | synonymous_variant | gnomAD | — | 2.06e-06 | — | — | 0.00 | chr17-48076034-T-C |
| 52 | — | frameshift_variant | gnomAD | — | 1.37e-06 | LoF | — | — | chr17-48076035-CCCTT-C |
| 52 | G→R | missense_variant | gnomAD | — | 2.74e-06 | damaging | ambiguous (0.54) | -7.55 | chr17-48076036-C-T |
| 53 | E→E | synonymous_variant | gnomAD | — | 6.86e-07 | — | — | 0.00 | chr17-48076031-C-T |
| 53 | E→Q | missense_variant | COSMIC | — | — | damaging | likely_benign (0.23) | -8.62 | COSV99848203 |
| 54 | E→D | missense_variant | gnomAD | — | 6.86e-07 | — | likely_benign (0.05) | -5.84 | chr17-48076028-C-G |
| 55 | S→S | synonymous_variant | gnomAD | — | 4.81e-06 | — | — | 0.00 | chr17-48076025-G-A |
| 55 | S→G | missense_variant | gnomAD | — | 6.86e-07 | — | likely_benign (0.06) | -5.94 | chr17-48076027-T-C |
| 55 | S→S | synonymous_variant | COSMIC | — | — | — | — | 0.00 | COSV99848275 |
| 57 | P→P | synonymous_variant | gnomAD | — | 2.07e-06 | — | — | 0.00 | chr17-48076019-T-C |
| 57 | P→L | missense_variant | gnomAD | — | 1.38e-06 | — | likely_benign (0.09) | -6.30 | chr17-48076020-G-A |
| 59 | K→N | missense_variant | gnomAD | — | 6.90e-07 | damaging | likely_pathogenic (0.93) | -10.37 | chr17-48076013-C-A |
| 59 | K→K | synonymous_variant | gnomAD | — | 1.38e-06 | — | — | 0.00 | chr17-48076013-C-T |
| 60 | K→N | missense_variant | gnomAD | — | 6.91e-07 | damaging | likely_pathogenic (0.91) | -10.25 | chr17-48076010-C-G |
| 60 | K→R | missense_variant | gnomAD | — | 1.38e-06 | — | likely_benign (0.09) | -5.37 | chr17-48076011-T-C |
| 60 | K→N | missense_variant | COSMIC | — | — | damaging | likely_pathogenic (0.91) | -10.25 | COSV56681336 |
| 60 | K→N | missense_variant | COSMIC | — | — | damaging | likely_pathogenic (0.91) | -10.25 | COSV99848297 |
| 61 | — | inframe_deletion | gnomAD | — | 2.77e-06 | — | — | — | chr17-48076007-TTTC-T |
| 61 | K→R | missense_variant | gnomAD | — | 6.92e-07 | — | likely_benign (0.09) | -6.75 | chr17-48076008-T-C |
| 62 | E→G | missense_variant | gnomAD | — | 6.92e-07 | damaging | likely_benign (0.33) | -8.81 | chr17-48076005-T-C |
| 63 | E→V | missense_variant | gnomAD | — | 1.39e-06 | damaging | likely_benign (0.30) | -8.24 | chr17-48076002-T-A |
| 63 | E→Q | missense_variant | gnomAD | — | 6.94e-07 | damaging | ambiguous (0.43) | -8.99 | chr17-48076003-C-G |
| 64 | S→S | synonymous_variant | gnomAD | — | 6.85e-07 | — | — | 0.00 | chr17-48075098-T-A |
| 64 | S→L | missense_variant | gnomAD | — | 6.85e-07 | — | likely_benign (0.08) | -6.09 | chr17-48075099-G-A |
| 65 | E→A | missense_variant | gnomAD | — | 6.85e-07 | damaging | ambiguous (0.50) | -9.43 | chr17-48075096-T-G |
| 67 | P→S | missense_variant | gnomAD | — | 6.84e-07 | — | ambiguous (0.53) | -5.76 | chr17-48075091-G-A |
| 67 | P→Q | missense_variant | COSMIC | — | — | damaging | likely_pathogenic (0.61) | -7.20 | COSV106082017 |
| 67 | P→S | missense_variant | COSMIC | — | — | — | ambiguous (0.53) | -5.76 | COSV56682997 |
| 68 | R→Q | missense_variant | gnomAD | — | 6.84e-07 | damaging | likely_pathogenic (0.92) | -8.69 | chr17-48075087-C-T |
| 68 | R→* | stop_gained | gnomAD | — | 2.05e-06 | LoF | — | — | chr17-48075088-G-A |
| 68 | R→* | stop_gained | COSMIC | — | — | LoF | — | — | COSV56682269 |
| 70 | F→S | missense_variant | gnomAD | — | 6.84e-07 | damaging | likely_pathogenic (1.00) | -10.44 | chr17-48075081-A-G |
| 71 | A→A | synonymous_variant | gnomAD | — | 1.37e-06 | — | — | 0.00 | chr17-48075077-A-C |
| 71 | A→A | synonymous_variant | gnomAD | — | 6.84e-07 | — | — | 0.00 | chr17-48075077-A-G |
| 72 | R→Q | missense_variant | gnomAD | — | 6.84e-07 | damaging | likely_pathogenic (0.99) | -8.12 | chr17-48075075-C-T |
| 72 | R→* | stop_gained | gnomAD | — | 6.84e-07 | LoF | — | — | chr17-48075076-G-A |
| 72 | R→R | synonymous_variant | COSMIC | — | — | — | — | 0.00 | COSV107309107 |
| 72 | R→L | missense_variant | COSMIC | — | — | damaging | likely_pathogenic (1.00) | -10.37 | COSV99847942 |
| 72 | R→Q | missense_variant | COSMIC | — | — | damaging | likely_pathogenic (0.99) | -8.12 | COSV56683114 |
| 72 | R→* | stop_gained | COSMIC | — | — | LoF | — | — | COSV104387334 |
| 73 | G→D | missense_variant | COSMIC | — | — | damaging | likely_pathogenic (0.89) | -9.62 | COSV105045685 |
| 75 | E→K | missense_variant | gnomAD | — | 6.84e-07 | damaging | likely_pathogenic (0.87) | -7.54 | chr17-48075067-C-T |
| 75 | E→E | synonymous_variant | COSMIC | — | — | — | — | 0.00 | COSV56681442 |
| 76 | P→P | synonymous_variant | gnomAD | — | 1.97e-04 | — | — | 0.00 | chr17-48075062-C-T |
| 76 | P→L | missense_variant | gnomAD | — | 7.52e-06 | damaging | likely_pathogenic (1.00) | -10.81 | chr17-48075063-G-A |
| 76 | P→P | synonymous_variant | COSMIC | — | — | — | — | 0.00 | COSV99848187 |
| 76 | P→Q | missense_variant | COSMIC | — | — | damaging | likely_pathogenic (1.00) | -12.12 | COSV56682499 |
| 76 | P→S | missense_variant | COSMIC | — | — | damaging | likely_pathogenic (1.00) | -9.69 | COSV99848079 |
| 77 | E→D | missense_variant | gnomAD | — | 1.37e-06 | damaging | likely_pathogenic (0.92) | -9.06 | chr17-48075059-C-A |
| 77 | E→E | synonymous_variant | gnomAD | — | 1.37e-06 | — | — | 0.00 | chr17-48075059-C-T |
| 77 | E→D | missense_variant | ClinVar | Uncertain significance | — | damaging | likely_pathogenic (0.92) | -9.06 | ClinVar:4648787 |
| 77 | E→K | missense_variant | COSMIC | — | — | damaging | likely_pathogenic (1.00) | -12.12 | COSV56681862 |
| 78 | R→R | synonymous_variant | gnomAD | — | 6.84e-07 | — | — | 0.00 | chr17-48075056-C-T |
| 78 | R→Q | missense_variant | COSMIC | — | — | damaging | likely_pathogenic (0.98) | -7.87 | COSV56682101 |
| 78 | R→L | missense_variant | COSMIC | — | — | damaging | likely_pathogenic (0.99) | -10.12 | COSV99848174 |
| 79 | I→V | missense_variant | gnomAD | — | 6.84e-07 | damaging | likely_pathogenic (0.96) | -7.69 | chr17-48075055-T-C |
| 81 | G→A | missense_variant | COSMIC | — | — | damaging | likely_pathogenic (1.00) | -12.44 | COSV56682437 |
| 83 | T→T | synonymous_variant | gnomAD | — | 2.74e-06 | — | — | 0.00 | chr17-48075041-T-C |
| 85 | S→S | synonymous_variant | gnomAD | — | 6.84e-07 | — | — | 0.00 | chr17-48075035-G-A |
| 86 | S→S | synonymous_variant | COSMIC | — | — | — | — | 0.00 | COSV56681832 |
| 88 | E→Q | missense_variant | gnomAD | — | 6.84e-07 | damaging | likely_pathogenic (0.92) | -10.00 | chr17-48075028-C-G |
| 89 | L→L | synonymous_variant | gnomAD | — | 9.59e-06 | — | — | 0.00 | chr17-48075023-G-A |
| 90 | M→L | missense_variant | gnomAD | — | 6.85e-07 | damaging | likely_pathogenic (0.91) | -10.25 | chr17-48075022-T-A |
| 90 | M→V | missense_variant | gnomAD | — | 6.85e-07 | damaging | likely_pathogenic (0.99) | -11.50 | chr17-48075022-T-C |
| 90 | — | inframe_deletion | COSMIC | — | — | — | — | — | COSV56681204 |
| 91 | F→F | synonymous_variant | gnomAD | — | 6.85e-07 | — | — | 0.00 | chr17-48075017-G-A |
| 92 | L→P | missense_variant | gnomAD | — | 6.85e-07 | damaging | likely_pathogenic (1.00) | -13.06 | chr17-48075015-A-G |
| 92 | L→L | synonymous_variant | COSMIC | — | — | — | — | 0.00 | COSV99848148 |
| 92 | L→L | synonymous_variant | COSMIC | — | — | — | — | 0.00 | COSV56681918 |
| 93 | M→I | missense_variant | gnomAD | — | 6.86e-07 | damaging | likely_pathogenic (0.97) | -8.12 | chr17-48075011-C-A |
| 93 | M→T | missense_variant | gnomAD | — | 6.86e-07 | damaging | likely_pathogenic (1.00) | -11.81 | chr17-48075012-A-G |
| 94 | K→K | synonymous_variant | gnomAD | — | 6.86e-07 | — | — | 0.00 | chr17-48075008-T-C |
| 94 | K→T | missense_variant | gnomAD | — | 6.86e-07 | damaging | likely_pathogenic (1.00) | -11.87 | chr17-48075009-T-G |
| 96 | K→R | missense_variant | gnomAD | — | 8.28e-06 | — | likely_benign (0.14) | -6.94 | chr17-48071577-T-C |
| 97 | N→K | missense_variant | COSMIC | — | — | damaging | likely_pathogenic (0.89) | -8.94 | COSV99848229 |
| 98 | S→F | missense_variant | COSMIC | — | — | damaging | likely_pathogenic (0.99) | -10.37 | COSV56681364 |
| 101 | A→A | synonymous_variant | gnomAD | — | 6.87e-07 | — | — | 0.00 | chr17-48071561-A-G |
| 102 | D→D | synonymous_variant | gnomAD | — | 6.87e-07 | — | — | 0.00 | chr17-48071558-G-A |
| 104 | V→L | missense_variant | gnomAD | — | 6.86e-07 | damaging | likely_pathogenic (1.00) | -11.25 | chr17-48071554-C-G |
| 106 | A→A | synonymous_variant | gnomAD | — | 6.85e-07 | — | — | 0.00 | chr17-48071546-G-A |
| 106 | A→D | missense_variant | gnomAD | — | 6.85e-07 | damaging | likely_pathogenic (1.00) | -12.00 | chr17-48071547-G-T |
| 106 | A→V | missense_variant | COSMIC | — | — | damaging | likely_pathogenic (1.00) | -10.62 | COSV99848117 |
| 108 | E→* | stop_gained | gnomAD | — | 6.85e-07 | LoF | — | — | chr17-48071542-C-A |
| 108 | E→K | missense_variant | COSMIC | — | — | damaging | likely_pathogenic (0.99) | -10.94 | COSV56681808 |
| 110 | N→S | missense_variant | gnomAD | — | 6.84e-07 | damaging | likely_pathogenic (0.91) | -12.44 | chr17-48071535-T-C |
| 111 | V→F | missense_variant | gnomAD | — | 2.74e-06 | damaging | likely_pathogenic (0.57) | -8.73 | chr17-48071533-C-A |
| 111 | V→V | synonymous_variant | COSMIC | — | — | — | — | 0.00 | COSV56681933 |
| 112 | K→K | synonymous_variant | COSMIC | — | — | — | — | 0.00 | COSV56682664 |
| 113 | C→C | synonymous_variant | gnomAD | — | 1.37e-06 | — | — | 0.00 | chr17-48071525-G-A |
| 113 | C→S | missense_variant | gnomAD | — | 6.84e-07 | damaging | likely_pathogenic (1.00) | -5.97 | chr17-48071526-C-G |
| 114 | P→L | missense_variant | gnomAD | — | 6.84e-07 | damaging | likely_pathogenic (1.00) | -10.62 | chr17-48071523-G-A |
| 115 | Q→Q | synonymous_variant | gnomAD | — | 6.84e-07 | — | — | 0.00 | chr17-48071519-C-T |
| 115 | Q→* | stop_gained | gnomAD | — | 6.84e-07 | LoF | — | — | chr17-48071521-G-A |
| 116 | V→F | missense_variant | gnomAD | — | 2.05e-06 | damaging | likely_pathogenic (0.99) | -12.05 | chr17-48071518-C-A |
| 118 | I→I | synonymous_variant | COSMIC | — | — | — | — | 0.00 | COSV56681722 |
| 118 | I→T | missense_variant | COSMIC | — | — | damaging | likely_pathogenic (1.00) | -11.06 | COSV56682067 |
| 119 | S→S | synonymous_variant | gnomAD | — | 2.05e-06 | — | — | 0.00 | chr17-48071507-G-T |
| 120 | F→F | synonymous_variant | gnomAD | — | 6.84e-07 | — | — | 0.00 | chr17-48071504-G-A |
| 121 | Y→Y | synonymous_variant | gnomAD | — | 1.37e-06 | — | — | 0.00 | chr17-48071501-A-G |
| 121 | Y→C | missense_variant | gnomAD | — | 6.84e-07 | damaging | likely_pathogenic (1.00) | -11.94 | chr17-48071502-T-C |
| 124 | R→K | missense_variant | COSMIC | — | — | damaging | likely_pathogenic (0.99) | -11.06 | COSV56682552 |
| 126 | T→T | synonymous_variant | gnomAD | — | 4.11e-06 | — | — | 0.00 | chr17-48071486-C-T |
| 126 | T→K | missense_variant | COSMIC | — | — | damaging | likely_pathogenic (0.99) | -13.94 | COSV56681836 |
| 128 | H→H | synonymous_variant | gnomAD | — | 6.84e-07 | — | — | 0.00 | chr17-48071480-A-G |
| 129 | S→S | synonymous_variant | gnomAD | — | 6.84e-07 | — | — | 0.00 | chr17-48071477-G-A |
| 129 | S→F | missense_variant | gnomAD | — | 6.84e-07 | damaging | likely_pathogenic (0.97) | -11.12 | chr17-48071478-G-A |
| 129 | S→A | missense_variant | gnomAD | — | 6.84e-07 | damaging | likely_benign (0.27) | -9.06 | chr17-48071479-A-C |
| 129 | S→P | missense_variant | gnomAD | — | 6.84e-07 | damaging | likely_pathogenic (0.98) | -11.81 | chr17-48071479-A-G |
| 129 | S→T | missense_variant | gnomAD | — | 6.84e-07 | damaging | ambiguous (0.44) | -10.19 | chr17-48071479-A-T |
| 130 | Y→Y | synonymous_variant | gnomAD | — | 2.94e-05 | — | — | 0.00 | chr17-48071474-G-A |
| 130 | Y→* | stop_gained | gnomAD | — | 6.85e-07 | LoF | — | — | chr17-48071474-G-T |
| 131 | P→S | missense_variant | gnomAD | — | 6.85e-07 | damaging | likely_pathogenic (0.58) | -8.75 | chr17-48071473-G-A |
| 131 | P→S | missense_variant | COSMIC | — | — | damaging | likely_pathogenic (0.58) | -8.75 | COSV108798400 |
| 132 | S→S | synonymous_variant | gnomAD | — | 4.11e-06 | — | — | 0.00 | chr17-48071468-C-T |
| 132 | S→L | missense_variant | gnomAD | — | 2.05e-06 | damaging | likely_benign (0.21) | -8.42 | chr17-48071469-G-A |
| 132 | S→L | missense_variant | ClinVar | Uncertain significance | — | damaging | likely_benign (0.21) | -8.42 | ClinVar:3138046 |
| 132 | S→S | synonymous_variant | COSMIC | — | — | — | — | 0.00 | COSV56681272 |
| 132 | S→* | stop_gained | COSMIC | — | — | LoF | — | — | COSV99848014 |
| 133 | E→K | missense_variant | COSMIC | — | — | damaging | likely_pathogenic (0.75) | -12.00 | COSV104555665 |
| 134 | D→E | missense_variant | gnomAD | — | 4.80e-06 | — | likely_benign (0.09) | -5.24 | chr17-48071462-A-C |
| 134 | D→D | synonymous_variant | gnomAD | — | 6.85e-07 | — | — | 0.00 | chr17-48071462-A-G |
| 134 | D→Y | missense_variant | gnomAD | — | 6.85e-07 | damaging | likely_pathogenic (0.71) | -11.56 | chr17-48071464-C-A |
| 134 | D→N | missense_variant | gnomAD | — | 1.37e-06 | damaging | likely_benign (0.31) | -10.18 | chr17-48071464-C-T |
| 134 | D→E | missense_variant | ClinVar | Uncertain significance | — | — | likely_benign (0.09) | -5.24 | ClinVar:2517719 |
| 136 | — | inframe_deletion | gnomAD | — | 2.06e-06 | — | — | — | chr17-48071456-GTCA-G |
| 136 | D→N | missense_variant | COSMIC | — | — | damaging | likely_benign (0.19) | -8.62 | COSV99848218 |
| 137 | K→N | missense_variant | COSMIC | — | — | damaging | likely_pathogenic (0.72) | -9.62 | COSV99848111 |
| 137 | — | frameshift_variant | COSMIC | — | — | LoF | — | — | COSV56682928 |
| 138 | K→* | stop_gained | gnomAD | — | 6.85e-07 | LoF | — | — | chr17-48071452-T-A |
| 139 | D→V | missense_variant | gnomAD | — | 6.86e-07 | damaging | ambiguous (0.45) | -9.47 | chr17-48071448-T-A |
| 139 | — | frameshift_variant | gnomAD | — | 6.86e-07 | LoF | — | — | chr17-48071449-C-CT |
| 139 | — | frameshift_variant | gnomAD | — | 6.86e-07 | LoF | — | — | chr17-48071449-CTTTT-C |
| 139 | D→H | missense_variant | COSMIC | — | — | damaging | likely_pathogenic (0.73) | -10.72 | COSV56682228 |
| 139 | D→Y | missense_variant | COSMIC | — | — | damaging | likely_pathogenic (0.65) | -9.97 | COSV56681987 |
| 140 | — | frameshift_variant | gnomAD | — | 6.86e-07 | LoF | — | — | chr17-48071445-TC-T |
| 140 | D→Y | missense_variant | gnomAD | — | 1.37e-06 | damaging | likely_pathogenic (0.68) | -10.43 | chr17-48071446-C-A |
| 140 | D→N | missense_variant | gnomAD | — | 6.86e-07 | — | likely_benign (0.27) | -7.43 | chr17-48071446-C-T |
| 140 | — | frameshift_variant | gnomAD | — | 6.86e-07 | LoF | — | — | chr17-48071446-CA-C |
| 142 | N→K | missense_variant | gnomAD | — | 6.87e-07 | damaging | likely_pathogenic (0.68) | -7.21 | chr17-48071438-G-C |
| 142 | — | frameshift_variant | gnomAD | — | 1.37e-06 | LoF | — | — | chr17-48071439-TTCTTG-T |
| 142 | N→K | missense_variant | COSMIC | — | — | damaging | likely_pathogenic (0.68) | -7.21 | COSV56682095 |
205 variants in the shared canonical core.
LoF = frameshift / stop-gain / splice-disrupting — inherently loss-of-function, flagged by consequence (AlphaMissense and ESM-C score only missense/substitutions, so they are blank here by design, not by absence of impact). AlphaMissense is computed in the canonical reading frame, so it scores the shared core but reads N/A across an isoform-unique extension — use ESM-C ΔLLR there.