← SkillSafe / Sequon Desk

Every sequon, every liability motif, and the call on each one

Paste the sequence you are about to order DNA for. Before anything leaves this page, your browser reads it: mass, theoretical pI, net charge, extinction coefficient, every N-X-S/T sequon with the proline-blocked near-misses kept separate, the O-glycosylation hotspots, the deamidation and oxidation motifs, cysteine parity, and every hydrophobic and charge patch. Then pick a lane.

Both examples ship with a saved model run for every lane, so you can see all three documents without signing in and without spending a credit.

nothing pasted yet
Drag a .fasta or .txt in, or Nothing is uploaded until you press a lane's run button.

One chain or several. FASTA headers, residue numbering, whitespace, gap characters and a trailing stop codon are all handled; a stop codon in the middle is reported as a truncated construct rather than removed.

These four change the grading, not just the wording: a sequon in an E. coli construct is information rather than a liability, and a C-terminal lysine on an IgG is expected heterogeneity rather than a defect. The prescan applies them before anything is sent.

What else you know about this molecule
Paste a sequence to price the run.

What this does, and what it cannot do

The prescan is a real sequence reader, not a keyword search. It walks the paste character by character so that every position it reports is a residue index in the ungapped chain: FASTA headers, embedded residue numbering, whitespace and alignment gaps are removed and counted, a trailing stop codon is dropped as the translated stop it is, and a stop codon in the middle is reported as a truncated construct instead. The glycosylation rule is N-X-S/T with X not proline, which is why NPS and NPT appear in their own list as near-sequons rather than as sites — a rule written as N.[ST] would call them glycosylated, and they never are. N-X-C is listed separately again, as something to check rather than as a site.

Grading depends on the four context fields and on the numbers the scan already computed. A sequon in an E. coli construct is information, because there is no N-glycosylation machinery to use it. A C-terminal lysine on an IgG is expected, characterised heterogeneity. A hydrophobic patch is graded against the net charge density of the same chain, because hydrophobicity and charge only mean something read together. And cysteine parity is reported for what it is: an even count is consistent with complete pairing and is not evidence of it, while a single cysteine is definitely unpaired.

What it cannot do: a sequence carries no structure. Solvent exposure, CDR boundaries, epitope overlap and real glycan occupancy are all unknown here, and every lane is instructed to name the cheapest experiment or prediction that would settle a call rather than assert one. No numbering scheme is assigned. The O-glycosylation list is a proline-context heuristic, not a predictor. The theoretical pI comes from the Bjellqvist pKa set — the one ExPASy ProtParam uses — and can differ from ProtParam by a couple of tenths because ProtParam also applies a position-specific N-terminal value. Nothing here is therapeutic, clinical or regulatory advice, and nothing here is executed: the page reads text and does arithmetic on it.

Nothing to hand? Load the , which carries a CDR sequon, a proline-blocked near-sequon, an NG motif and an odd cysteine count, or the , where the sequons are information and the hydrophobic patch is the real finding. Both replay a saved run in every lane, for free.