Every sequon, every liability motif, and the call on each one
Paste the sequence you are about to order DNA for. Before anything leaves this page, your browser reads it: mass, theoretical pI, net charge, extinction coefficient, every N-X-S/T sequon with the proline-blocked near-misses kept separate, the O-glycosylation hotspots, the deamidation and oxidation motifs, cysteine parity, and every hydrophobic and charge patch. Then pick a lane.
Both examples ship with a saved model run for every lane, so you can see all three documents without signing in and without spending a credit.
What this does, and what it cannot do
The prescan is a real sequence reader, not a keyword search. It walks the paste character by
character so that every position it reports is a residue index in the ungapped chain: FASTA
headers, embedded residue numbering, whitespace and alignment gaps are removed and counted, a
trailing stop codon is dropped as the translated stop it is, and a stop codon in the
middle is reported as a truncated construct instead. The glycosylation rule is
N-X-S/T with X not proline, which is why NPS and
NPT appear in their own list as near-sequons rather than as sites — a rule written
as N.[ST] would call them glycosylated, and they never are. N-X-C is
listed separately again, as something to check rather than as a site.
Grading depends on the four context fields and on the numbers the scan already computed. A sequon in an E. coli construct is information, because there is no N-glycosylation machinery to use it. A C-terminal lysine on an IgG is expected, characterised heterogeneity. A hydrophobic patch is graded against the net charge density of the same chain, because hydrophobicity and charge only mean something read together. And cysteine parity is reported for what it is: an even count is consistent with complete pairing and is not evidence of it, while a single cysteine is definitely unpaired.
What it cannot do: a sequence carries no structure. Solvent exposure, CDR boundaries, epitope overlap and real glycan occupancy are all unknown here, and every lane is instructed to name the cheapest experiment or prediction that would settle a call rather than assert one. No numbering scheme is assigned. The O-glycosylation list is a proline-context heuristic, not a predictor. The theoretical pI comes from the Bjellqvist pKa set — the one ExPASy ProtParam uses — and can differ from ProtParam by a couple of tenths because ProtParam also applies a position-specific N-terminal value. Nothing here is therapeutic, clinical or regulatory advice, and nothing here is executed: the page reads text and does arithmetic on it.
Nothing to hand? Load the , which carries a CDR sequon, a proline-blocked near-sequon, an NG motif and an odd cysteine count, or the , where the sequons are information and the hydrophobic patch is the real finding. Both replay a saved run in every lane, for free.