Educational resource
Sequence notation guide
Every compound record stores two related sequence fields. They serve different jobs: one is compact and searchable; the other is human-readable and chemically explicit.
1-letter sequence
The Sequence (1-letter) field uses the IUPAC/IUBMB one-letter amino-acid alphabet (A, C, D, E, F, G, H, I, K, L, M, N, P, Q, R, S, T, V, W, Y). It is the primary string used by site search and sequence filters.
- Spaces may separate chains (for example insulin A and B).
- Only uppercase L-proteinogenic letters are expected in the searchable core.
- Non-canonical residues that lack a standard letter are shown as
X; the display field explains whatXstands for. - Hyphens, three-letter names, and terminal annotations are stripped from this field so matching stays predictable.
Example (oxytocin core): CYIQNCPLG.
Display notation
The Sequence notation (display) field is the authoritative human-readable form. It may include chemistry that cannot be expressed in twenty letters alone:
- C-terminal amide as
-NH2, or ethylamide as-NHEt - N-terminal pyroglutamate as
pGluorpE - Disulfides as
Cys1–Cys6(residue numbers in the mature peptide) - D-residues as a leading
dorD-prefix (dF,D-Trp) - Side-chain adducts in parentheses on the modified residue, e.g.
K(Nε-γ-Glu-palmitoyl) - Reduced termini such as threoninol as
-ol - Non-proteinogenic residues by name (Aib, Orn, MeBmt, Mpr, and others)
Example (oxytocin display): Cys-Tyr-Ile-Gln-Asn-Cys-Pro-Leu-Gly-NH2 with disulfide Cys1–Cys6.
Reading order and termini
Sequences are written N→C (amino terminus on the left, carboxyl or modified C-terminus on the right), matching standard peptide literature. Free N-termini are usually left unmarked. Free C-termini may be unmarked or written as -OH when contrast with an amidated analogue matters.
What the 1-letter field is not
It is not a SMILES string, not a FASTA header, not a full chemical graph, and not a substitute for stereochemistry. Protecting groups, counter-ions, hydrates, and formulation salts are out of scope unless listed under modifications or structure notes. When display notation and the 1-letter code disagree, trust the display notation and the modifications taxonomy.
Chains, rings, and topology
A polymer can be written as one or more linear residue strings and still be classified cyclic if a covalent ring exists (disulfide, thioether, side-chain lactam, or head-to-tail amide). Insulin is stored as two chains and classified cyclic because of inter- and intrachain disulfides. Residue numbering for disulfide pairs follows the mature peptide as cited in the source, not the prohormone precursor.
Practical checklist
- Use 1-letter for search and length counting of standard residues.
- Use display notation to confirm termini, stereochemistry, and adducts.
- Use modification tags to filter amidated, disulfide-bridged, or lipidated sets.
- Confirm identity with CAS / PubChem before analytical work.