Skip to main content
10% off your first order — code WELCOME10 at checkout

Search

Peptide Sequence Notation: Three-Letter, One-Letter and Modified Residues

Peptide Sequence Notation: Three-Letter, One-Letter and Modified Residues

Last reviewed 9 August 2026

The identification table in a library entry gives a compound’s sequence twice, once in three-letter symbols and once in one. Ac-Leu-Lys-Lys-Thr-Glu-Thr-Gln-OH and Ac-LKKTETQ are the same seven residues under two published conventions, and neither is a house style. Almost every character in such a string is load-bearing: a hyphen is a chemical bond, a prefix is a functional group, and an absent prefix is itself an assertion. Two compounds in this catalogue are separated from each other by exactly one terminal symbol.

The governing text is Nomenclature and Symbolism for Amino Acids and Peptides, the 1983 recommendations of the IUPAC-IUB Joint Commission on Biochemical Nomenclature, reprinted in European Journal of Biochemistry in 1984 and elsewhere. It runs in three parts — nomenclature, symbolism, and modification of named peptides — extended by commission newsletters in 1985, 1986, 1989, 1999 and 2009. Rule numbers below are from that document.

Two systems, built for different jobs

The three-letter system occupies rules 3AA-14 to 3AA-19; the one-letter system occupies 3AA-20 and 3AA-21. They are not interchangeable, and the recommendation says so. The one-letter code was approved in 1968 on proposals from a subcommittee of W. E. Cohn, M. O. Dayhoff, R. V. Eck and B. Keil, and the 1983 text carries it forward “with no substantial change”. Its purpose is compression, and rule 3AA-20.2 limits it accordingly: it “is less easily understood than the three-letter system by those not familiar with it, so it should not be used in simple text or in reporting experimental details of sequence determination”, and is recommended instead for comparing long sequences in tables and lists.

The letter assignments are not arbitrary, and the reasoning explains the gaps. Six amino acids took their own initial uncontested; where names shared an initial the letter went to the most frequent and structurally simplest, F and R were chosen phonetically, and tryptophan took W because, in the commission’s words, “the double ring of the molecule is associated with the bulky letter W”. U and O were deliberately avoided, U because it is confused with V in handwriting and O with G, Q, C, D and zero; J was avoided as absent from several languages.

One-letter symbols, in alphabetical order of symbol. Sec and Pyl were added by commission newsletters in 1999 and 2009; Xle is a sequence-database code rather than a 1983 assignment.
One-letterThree-letterAmino acid or meaning
AAlaalanine
BAsxaspartic acid or asparagine
CCyscysteine
DAspaspartic acid
EGluglutamic acid
FPhephenylalanine
GGlyglycine
HHishistidine
IIleisoleucine
JXleleucine or isoleucine, not distinguished
KLyslysine
LLeuleucine
MMetmethionine
NAsnasparagine
OPylpyrrolysine
PProproline
QGlnglutamine
RArgarginine
SSerserine
TThrthreonine
USecselenocysteine
VValvaline
WTrptryptophan
XXaaunknown or atypical amino acid
YTyrtyrosine
ZGlxglutamic acid or glutamine, and substances such as 5-oxoproline that yield glutamic acid on acid hydrolysis

B, Z and X are ambiguity symbols from the original set, for positions analysis has not resolved. The 1999 newsletter recommended “Sec as the three-letter symbol, and U as the one-letter symbol, for selenocysteine”, noting that J was unavailable because it was already used in NMR work for leucine and isoleucine signals, and conceding that “it is a disadvantage that U also stands for uracil”. The 2009 newsletter added Pyl and O for pyrrolysine. Xle and J are a database convention, recorded in the INSDC sequence-description codes.

Direction, and what a hyphen is doing

Both systems read left to right from the free amino group; rule 3AA-21.1 states it for the one-letter code. In the three-letter system the hyphen is not punctuation. Rule 3AA-16.1 works it through: a hyphen on the right of a symbol removes the hydroxyl from the 1-carboxyl group, a hyphen on the left removes a hydrogen from the 2-amino group, and both can apply at once. So Gly- is an acyl group, -Gly is a residue with a free carboxyl, and -Gly- is an internal residue — which is why Gly-Glu is a distinct thing from -Gly-Glu-, a fragment of something larger.

Rule 3AA-17.6 lets the ends be made explicit where it matters: H-Ala contrasts with Ac-Ala, and Ala-OH with Ala-OMe. The fifteen-residue sequence in the BPC-157 entry can be written Gly-Glu-Pro-Pro-Pro-Gly-Lys-Pro-Ala-Asp-Asp-Ala-Gly-Leu-Val or H-Gly-…-Val-OH. Same molecule; the second form says out loud that neither terminus is modified.

Acetyl at one end, amide at the other

Ac- for acetyl appears in the recommendation’s table of non-urethane substituents for nitrogen, oxygen or sulfur. The C-terminal amide is covered by the 1985 newsletter, which gives Gly-NH2 for glycine amide; the main text’s own worked example is thyroliberin, Glp-His-Pro-NH2, a three-residue string carrying both a non-standard residue and a terminal amide.

Both are common here. Esposito and colleagues identified the N-terminally acetylated 17–23 fragment of thymosin β4 (Ac-LKKTETQ) in TB-500 by high-performance liquid chromatography with high-resolution mass spectrometry, and separately synthesised that fragment by solid-phase peptide synthesis as a reference standard; the library entry gives it as Ac-Leu-Lys-Lys-Thr-Glu-Thr-Gln-OH. Birk and colleagues give the structure of SS-31 as “D-Arg-Dmt-Lys-Phe-NH2”, which ends in an amide rather than a free acid; see the SS-31 entry.

Neither modification alters the residue list, and both alter the mass an analyst confirms against. Acetylation replaces one hydrogen on the terminal amino nitrogen with an acetyl group, a net addition of C2H2O; on the NIST relative atomic masses of the principal isotopes — 12C exactly 12, 1H 1.00782503223, 14N 14.00307400443, 16O 15.99491461957 — that is +42.0106 to the monoisotopic mass. Amidation exchanges a terminal hydroxyl for an amino group: 16.01872 in, 17.00274 out, a net −0.9840. A sequence quoted without its terminal symbols will not reconcile with a measured mass, which is one of the things a certificate of analysis is read for.

The clearest illustration in this catalogue is a pair. PubChem records melanotan II as C50H69N15O9, CAS 121062-08-6, with the symbol string Ac-Nle-cyclo[Asp-His-D-Phe-Arg-Trp-Lys]-NH2. It records bremelanotide as C50H68N14O10, CAS 189691-06-3: the identical ring, written as a (2→7) lactam and ending in a free carboxylic acid rather than an amide. One terminal symbol separates two registered compounds.

L, D, and what the one-letter system cannot record

Rule 3AA-14.5 is unambiguous: “Amino-acid symbols denote the L configuration of chiral amino acids unless otherwise indicated by the presence of D or DL before the symbol and separated from it with a hyphen.” Rule 3AA-19.2 adds that the hyphen may be dropped to keep residue counts legible, and that DL denotes a racemate and so should not appear in a peptide with more than one chiral residue. Part 3 supplies retro– for a reversed sequence and ent– for the whole enantiomer.

All of that lives in the three-letter system. The one-letter rules, read end to end, contain no provision for configuration at all — nothing in 3AA-21 marks a D residue. Lower-case letters are a convention of some databases and drawing tools, not a rule in the recommendation, and they need defining wherever they appear. Hence a table offering only a one-letter string is under-specified for a compound like ipamorelin, given by Raun and colleagues as Aib-His-D-2-Nal-D-Phe-Lys-NH2: two of its five residues are D, one is achiral, two are not proteinogenic amino acids, and the C-terminus is an amide.

Residues outside the twenty

Rule 3AA-15.2 sets a standing obligation: “Symbols for less common amino acids should be defined in each publication in which they appear.” Four examples show why.

  • Dmt. Birk and colleagues define it in the paper itself — “Dmt = 2′,6′-dimethylTyr” — a tyrosine bearing two ring methyl groups. No one-letter symbol exists, so it is written in full in both notations.
  • Nle. Norleucine, in both melanocortin compounds above. Rule 3AA-15.2.3 recommends that this use of ‘nor’ “should be progressively abandoned” along with the symbols Nva and Nle, in favour of Ape and Ahx. Four decades on, neither the literature nor the chemical registries have followed. A recommendation and settled practice can diverge, and a reader meets both.
  • Aib. 2-Amino-2-methylpropanoic acid, PubChem CID 6119, C4H9NO2. Its alpha carbon carries two methyl groups, so it is not a stereocentre and no D or L prefix can apply to it.
  • Nal. Raun and colleagues write D-2-Nal; Hruby and colleagues write D-2′-naphthylalanine for the same class of residue. Different locant styles for one idea, which is the situation 3AA-15.2 exists to contain.

Substitution on a side chain or a ring uses parentheses immediately after the residue symbol, with locants where context does not supply them: 3,5-diiodotyrosine is Tyr(3,5-I2). The position of those parentheses carries meaning, and the recommendation flags the trap outright — Asp-OMe is the ester at C-1, Asp(OMe) the side-chain ester. The 2009 newsletter extends the logic: Lys(Me)- puts the methyl on lysine’s own side-chain nitrogen, Lys-(Me)Ala- puts it on the nitrogen of the following alanine.

Rings, and naming by reference to a parent

A cyclic sequence is written with the residues in parentheses preceded by cyclo. Where the ring is made only of ordinary peptide bonds it is homodetic; where a linkage is an isopeptide, disulfide or ester bond it is heterodetic. The melanocortin ring above is the second kind, the registry name recording a cyclic (2→7) peptide, or a (2→7) lactam, closed between two side chains rather than through the backbone. A one-letter string cannot express that.

Part 3 of the recommendation covers compounds named by reference to a parent sequence: square brackets with a superscript position for a replacement, as in [Cit8]vasopressin; des– for a removal; endo– for an insertion; and a fragment written as parent-(p–q)-peptide. The title of Hruby and colleagues’ 1995 paper contains “Ac-Nle4-cyclo[Asp5, D-Phe7,Lys10] alpha-melanocyte-stimulating hormone-(4-10)-NH2”: an N-terminal acetyl, norleucine at position 4, a ring closed across positions 5 and 10, D-phenylalanine at 7, a span of residues 4 to 10 of the parent, and a C-terminal amide. Every element is one of the rules above, applied once.

The same device is also used loosely. Magrì and colleagues title Semax “a synthetic analog of ACTH(4-10)”, then open the abstract by calling it “a heptapeptide (Met-Glu-His-Phe-Pro-Gly-Pro) that encompasses the sequence 4-7 of N-terminal domain of the adrenocorticotropic hormone and a C-terminal Pro-Gly-Pro tripeptide”. Both are accurate and describe one molecule, but they cite different ranges, because a bare fragment range identifies a segment of a parent rather than the compound in hand. Only the residue list settles it — Kost and colleagues give it for Semax and for Selank, Thr-Lys-Pro-Arg-Pro-Gly-Pro, in one sentence. The two are compared separately.

What a sequence string does not carry

  • Salt form. Notation records the peptide, not the counterion — and the counterion is part of what is weighed.
  • Purity and net peptide content. Neither is an identity attribute and neither appears in the sequence. Both belong to the certificate.
  • A coordinated metal. GHK is Gly-His-Lys, PubChem CID 73587, C14H24N6O4. The copper in the designation GHK-Cu has no residue symbol and no position in the string, because a sequence is a statement about a peptide and not about a complex.
  • The leucine-isoleucine distinction, by mass alone. Both are C6H13NO2 (PubChem CIDs 6106 and 6306) and therefore isobaric, which is why Xle exists. Configuration away from the alpha carbon is a second gap: threonine and isoleucine each carry a further stereocentre, and the recommendation deals with their allo forms separately, as aThr and aIle.
  • A mixture. A blended product has no single sequence, only a list of separate molecules and their masses — the subject of a separate article.

Read in order — direction, then both termini, then any D prefix, then any symbol outside the twenty, then any ring — a sequence line resolves into a set of checkable statements, each testable against a stated formula and a stated mass. Terms used above are collected in the glossary, and the structural distinction between two of the compounds named here is drawn out separately.

References

  1. IUPAC-IUB Joint Commission on Biochemical Nomenclature (JCBN). Nomenclature and symbolism for amino acids and peptides. Recommendations 1983. Eur J Biochem. 1984;138(1):9–37. Nomenclature recommendation. PMID 6692818. DOI 10.1111/j.1432-1033.1984.tb07877.x.
  2. IUPAC-IUB JCBN and NC-IUB. Newsletter 1985, item on amide derivatives of amino acids and peptides. Eur J Biochem. 1985;146:237–239. Nomenclature recommendation.
  3. IUPAC-IUB JCBN and NC-IUBMB. Newsletter 1999, item on selenocysteine. Eur J Biochem. 1999;264:607–609. Nomenclature recommendation.
  4. IUPAC JCBN and NC-IUBMB. Newsletter 2009, items on pyrrolysine and on symbolism for di-substituted amino-acid residues. Nomenclature recommendation.
  5. DNA Data Bank of Japan, for the INSDC. Codes used in sequence description: amino acid codes. Sequence-database standard.
  6. National Institute of Standards and Technology. Atomic Weights and Isotopic Compositions. Reference data.
  7. Birk AV, Liu S, Soong Y, et al. The mitochondrial-targeted compound SS-31 re-energizes ischemic mitochondria by interacting with cardiolipin. J Am Soc Nephrol. 2013;24(8):1250–61. Model: in vitro and rodent. PMID 23813215. DOI 10.1681/ASN.2012121216.
  8. Esposito S, Deventer K, Goeman J, Van der Eycken J, Van Eenoo P. Synthesis and characterization of the N-terminal acetylated 17-23 fragment of thymosin beta 4 identified in TB-500, a product suspected to possess doping potential. Drug Test Anal. 2012;4(9):733–8. Model: analytical chemistry. PMID 22962027. DOI 10.1002/dta.1402.
  9. Raun K, Hansen BS, Johansen NL, Thøgersen H, Madsen K, Ankersen M, Andersen PH. Ipamorelin, the first selective growth hormone secretagogue. Eur J Endocrinol. 1998;139(5):552–61. Model: in vitro (primary rat pituitary cells), rodent (anaesthetised rats) and non-rodent animal (conscious swine). PMID 9849822. DOI 10.1530/eje.0.1390552.
  10. Hruby VJ, Lu D, Sharma SD, et al. Cyclic lactam alpha-melanotropin analogues of Ac-Nle4-cyclo[Asp5, D-Phe7,Lys10] alpha-melanocyte-stimulating hormone-(4-10)-NH2 with bulky aromatic amino acids at position 7 show high antagonist potency and selectivity at specific melanocortin receptors. J Med Chem. 1995;38(18):3454–61. Model: ex vivo amphibian tissue (classical Rana pipiens frog skin assay) and in vitro cell lines expressing cloned human and mouse melanocortin receptors. PMID 7658432. DOI 10.1021/jm00018a005.
  11. Magrì A, Tabbì G, Giuffrida A, et al. Influence of the N-terminus acetylation of Semax, a synthetic analog of ACTH(4-10), on copper(II) and zinc(II) coordination and biological properties. J Inorg Biochem. 2016;164:59–69. Model: in vitro (copper(II) and zinc(II) coordination chemistry; SH-SY5Y neuroblastoma cell line). PMID 27586814. DOI 10.1016/j.jinorgbio.2016.08.013.
  12. Kost NV, Sokolov OIu, Gabaeva MV, Grivennikov IA, Andreeva LA, Miasoedov NF, Zozulia AA. Semax and selank inhibit the enkephalin-degrading enzymes from human serum. Bioorg Khim. 2001;27(3):180–3. Model: in vitro, human serum enzymes. PMID 11443939. DOI 10.1023/a:1011373002885.
  13. National Center for Biotechnology Information. PubChem Compound Summaries for CID 92432 (melanotan II), CID 9941379 (bremelanotide), CID 6119 (2-aminoisobutyric acid), CID 73587 (glycyl-L-histidyl-L-lysine), CID 6106 (L-leucine) and CID 6306 (L-isoleucine). Chemical registry records.
Back to Top
Product has been added to your cart
Compare (0)