Peptide Codons: Why a Codon Specifies an Amino Acid, Not a Peptide
The phrase peptide codon is loose speech, and the correction is short enough to state in one line: a codon is a triplet of nucleotides that specifies a single amino acid or a stop signal, while a peptide is the product of a run of such triplets read in one frame. Nothing in the genetic code names a peptide as a unit. A five-residue peptide requires five sense codons plus a stop, and the mapping from peptide back to nucleotide sequence is one-to-many rather than one-to-one, which is the fact that trips people up when they try to reverse the table.
Getting the direction right matters for three practical reasons. It explains why the code is described as degenerate rather than ambiguous, it explains why the same peptide can be encoded by many different messenger sequences, and it explains why a change in the third position of a codon often changes nothing in the product. The residue-level conventions used below are set out in our peptide structure and classification reference, and the mechanical side of reading triplets is covered in how transfer RNA reads a codon during synthesis.
Sixty-four triplets, sixty-one sense, three stops
Four bases taken three at a time give 64 possible triplets, and the standard code assigns 61 of them to the 20 standard residues and leaves three as termination signals: UAA, UAG and UGA, conventionally named ochre, amber and opal. No transfer RNA reads a stop; release factors recognize them instead and hydrolyze the completed chain from the final transfer RNA. AUG is the usual start triplet and codes for methionine, so the same triplet both marks where translation begins and contributes the first residue, which is why many mature proteins do not begin with methionine once post-translational processing has removed it.
Degeneracy is organized, not random. Where a residue has four codons they usually share the first two bases and differ only in the third, and the third position is where the wobble hypothesis applies: the pairing between the third base of the codon and the first base of the anticodon is geometrically relaxed, and inosine in the anticodon widens it further, so one transfer RNA can read several triplets. Two documented recoding exceptions exist as well, where UGA can specify selenocysteine and UAG pyrrolysine given the right sequence context, which is worth knowing so that a table that shows only three stops is not mistaken for the whole story.
| Grouping | Codons | Amino acids | Notes |
|---|---|---|---|
| Fourfold degenerate families | GCN, CCN, ACN, GGN, GUN | Ala, Pro, Thr, Gly, Val | Third base is silent; one tRNA often reads all four |
| Twofold families | AAA and AAG, GAA and GAG, UUU and UUC | Lys, Glu, Phe and others | Purine ending versus pyrimidine ending distinguishes the pair |
| Six codon families | UUA, UUG, CUN; UCN, AGU, AGC; CGN, AGA, AGG | Leu, Ser, Arg | A four codon block plus a separate two codon block |
| Three codons | AUU, AUC, AUA | Ile | One residue, three triplets, no fourfold pattern |
| Single codon residues | AUG, UGG | Met, Trp | Also the start triplet in the case of AUG |
| Termination signals | UAA, UAG, UGA | None | Read by release factors, not by tRNA |
| Recoded exceptions | UGA, UAG in context | Selenocysteine, pyrrolysine | Requires specific sequence elements in the messenger |
A worked translation with residue masses
Take a short messenger stretch and read it in triplets from the start: 5 prime AUG GCU UAC GGU UUU UAA 3 prime. The codons translate to methionine, alanine, tyrosine, glycine and phenylalanine, giving the pentapeptide Met Ala Tyr Gly Phe, written MA YGF in one-letter code, and the sixth triplet is a stop that adds nothing to the chain. This is the concrete version of the point made at the top: one codon, one residue, and five residues from five sense triplets plus one termination triplet. Nothing in the codon table produces a peptide in a single step.
The mass arithmetic follows directly. Summing monoisotopic residue masses gives 131.04049 for methionine, 71.03711 for alanine, 163.06333 for tyrosine, 57.02146 for glycine and 147.06841 for phenylalanine, a running total of 569.23080 daltons for the five residues. A peptide of five residues has four peptide bonds and retains one water molecule, so the neutral mass is 569.23080 plus 18.01056, or about 587.24 daltons. The usual shortcut of 110 daltons per residue gives 568 here, about 19 daltons low, which is a fair illustration of why the average figure is an estimate and not a calculation.
Shifting the frame shows what the start triplet is doing. Read the same string one base later and the triplets become UGG CUU ACG GUU UUU, which translate to a completely different run of residues. The frame, set by where translation begins, is therefore part of the information, and an insertion or deletion that is not a multiple of three changes every residue downstream of it. This is the sense in which a peptide is encoded by a sequence of codons rather than by any individual codon.
| Position | Codon | Residue | One letter | Residue mass in Da | Running total in Da |
|---|---|---|---|---|---|
| 1 | AUG | Methionine | M | 131.04049 | 131.04049 |
| 2 | GCU | Alanine | A | 71.03711 | 202.07760 |
| 3 | UAC | Tyrosine | Y | 163.06333 | 365.14093 |
| 4 | GGU | Glycine | G | 57.02146 | 422.16239 |
| 5 | UUU | Phenylalanine | F | 147.06841 | 569.23080 |
| 6 | UAA | Stop, no residue added | - | 0 | 587.24 with the retained water |
Why the phrase is loose, and why the degeneracy matters
Calling a codon a peptide codon compresses two different levels into one word. The code is a lookup table from triplets to residues; the peptide is the output of reading many entries in that table in order. The distinction has a measurable consequence. Reverse translation is degenerate: the pentapeptide above has one codon for methionine, four for alanine, two for tyrosine, four for glycine and two for phenylalanine, so 1 times 4 times 2 times 4 times 2 gives 64 different messenger sequences that all encode the same five residues. Reversing the table is therefore not a lookup, and any claim to have found the sequence for a peptide is a claim about one of many possibilities.
That degeneracy is not noise; it is used. Codon choice affects translation efficiency and accuracy because transfer RNA abundances differ between organisms, so the same protein coding sequence is routinely rewritten for a different expression host while the residue sequence is held constant. It also means that a mutation in the third position is often silent at the residue level, which is why sequence comparisons distinguish synonymous from nonsynonymous change. Finally, the codon account covers only ribosomal synthesis: short peptides made by non-ribosomal synthetases, and peptides produced by hydrolysis of a larger protein, are assembled without any codon at all. For a labeled example of where short peptides in commerce actually come from, see how peptide content is described on supplement labels.
Frequently asked questions
Does one codon make a peptide?
No. One codon specifies one residue, or signals termination. A peptide of n residues requires n sense codons read in one frame, plus a stop. The phrase peptide codon compresses two levels into one word, and keeping them apart is what makes the code table usable.
How many codons are there and how many are stops?
Sixty-four triplets in total, sixty-one of which specify residues and three of which are termination signals: UAA, UAG and UGA. AUG specifies methionine and also serves as the usual start. Two context-dependent exceptions allow UGA and UAG to specify additional residues in some organisms.
Why can different DNA sequences give the same peptide?
Because the code is degenerate: most residues are specified by two, four or six codons that usually differ at the third position. The pentapeptide Met Ala Tyr Gly Phe, for example, can be written 64 different ways. That freedom is what codon optimization exploits when a sequence is rewritten for a new host.
Related reading
tRNA in Polypeptide Synthesis: Translation, Step by Step
How transfer RNA works in translation: aminoacyl-tRNA synthetase charging, codon recognition, the A, P and E sites, pept
Peptides Also Known As: The Synonym Problem
Why one peptide carries a trade name, a research code, a sequence abbreviation, a CAS number and an INN, and how to reco
Peptides in Body Building: Verification Programs and Label Categories
How peptides appear in bodybuilding products, what NSF Certified for Sport, Informed Choice and Informed Sport verify, a
Sources & further reading
- NCBI Bookshelf — https://www.ncbi.nlm.nih.gov/books/
- PDB-101: educational resources on molecular structure — https://pdb101.rcsb.org/
- Nature Scitable — https://www.nature.com/scitable/
This page is part of the Peptide Structure, Classification & Scientific Terminology guide.
Questions about method, arithmetic or sourcing on this page? Message the editorial desk.
Message us