Peptides Also Known As: The Synonym Problem
One molecule can carry half a dozen names at the same time, and none of them is a mistake. A trade name is chosen by the company that sells the material. A research code is assigned during development and often stays in the literature long after a trade name appears. A sequence abbreviation describes the chain itself. A registry number identifies a structure uniquely. An international nonproprietary name is issued for use in prescribing and regulation. A catalog code belongs to one supplier. When a reader meets two names and wants to know whether they describe the same substance, the question is not which name is right but which kind of identifier is being used and who issued it.
That distinction is practical rather than academic. A trade name can be reused in another market for a different formula. A catalog code can be retired when a product line ends. A registry number is stable but silent about modifications and salt form. A sequence is precise about residues and silent about the counter-ion. Reconciling two names therefore means comparing the most structural identifier available first, then checking everything else against it. The vocabulary used on this site is defined in peptide structure and classification terminology, and the labelling category that usually accompanies a research listing is explained in what research use only means on a label.
The kinds of name a peptide can carry
Trade names are marketing property. They are capitalized by their owners, lowercase in our house style, and they carry no chemical information at all; the same trade name in two countries can legally refer to two different formulas. Research or development codes are the opposite in one respect: they are stable enough to persist in the literature for decades, and they usually encode the sponsor in a letter prefix followed by a number. Catalog codes sit below both, identifying an item as one supplier lists it, which means the same chain can have a dozen catalog codes across a dozen sellers while being one molecule.
International nonproprietary names are different in kind because they are issued by a body with a procedure. The World Health Organization assigns an INN per substance, published in lists, chosen to be distinctive and to signal the class through a shared stem. That stem is why so many generic names end in the same syllables: -tide marks peptides and glycopeptides, -relin marks peptides that stimulate hormone release, -morelin marks growth hormone secretagogues, -mab marks monoclonal antibodies, and -tinib marks kinase inhibitors. The shared ending lets a reader infer a family, and it lets the issuing body test new names against existing ones to reduce confusion between similar-sounding medicines.
Registry and database identifiers are the practical keys. A CAS registry number is assigned to a substance by Chemical Abstracts Service and is stable, though different salts, stereoisomers and hydrates receive different numbers, so the number identifies a specific form rather than a family. PubChem assigns its own compound identifier, and UniProt assigns an accession to a protein sequence, with entry versions changing as annotation grows while the accession stays put. These are the identifiers to reach for when two names have to be reconciled, because they are issued by organizations that maintain them.
| Identifier type | Example format | Who issues it | How stable |
|---|---|---|---|
| International nonproprietary name | semaglutide, tesamorelin | World Health Organization | Stable, one per substance; the stem signals the class |
| CAS registry number | Digits, hyphen, two digits, hyphen, check digit | Chemical Abstracts Service | Stable per form; salts and stereoisomers get separate numbers |
| PubChem compound identifier | 4660557 | PubChem at NCBI | Stable per structure record; some records are merged over time |
| UniProt accession | P01308 | UniProt Consortium | Stable per sequence; the entry version changes as annotation grows |
| Sequence notation | GIVEQCCTSICSLYQLENYCN | Authors, journals and suppliers | Precise about residues; silent on counter-ion and purity |
| Research development code | LY3298176 | The sponsoring organization | Persists in the literature; not a naming standard |
| Supplier catalog code | Letters and digits, supplier specific | The selling company | Stable only while the item is listed |
Writing a sequence so that it identifies one molecule
A sequence is the most informative identifier available for a chain, provided it is written completely. The one-letter code is standard, read from the amino terminus to the carboxyl terminus: the human insulin A chain is GIVEQCCTSICSLYQLENYCN and the B chain is FVNQHLCGSHLVEALYLVCGERGFFYTPKA, 21 and 30 residues respectively. The three-letter form is used where clarity about a modified residue matters. What turns a string of letters into a full description is the list of modifications: N-terminal acetylation, C-terminal amidation, the disulfide connectivities written as residue pairs, any non-standard residues such as norleucine or aminoisobutyric acid, and any label or linker such as biotin or a fluorophore.
A mass check is then the fastest way to catch a mismatch. With an average residue mass of about 110 Da and 18.01056 Da of water lost per bond formed, the mass of a chain can be estimated by hand before it is compared with a stated value, and each disulfide subtracts about 2.016 Da because two hydrogen atoms are lost on oxidation. Where an estimated mass and a stated mass disagree, the usual explanations are a counter-ion, a missing terminal amide, an undeclared modification, or simply a wrong sequence. Chirality is part of the same check: residues are L unless a D form is written explicitly, and glycine is the one achiral common residue.
- Confirm the direction: sequences are written amino terminus to carboxyl terminus unless stated otherwise.
- List every modification, including terminal acetylation or amidation and each disulfide as a residue pair.
- Name non-standard residues explicitly, and mark D residues where they occur.
- Recompute the expected mass from residue masses, subtract water per bond and 2.016 Da per disulfide.
- Compare the estimate with the stated mass, and treat an unexplained gap as a question rather than a rounding error.
Reconciling two names for one molecule
Work from the most structural identifier toward the least. Take the sequence from each source and align them; if they match, compare the modification lists, then the salt or counter-ion form, then the stereochemistry. Differences at those levels are real differences between products even when the chain is identical: an acetate salt and a trifluoroacetate salt of the same chain are not the same article, and a free acid is not the same article as the amide. Only then confirm with the registry data, checking whether the CAS number, the PubChem record and, for proteins, the UniProt accession all point at the structure the sequences describe.
The residual cases are where the effort is repaid. Some names denote a family or a mixture rather than a molecule, and no amount of cross-referencing will reconcile them with a single sequence: hydrolyzed collagen peptides is a distribution of fragments defined by a process, not one structure, which is why the declared composition matters so much in that category, as set out in how hydrolyzed collagen is described. Other names are repackaging, where a supplier applies its own word to a known compound. If no source states the sequence, the honest conclusion is that identity is unestablished, and a batch-specific test is the only thing that settles it, which is what independent testing against a certificate of analysis is for.
Frequently asked questions
Why do so many generic drug names end in the same few letters?
Because the ending is a stem chosen to signal the class. The World Health Organization assigns international nonproprietary names using published stems, so -tide marks peptides and glycopeptides, -mab marks monoclonal antibodies and -tinib marks kinase inhibitors. Shared endings help prescribers recognize a family and let the issuing body screen new names against existing ones to reduce confusion.
Is a CAS number enough to identify a peptide?
It is the most reliable single key, but not a complete description. A CAS number identifies a specific form, so different salts, stereoisomers and hydrates of the same chain carry different numbers. Use it to confirm the structure, then check the sequence, the modification list and the counter-ion form, since those determine what the material actually is in a vial.
How do you know whether two product names are the same molecule?
Compare sequences first, then modifications, then salt form, then confirm with a registry identifier such as a CAS number, a PubChem record or a UniProt accession. If no source publishes a sequence, the identity cannot be established from the names alone, and a batch-specific analytical test is needed. Where the name describes a mixture rather than one chain, the comparison cannot be made at all.
Related reading
RUO Peptide Meaning: What Research Use Only Actually Says
What research use only and not for human consumption mean, what they do not do, and how RUO differs from IVD, GMP and co
Glutamina Peptide: Glutamine Chemistry Inside a Chain
No molecule called glutamina peptide is registered, so this page covers glutamine chemistry: the side-chain amide, deami
Hydrolyzed Collagen Peptides: What Top Rated Labels Can Be Checked Against
How to read a hydrolyzed collagen peptides label: type, source, declared content, amino-acid profile and testing marks.
Sources & further reading
- World Health Organization: International Nonproprietary Names — https://www.who.int/medicines/services/inn/en/
- CAS Common Chemistry — https://commonchemistry.cas.org/
- UniProt entry P01308, human insulin — https://www.uniprot.org/uniprotkb/P01308
This page is part of the Peptide Structure, Classification & Scientific Terminology guide.
Questions about method, arithmetic or sourcing on this page? Message the editorial desk.
Message us