Explainer

The Genetic Code: Chemistry of Information

Biochemistry & the Chemistry of LifeIntermediate7 min read
On this page
  1. The alphabet: four bases
  2. The problem: 4 letters, 20 amino acids
  3. Cracking the code
  4. Reading the codon table
  5. The adapter: transfer RNA
  6. Wobble: fewer tRNAs than codons
  7. Translation in brief
  8. Mutations: when the text changes
  9. Is the code universal?
  10. Common misconceptions
  11. Key takeaways

DNA is a chemical, made of carbon, hydrogen, oxygen, nitrogen and phosphorus, yet it stores the information to build a human being. How can a molecule carry information? The answer is the genetic code: a set of rules that translates a sequence of chemical bases into a sequence of amino acids. It’s one of the great discoveries of the twentieth century, and it rests on simple chemistry — hydrogen bonding, molecular shape and a few clever adapter molecules.

The alphabet: four bases

DNA is a chain of nucleotides. Each nucleotide contains a phosphate, a sugar (deoxyribose) and one of four bases: adenine (A), guanine (G), cytosine (C) and thymine (T). RNA uses ribose and swaps thymine for uracil (U). See nucleic acids and DNA vs RNA.

The sugar–phosphate backbone is the same all along the chain; it’s the order of the bases that carries the information. It’s like a sentence written with a four-letter alphabet.

The information can be copied accurately because the bases pair by hydrogen bonding in a precise way: A with T (or U), with two hydrogen bonds, and G with C, with three (see base pairing). The shapes and positions of hydrogen-bond donors and acceptors mean that only these pairs fit neatly inside the double helix.

The problem: 4 letters, 20 amino acids

Proteins are built from 20 standard amino acids (see amino acids). How can four bases specify twenty different amino acids?

  • One base per amino acid: only 4 possibilities — not enough.
  • Two bases: 4 × 4 = 16 — still not enough.
  • Three bases: 4 × 4 × 4 = 64 — more than enough.

So the code must use triplets of bases. Each triplet of bases in messenger RNA is called a codon. Experiments by Francis Crick, Sydney Brenner and colleagues in 1961, using mutations that added or deleted bases, showed that the code is indeed read in groups of three, from a fixed starting point, without gaps or overlaps.

Cracking the code

The first codon was identified in 1961 by Marshall Nirenberg and Heinrich Matthaei. They made an artificial RNA containing only uracil (poly-U: UUUUUU…) and added it to a cell-free extract of bacteria that could make proteins. The extract produced a protein made only of phenylalanine. So UUU must code for phenylalanine.

Over the next few years, Nirenberg, Har Gobind Khorana (who synthesised RNAs with defined repeating sequences) and others worked out all 64 codons. Nirenberg, Khorana and Robert Holley (who determined the structure of a transfer RNA) shared the 1968 Nobel Prize in Physiology or Medicine.

Reading the codon table

The genetic code is usually shown as a table of mRNA codons. A few key features:

Feature Example
Start codon AUG — codes for methionine and marks where translation begins
Stop codons UAA, UAG, UGA — code for no amino acid; they end the protein
Most amino acids have several codons Leucine: UUA, UUG, CUU, CUC, CUA, CUG (6 codons)
Some have only one Methionine (AUG), tryptophan (UGG)

61 codons code for amino acids and 3 are stop signals. Because several codons can code for the same amino acid, the code is called degenerate (or redundant). It is not ambiguous: each codon codes for only one thing.

Look closely and a pattern appears: codons for the same amino acid usually differ only in the third base. GCU, GCC, GCA and GCG all code for alanine. This makes the code tolerant of errors: many mutations in the third position don’t change the amino acid at all.

There’s chemical logic in the table too. Codons with U in the second position mostly code for hydrophobic amino acids (phenylalanine, leucine, isoleucine, methionine, valine). A mistake in the first or third position often swaps one hydrophobic amino acid for another similar one, minimising damage to protein folding (see protein folding).

The adapter: transfer RNA

Codons don’t bind amino acids directly — there’s no chemical reason why UUU should attract phenylalanine. Instead, the link is made by adapter molecules called transfer RNAs (tRNAs), as Crick predicted in 1955.

Each tRNA is a short RNA chain (about 76 nucleotides) that folds into an L-shape. It has two important ends:

  1. An anticodon loop, with three bases that pair with a codon on mRNA by hydrogen bonding.
  2. An acceptor end, where the matching amino acid is attached by an ester bond.

The crucial step is attaching the correct amino acid to each tRNA. This is done by enzymes called aminoacyl-tRNA synthetases — typically one for each amino acid. Each synthetase recognises both its amino acid and the correct tRNAs, and uses ATP to attach them. These enzymes are the true “translators” of the genetic code; some even proofread, removing wrongly attached amino acids. For example, the enzyme for isoleucine checks for, and removes, valine, which differs by just one CH₂ group.

Wobble: fewer tRNAs than codons

There are 61 amino-acid codons, but cells don’t need 61 different tRNAs. Crick’s wobble hypothesis (1966) explained why: the pairing between the third base of the codon and the first base of the anticodon is less strict. For example, a G in the anticodon can pair with either C or U. Some tRNAs contain a modified base, inosine, which can pair with U, C or A. So one tRNA can read several codons that differ only in the third position.

Translation in brief

  1. Transcription: DNA is copied into mRNA in the nucleus.
  2. Initiation: the ribosome binds the mRNA and finds the AUG start codon, which sets the reading frame.
  3. Elongation: tRNAs bring amino acids one by one; the ribosome joins them with peptide bonds (see peptide bonds).
  4. Termination: at a stop codon, a release factor frees the finished protein.

For the full process, see protein synthesis.

Mutations: when the text changes

Because the code is read in triplets, different kinds of mutation have very different effects:

  • Silent: the codon changes but still codes for the same amino acid (e.g. GCU → GCC, both alanine).
  • Missense: one amino acid is replaced by another. In sickle cell disease, a single base change turns GAG (glutamic acid) into GUG (valine) in a haemoglobin chain, swapping a charged amino acid for a hydrophobic one, which makes the proteins stick together (see haemoglobin).
  • Nonsense: a codon becomes a stop codon, cutting the protein short.
  • Frameshift: inserting or deleting a number of bases that isn’t a multiple of three shifts the reading frame, scrambling every codon after it.

Is the code universal?

Almost. The same codon table is used by nearly all living things, from bacteria to humans — powerful evidence for a common ancestor. That’s why a human insulin gene can be inserted into bacteria and read correctly to make human insulin (see insulin). There are small variations: mitochondria use slightly different rules (in human mitochondria, UGA codes for tryptophan rather than stop), and a few microbes have reassigned codons. Some organisms even use extra amino acids: selenocysteine, which contains selenium, is inserted at certain UGA codons in special contexts.

Common misconceptions

  • “DNA codes for amino acids directly.” DNA is transcribed into mRNA; tRNAs and synthetases make the link to amino acids.
  • “Each amino acid has one codon.” Most have two to six.
  • “Stop codons code for a stop amino acid.” They code for none; a protein factor recognises them.
  • “All mutations change the protein.” Silent mutations don’t.

Key takeaways

  • Information is stored in the order of four bases; copying relies on hydrogen-bonded base pairing.
  • The code uses triplets (codons): 64 codons, 61 for amino acids and 3 stops; AUG starts.
  • The code is degenerate but unambiguous, and nearly universal.
  • tRNAs carry anticodons and amino acids; aminoacyl-tRNA synthetases do the actual translating.
  • Wobble lets one tRNA read several codons; mutations can be silent, missense, nonsense or frameshift.

To see how scientists now edit this code directly, read CRISPR: the chemistry of gene editing.

Advertisement

More from this topic: Biochemistry & the Chemistry of Life