Explainer

Transcription and Translation from a Chemist's View

Biochemistry & the Chemistry of LifeAdvanced7 min read
On this page
  1. Part 1: Transcription
  2. Part 2: The genetic code
  3. Part 3: Charging transfer RNA
  4. Part 4: The ribosome and the peptide bond
  5. Part 5: Termination and folding
  6. What does a protein cost?
  7. Key takeaways

The instructions for building every protein in your body are written in DNA, a four-letter code of bases. Proteins, however, are made of twenty amino acids linked by peptide bonds. Getting from one to the other takes two steps: transcription, which copies a gene into RNA, and translation, which reads the RNA and builds a protein. Biology courses usually describe these as information flow. This article looks at the chemistry: which bonds form, how the right amino acid gets attached to the right RNA, and what it all costs in energy.

Part 1: Transcription

The template and the product

A gene is a stretch of double-stranded DNA. Only one strand, the template strand (also called the antisense strand), is read. The RNA made is complementary to the template and so has the same sequence as the other strand, the coding strand, except with uracil (U) in place of thymine (T) (see DNA vs RNA).

Example:

  • Coding strand: 5′-ATG GCC TTA-3′
  • Template strand: 3′-TAC CGG AAT-5′
  • mRNA: 5′-AUG GCC UUA-3′

RNA polymerase

The enzyme RNA polymerase:

  1. binds to a promoter, a DNA sequence just before the gene;
  2. separates a short region of the double helix (a “transcription bubble” about 14 base pairs long);
  3. reads the template 3′→5′ and builds RNA 5′→3′, pairing each ribonucleotide with the template base;
  4. releases the RNA when it reaches a termination signal.

The chemistry of joining nucleotides is the same as in DNA replication: the 3′-OH of the growing RNA attacks the α-phosphate of an incoming ribonucleoside triphosphate (ATP, UTP, GTP or CTP), forming a phosphodiester bond and releasing pyrophosphate, whose hydrolysis drives the reaction forward. Unlike DNA polymerase, RNA polymerase doesn’t need a primer; it can start a new chain from scratch. It also proofreads less, which is acceptable because RNA molecules are temporary copies.

Processing mRNA (in eukaryotes)

In human cells, the first RNA copy is modified before it leaves the nucleus:

  • A 5′ cap (a modified guanine nucleotide, attached by an unusual 5′-to-5′ triphosphate link) protects the RNA and helps ribosomes find it.
  • A poly-A tail of about 200 adenine nucleotides is added to the 3′ end, protecting it from degradation.
  • Splicing removes non-coding sections (introns) and joins the coding sections (exons). The chemistry is a pair of transesterification reactions: an –OH group attacks a phosphodiester bond, swapping one ester link for another, so no energy source is needed for the bond exchange itself. The machinery that catalyses it, the spliceosome, uses RNA at its catalytic core.

Part 2: The genetic code

The mRNA is read in groups of three bases called codons. With four possible bases in each of three positions, there are 4³ = 64 codons:

  • 61 code for amino acids;
  • 3 are stop codons (UAA, UAG, UGA);
  • AUG codes for methionine and also serves as the usual start codon.

Because 61 codons code for only 20 amino acids, most amino acids have several codons. The code is degenerate, but not ambiguous: each codon means exactly one thing. Different codons for the same amino acid often differ only in the third base, which is why “wobble” base pairing at that position (see base pairing) lets one tRNA read several codons.

The code is almost universal, shared by bacteria, plants and animals, which is strong evidence that all life shares a common origin.

Part 3: Charging transfer RNA

Here’s the chemical heart of translation. Nothing in an amino acid “recognises” a codon. Instead, the link between code and amino acid is made by transfer RNA (tRNA) and a family of enzymes that attach the right amino acid to the right tRNA.

Each tRNA is a folded RNA about 76 nucleotides long, with:

  • an anticodon loop, three bases that pair with the codon on mRNA;
  • a 3′ end (always ending in the sequence CCA) where an amino acid is attached.

The enzymes aminoacyl-tRNA synthetases (one for each amino acid) “charge” tRNAs in two steps:

  1. Activation: the amino acid’s carboxyl group attacks ATP, forming an aminoacyl-adenylate (the amino acid linked to AMP by a high-energy mixed anhydride bond) and releasing pyrophosphate: amino acid + ATP → aminoacyl-AMP + PPᵢ
  2. Transfer: the amino acid is moved onto the 3′ end of the correct tRNA, forming an ester bond: aminoacyl-AMP + tRNA → aminoacyl-tRNA + AMP

The ester bond between the amino acid and tRNA is a high-energy bond. It stores the energy that will later form the peptide bond.

Accuracy here is crucial: if the synthetase attaches the wrong amino acid, the ribosome has no way to notice, because it only checks codon–anticodon pairing. Many synthetases therefore have editing sites that hydrolyse wrongly attached amino acids, giving error rates of roughly 1 in 10,000 or better.

Part 4: The ribosome and the peptide bond

The ribosome is a huge complex of RNA and proteins, with two subunits. It has three sites for tRNAs, called A (aminoacyl), P (peptidyl) and E (exit).

The elongation cycle:

  1. Decoding: a charged tRNA whose anticodon pairs with the codon in the A site is delivered, with the help of a protein factor that hydrolyses GTP. The ribosome checks the geometry of the codon–anticodon pair to reject wrong tRNAs.
  2. Peptide bond formation: the amino group of the amino acid on the A-site tRNA acts as a nucleophile and attacks the ester carbonyl linking the growing chain to the P-site tRNA. The whole chain is transferred onto the new amino acid, forming a new peptide bond (see the peptide bond). The P-site tRNA is left empty.
  3. Translocation: the ribosome moves along the mRNA by one codon, powered by another GTP hydrolysis. The empty tRNA moves to the E site and leaves, the tRNA with the chain moves to the P site, and the A site is free for the next codon.

The chain therefore grows from its N-terminus to its C-terminus, and the energy for each peptide bond comes from the high-energy ester bond made during tRNA charging.

The ribosome is a ribozyme

In 2000, high-resolution structures of the ribosome revealed that the site where peptide bonds form, the peptidyl transferase centre, is made entirely of ribosomal RNA. No protein side chain is close enough to take part in the chemistry. The ribosome catalyses the reaction mainly by holding the two substrates in exactly the right position and orientation, with help from an RNA hydroxyl group. This makes it a ribozyme and is strong evidence for the idea that RNA once did life’s catalysis (the “RNA world”). Venki Ramakrishnan, Thomas Steitz and Ada Yonath shared the 2009 Nobel Prize in Chemistry for ribosome structures.

Part 5: Termination and folding

When a stop codon reaches the A site, a release factor protein binds instead of a tRNA. It positions a water molecule to hydrolyse the ester bond between the finished chain and the last tRNA, releasing the protein. The chain then folds into its working shape, often with help from chaperones (see protein folding).

What does a protein cost?

Adding up the high-energy phosphate bonds used per amino acid:

  • 2 in charging the tRNA (ATP → AMP + PPᵢ, with PPᵢ then hydrolysed);
  • 1 GTP in delivering the tRNA to the A site;
  • 1 GTP in translocation.

That’s about 4 high-energy phosphate bonds per peptide bond, not counting transcription, proofreading or folding. A 300-amino-acid protein therefore costs at least 1,200 such bonds, which is why protein synthesis is one of the biggest energy costs in a growing cell (see ATP, the cell’s energy currency).

Key takeaways

  • Transcription: RNA polymerase reads the template strand 3′→5′ and builds RNA 5′→3′, using ribonucleoside triphosphates and releasing pyrophosphate; eukaryotic mRNA is capped, tailed and spliced.
  • The genetic code uses 64 codons: 61 for amino acids, 3 stops, with AUG as the start.
  • Aminoacyl-tRNA synthetases use ATP to attach each amino acid to its tRNA through a high-energy ester; this step, not the ribosome, links code to amino acid.
  • The ribosome, a ribozyme, forms each peptide bond by transferring the growing chain onto the incoming amino acid’s amino group.
  • Each peptide bond costs about four high-energy phosphate bonds. For the molecules involved, see nucleic acids.

Advertisement

More from this topic: Biochemistry & the Chemistry of Life