Explainer

DNA Replication: The Chemistry

Biochemistry & the Chemistry of LifeAdvanced7 min read
On this page
  1. The principle: semiconservative replication
  2. Step 1: opening the helix
  3. Step 2: the key reaction, making a phosphodiester bond
  4. Why DNA is only built 5′→3′
  5. Leading and lagging strands
  6. Accuracy: three layers of checking
  7. Replication in the test tube: PCR
  8. The end problem and telomeres
  9. Key takeaways

Before a human cell divides, it copies all of its DNA, roughly 6 billion base pairs across two sets of chromosomes, in a matter of hours, with an error rate of around one in a billion. Biology textbooks describe the enzymes involved; this article looks at the chemistry: which bonds form and break, where the energy comes from, and why the molecular machinery has the peculiar features it does.

The principle: semiconservative replication

Watson and Crick noted that the double helix suggested a copying mechanism: separate the two strands, and use each as a template for a new complementary strand, guided by base pairing. Each new double helix would then contain one old strand and one new strand. This is semiconservative replication.

In 1958, Matthew Meselson and Franklin Stahl tested it with a beautiful experiment:

  1. They grew bacteria for many generations on a medium containing heavy nitrogen, ¹⁵N, so all their DNA was “heavy”.
  2. They moved the bacteria to a medium with normal ¹⁴N, and sampled DNA after each generation.
  3. They separated DNA by density using density-gradient centrifugation in caesium chloride solution.

Results:

  • After one generation, all DNA had a single intermediate density: each helix had one heavy and one light strand.
  • After two generations, there were two bands: half intermediate, half light.

This is exactly what semiconservative replication predicts, and it ruled out the alternatives (conservative replication, where the original helix stays intact, and dispersive replication, where old and new pieces mix along each strand). It’s a classic example of using isotopes as tracers.

Step 1: opening the helix

The two strands are held together by hydrogen bonds between base pairs and by base stacking. At body temperature, they don’t separate by themselves.

  • Helicase enzymes use energy from ATP hydrolysis to push apart the strands, moving along the DNA and breaking hydrogen bonds between base pairs. This creates a Y-shaped replication fork.
  • Single-strand binding proteins coat the exposed strands to stop them re-pairing or folding back on themselves.
  • Topoisomerases relieve the twisting strain that builds up ahead of the fork. Unwinding a right-handed helix over-winds the DNA ahead, like pulling apart the strands of a rope. Topoisomerases cut one or both strands, let them rotate, and rejoin them. (Some antibiotics, such as ciprofloxacin, work by blocking bacterial topoisomerases.)

Step 2: the key reaction, making a phosphodiester bond

The enzyme DNA polymerase builds the new strand. Its building blocks are deoxynucleoside triphosphates (dNTPs): dATP, dTTP, dGTP and dCTP. Each has three phosphate groups attached in a chain to carbon 5′ of the sugar, labelled α (nearest the sugar), β and γ.

The reaction:

  1. The polymerase holds the template strand and the growing new strand, with its free 3′-OH group at the end.
  2. An incoming dNTP pairs with the next template base. The polymerase’s active site checks that the pair has the correct shape.
  3. The 3′-OH group of the growing strand acts as a nucleophile, attacking the α-phosphorus of the incoming dNTP.
  4. A new phosphodiester bond forms, and the β and γ phosphates leave together as pyrophosphate (PPᵢ, P₂O₇⁴⁻).
  5. Two Mg²⁺ ions in the active site help: one activates the 3′-OH, and both stabilise the negative charges that build up on the phosphates during the reaction (see cofactors and coenzymes).

(new strand)–3′-OH + dNTP → (new strand, one longer)–3′-OH + PPᵢ

Where the energy comes from

Forming the phosphodiester bond by itself is only slightly favourable. But pyrophosphate is immediately hydrolysed by the enzyme pyrophosphatase into two phosphate ions:

PPᵢ + H₂O → 2Pᵢ

This second reaction releases a lot of free energy, and because it removes a product, it pulls the polymerisation forward, making it effectively irreversible (the same logic as Le Chatelier’s principle). So each nucleotide added costs the equivalent of two high-energy phosphate bonds.

Why DNA is only built 5′→3′

Because the reaction needs a free 3′-OH on the growing strand to attack the incoming nucleotide, DNA polymerase can only add nucleotides to the 3′ end. New strands therefore always grow in the 5′→3′ direction.

There’s a chemical logic to this. The energy for each bond is carried by the incoming nucleotide’s triphosphate. If a mistake is made and the last nucleotide is removed by proofreading, the growing strand still ends in a plain 3′-OH, ready to try again. If synthesis went the other way, the high-energy triphosphate would have to sit on the end of the growing chain, and removing a mistake would remove the energy needed for the next step.

Polymerases also can’t start a strand from scratch: they can only extend an existing chain. So an enzyme called primase first lays down a short RNA primer (about 10 nucleotides), which provides the first 3′-OH.

Leading and lagging strands

The two template strands are antiparallel, but polymerase only works 5′→3′. At a replication fork, this creates an asymmetry:

  • The leading strand is built continuously in the same direction the fork is moving. One primer is enough.
  • The lagging strand has to be built in the opposite direction to fork movement. It’s made in short pieces, called Okazaki fragments (about 1,000–2,000 nucleotides in bacteria and 100–200 in human cells), each started with its own RNA primer.

Later:

  • the RNA primers are removed and replaced with DNA;
  • DNA ligase joins the fragments by forming the final phosphodiester bond. Ligase needs its own energy source (ATP in humans), because there’s no triphosphate at the gap to supply it.

Accuracy: three layers of checking

  1. Base selection. The polymerase’s active site accepts a new nucleotide only if the base pair has the right geometry. This alone gives an error rate of about 1 in 10⁴ to 10⁵.
  2. Proofreading. Many polymerases have a second active site that removes a wrongly paired nucleotide from the 3′ end by hydrolysing its phosphodiester bond (a 3′→5′ exonuclease activity). This improves accuracy by about 100 to 1,000 times.
  3. Mismatch repair. After replication, other enzymes scan the new DNA for mismatches, identify the new strand and replace the wrong section. Final error rates are around 1 in 10⁹ to 10¹⁰.

Replication in the test tube: PCR

The polymerase chain reaction copies DNA in the lab using the same chemistry, driven by temperature instead of helicase:

  1. Denature at about 95 °C: the strands separate.
  2. Anneal at about 50–65 °C: short DNA primers pair with the target sequence.
  3. Extend at about 72 °C: a heat-stable DNA polymerase (such as Taq, from a hot-spring bacterium) adds dNTPs to the primers.

Each cycle doubles the target DNA, so 30 cycles can produce about a billion copies. PCR is used in diagnosis, forensic science, research and paternity testing (see forensic chemistry). The need for a heat-stable enzyme is a nice application of how temperature affects enzymes.

The end problem and telomeres

On linear chromosomes, the very end of the lagging strand can’t be completed once the last RNA primer is removed, so chromosomes would shorten slightly with every round of replication. Chromosome ends are protected by telomeres, repeated sequences (TTAGGG in humans) that act as buffers. The enzyme telomerase carries its own RNA template to extend them, and it’s active in stem cells and many cancer cells.

Key takeaways

  • DNA replication is semiconservative, as shown by Meselson and Stahl using ¹⁵N and ¹⁴N.
  • Helicase separates the strands using ATP; topoisomerases relieve twisting.
  • DNA polymerase forms phosphodiester bonds when the 3′-OH of the growing strand attacks the α-phosphate of an incoming dNTP, releasing pyrophosphate, whose hydrolysis drives the reaction.
  • Synthesis runs only 5′→3′, needs an RNA primer, and is continuous on the leading strand but in Okazaki fragments on the lagging strand, joined by ligase.
  • Base selection, proofreading and mismatch repair give an error rate near 1 in a billion. For the structure being copied, see the structure of DNA.

Advertisement

More from this topic: Biochemistry & the Chemistry of Life