Explainer

Base Pairing: A–T and G–C

Biochemistry & the Chemistry of LifeIntermediate6 min read
On this page
  1. The four bases
  2. Hydrogen-bond donors and acceptors
  3. Why a purine always pairs with a pyrimidine
  4. Chargaff’s rules
  5. GC content and melting temperature
  6. Base pairing in RNA
  7. Base pairing in action
  8. When pairing goes wrong
  9. A worked example: complementary sequences
  10. Key takeaways

Complementary base pairing is the rule that makes heredity work: adenine pairs with thymine (or uracil in RNA), and guanine pairs with cytosine. It’s why DNA can be copied, why genes can be read into RNA, and why a PCR test or DNA fingerprint works. But why these pairs and not others? The answer is a precise match of shapes and hydrogen bonds, and it’s a lovely piece of chemistry.

The four bases

DNA uses four bases, in two families (see nucleic acids):

  • Purines (two fused rings): adenine (A) and guanine (G).
  • Pyrimidines (one ring): cytosine (C) and thymine (T); RNA uses uracil (U) instead of thymine.

The bases are flat, and along one edge each carries a specific pattern of atoms that can take part in hydrogen bonds.

Hydrogen-bond donors and acceptors

A hydrogen bond forms between:

  • a donor: a hydrogen atom attached to an electronegative atom (here, an N–H group), and
  • an acceptor: an electronegative atom with a lone pair (here, a C=O oxygen or a ring nitrogen).

Along the pairing edge, each base presents a pattern of donors (D) and acceptors (A). A base can only pair well with a partner whose pattern is the mirror image: donor facing acceptor, and acceptor facing donor.

A–T: two hydrogen bonds

  • Adenine offers an N–H donor (from its amino group) and a ring-nitrogen acceptor.
  • Thymine offers a C=O acceptor opposite adenine’s N–H, and an N–H donor opposite adenine’s ring nitrogen.
  • Result: two hydrogen bonds.

G–C: three hydrogen bonds

  • Guanine offers a C=O acceptor and two N–H donors along its edge.
  • Cytosine offers an N–H donor facing guanine’s C=O, and a ring nitrogen and a C=O acceptor facing guanine’s two donors.
  • Result: three hydrogen bonds.

Why not A–C or G–T?

Try pairing adenine with cytosine, and donors end up facing donors or acceptors facing acceptors. Those positions repel instead of attracting, so the pair doesn’t form properly. The same happens for G with T (in its normal form). The donor–acceptor patterns act like a chemical lock and key.

Why a purine always pairs with a pyrimidine

There’s a second, geometric reason for the pairing rules. In the double helix, the two sugar–phosphate backbones are a fixed distance apart. Every base pair has to span that gap:

  • purine + pyrimidine (three rings in total) fits exactly;
  • purine + purine (four rings) would be too wide;
  • pyrimidine + pyrimidine (two rings) would be too narrow to reach.

A–T and G–C pairs are almost exactly the same overall size and shape, with the sugar attachment points the same distance apart. That’s what allows any sequence of pairs to fit into a smooth, regular helix, a key insight Watson and Crick reached when building their model (see the structure of DNA).

Chargaff’s rules

Before the double helix was known, the biochemist Erwin Chargaff measured base compositions of DNA from many species and found:

  • the amount of A equals the amount of T;
  • the amount of G equals the amount of C;
  • so purines (A + G) equal pyrimidines (C + T);
  • but the proportion of A+T versus G+C varies between species.

Base pairing explains all of this: every A on one strand is matched by a T on the other, and every G by a C.

Worked example: if 30% of the bases in a double-stranded DNA sample are adenine, what are the percentages of the others?

  • T = A = 30%
  • A + T = 60%, so G + C = 40%
  • G = C = 20%

GC content and melting temperature

Separating the two strands of DNA by heating is called melting (or denaturation). The temperature at which half the DNA has separated is the melting temperature, T_m.

DNA with more G–C pairs has a higher T_m. Part of the reason is the extra hydrogen bond in each G–C pair. The larger part is base stacking: G–C pairs stack on neighbouring pairs more strongly than A–T pairs, through dispersion forces and the hydrophobic effect (see intermolecular forces).

This matters in practice:

  • Organisms living in hot environments often have GC-rich DNA or other stabilising adaptations.
  • In PCR, primers are designed with a suitable GC content so they bind at the right temperature.
  • Regions of DNA rich in A–T pairs, which separate more easily, are often found where strand separation starts, such as replication origins.

Base pairing in RNA

RNA follows the same rules, with uracil taking the place of thymine: A–U (two hydrogen bonds) and G–C (three).

Because RNA is single-stranded, it pairs with itself, forming hairpins and loops. It also often uses a G–U “wobble” pair, held by two hydrogen bonds, which is slightly weaker and a different shape. Wobble pairing is important in translation: it lets one transfer RNA recognise more than one codon, which is part of why the genetic code has “synonymous” codons for the same amino acid (see transcription and translation).

Base pairing in action

  • DNA replication: each strand is a template; the enzyme DNA polymerase adds only the nucleotide that pairs correctly with the template base (see DNA replication: the chemistry).
  • Transcription: RNA polymerase builds an RNA copy complementary to one DNA strand.
  • Translation: the anticodon of each tRNA base-pairs with a codon on mRNA.
  • Laboratory methods: PCR primers, DNA probes, DNA microarrays, gene editing guide RNAs and COVID-19 PCR tests all rely on complementary base pairing to find specific sequences.

When pairing goes wrong

Occasionally a base temporarily shifts into a rarer form (a tautomer), in which a hydrogen atom has moved from one position to another. Its donor–acceptor pattern changes, and it may pair with the wrong partner, for example a rare form of T pairing with G. If this happens during replication and isn’t corrected, the result is a point mutation.

Chemical damage can do the same:

  • deamination of cytosine turns it into uracil, which pairs with A instead of G;
  • some chemicals add groups to bases, changing their pairing.

Cells have several layers of defence. DNA polymerase checks each new pair and removes mismatches (proofreading), and separate repair systems scan the DNA afterwards. The final error rate in copying human DNA is roughly one mistake per billion bases or better.

A worked example: complementary sequences

Suppose one strand of a DNA double helix reads 5′-ATGCGTAC-3′. To write the other strand, pair each base (A with T, G with C) and reverse the direction, because the strands are antiparallel. The partner strand is 3′-TACGCATG-5′, which is conventionally rewritten 5′→3′ as 5′-GTACGCAT-3′. The RNA transcribed using the first strand as the template would be complementary to it, with U in place of T: 5′-GUACGCAU-3′. Practising this reversal carefully avoids one of the most common mistakes in genetics exam questions.

Key takeaways

  • A pairs with T (U in RNA) through two hydrogen bonds; G pairs with C through three.
  • Pairing depends on matching hydrogen-bond donors and acceptors and on purine–pyrimidine geometry that keeps the helix a constant width.
  • Chargaff’s rules (A = T, G = C) follow directly from pairing.
  • GC-rich DNA has a higher melting temperature, mainly because of stronger base stacking.
  • Base pairing drives replication, transcription, translation and lab techniques; rare mispairs cause mutations. For the bigger picture, see DNA vs RNA.

Advertisement

More from this topic: Biochemistry & the Chemistry of Life