On this page
In 2012, a paper by Jennifer Doudna, Emmanuelle Charpentier and their colleagues showed that a bacterial immune enzyme could be programmed to cut any chosen DNA sequence. Within a few years, CRISPR–Cas9 had become the standard tool for editing genes in laboratories worldwide, and in 2020 Doudna and Charpentier received the Nobel Prize in Chemistry. In late 2023, the first CRISPR-based therapy, for sickle cell disease and β-thalassaemia, was approved for patients. Behind the headlines is a beautiful piece of molecular chemistry: an enzyme guided by RNA base pairing, checking DNA letter by letter, then breaking two phosphodiester bonds with metal-ion catalysis.
Origins: a bacterial immune system
Bacteria are constantly attacked by viruses called bacteriophages. In the 1990s and 2000s, researchers noticed strange repeated sequences in bacterial genomes: short palindromic repeats, separated by unique “spacer” sequences. The name CRISPR stands for clustered regularly interspaced short palindromic repeats. The spacers turned out to match pieces of viral DNA.
The system works like a molecular memory:
- When a bacterium survives a viral infection, it cuts out a short piece of the virus’s DNA and stores it as a new spacer in its CRISPR array.
- The array is transcribed into RNA, which is processed into short CRISPR RNAs (crRNAs), each carrying one viral sequence.
- A crRNA teams up with a Cas (CRISPR-associated) protein. If the same virus returns, the crRNA recognises its DNA by base pairing, and the Cas protein cuts it, disabling the virus.
The Cas9 protein from Streptococcus pyogenes (SpCas9), a large protein of 1,368 amino acids, became the workhorse of gene editing.
The guide RNA: programming by base pairing
In bacteria, Cas9 needs two RNAs: the crRNA and a second one called tracrRNA. The 2012 paper showed they could be joined into a single guide RNA (sgRNA) of about 100 nucleotides. The guide has two parts:
- A 20-nucleotide spacer at one end, chosen by the scientist to match the target DNA.
- A scaffold that folds into a structure Cas9 binds tightly.
This is what makes CRISPR so powerful: to target a new gene, you don’t need to engineer a new protein, as with older tools. You just change 20 bases of RNA. Targeting relies on the same Watson–Crick base pairing that holds DNA together — A with T (U in RNA), G with C — via hydrogen bonds (see base pairing).
A 20-base sequence is specific enough to be almost unique in a genome. There are 4²⁰ (about 10¹²) possible 20-base sequences, while the human genome is about 3 × 10⁹ base pairs long.
Finding the target: the PAM
Cas9 can’t scan the genome by opening every DNA double helix and testing it against the guide — that would be far too slow. Instead, it first looks for a short sequence called the protospacer adjacent motif (PAM). For SpCas9 the PAM is 5′-NGG-3′ (any base followed by two guanines), which must sit right next to the target sequence.
Cas9 slides and hops along DNA, and amino acids in a PAM-interacting region form hydrogen bonds with the two G bases in the major groove of the helix. Only when it finds a PAM does it try to unwind the adjacent DNA. The PAM requirement also protects the bacterium: its own CRISPR array contains the spacer sequences but not the PAM, so Cas9 doesn’t cut the bacterium’s own genome.
Checking the match: the R-loop
Once bound at a PAM, Cas9 begins to unwind the double helix next to it. The guide RNA tries to pair with one DNA strand, base by base, starting from the end nearest the PAM. The structure formed — an RNA–DNA hybrid with the other DNA strand pushed aside — is called an R-loop.
The 8–12 bases closest to the PAM, called the seed region, are the most important. If they match, the R-loop extends along all 20 bases. Mismatches in the seed usually stop the process. Mismatches further away are tolerated more, which is one reason Cas9 sometimes cuts “off-target” sites with similar sequences.
The cut: two scissors, two strands
When the RNA–DNA hybrid is complete, Cas9 changes shape and activates two separate cutting sites (nuclease domains):
- The HNH domain cuts the DNA strand paired with the guide RNA.
- The RuvC domain cuts the other (displaced) strand.
Each cut breaks a phosphodiester bond in the DNA backbone by hydrolysis — water attacks a phosphorus atom, breaking the bond between phosphate and sugar. Both domains use magnesium ions (Mg²⁺), which position the water, activate it and stabilise the negative charge that builds up during the reaction (see cofactors and coenzymes). The two cuts happen at the same spot, about three base pairs upstream of the PAM, leaving a double-strand break with blunt ends.
The edit: the cell does the rest
Cas9 only cuts. The actual edit happens when the cell repairs the break. Cells have two main repair routes:
- Non-homologous end joining (NHEJ): the cell glues the broken ends back together. This is fast but error-prone, often adding or deleting a few bases. If this happens in a gene, it usually causes a frameshift — shifting how the code is read in triplets — and switches the gene off (see the genetic code). This is the easiest way to “knock out” a gene.
- Homology-directed repair (HDR): if the scientist supplies a DNA template matching the sequences on either side of the cut but containing a desired change, the cell can copy it into the genome. This allows precise corrections but is less efficient, especially in cells that aren’t dividing.
Beyond cutting: newer editors
Because double-strand breaks can cause unwanted deletions or rearrangements, researchers have engineered Cas9 into gentler tools:
- Dead Cas9 (dCas9): with both nuclease domains disabled by single amino acid changes, Cas9 still binds DNA but doesn’t cut. Attached to other proteins, it can switch genes on or off, or carry fluorescent tags to label locations.
- Base editors (developed in David Liu’s lab from 2016): a Cas9 that cuts only one strand (a “nickase”) is joined to an enzyme that chemically changes a base. Cytosine base editors use a deaminase to remove an amine group from cytosine, converting it to uracil, which the cell then copies as thymine — turning a C–G pair into T–A. Adenine base editors convert A to inosine, read as G, turning A–T into G–C. No double-strand break is needed.
- Prime editors (2019): a nickase Cas9 fused to a reverse transcriptase, with an extended guide RNA that carries both the targeting sequence and a template for the new sequence. It can write small insertions, deletions and all types of base swaps.
CRISPR in medicine
The first approved CRISPR therapy (exa-cel, marketed as Casgevy) treats sickle cell disease and β-thalassaemia. Doctors collect a patient’s blood stem cells, edit them outside the body to disrupt a genetic “switch” that normally turns off fetal haemoglobin production after birth, then return the cells. The reactivated fetal haemoglobin compensates for the faulty adult haemoglobin (see haemoglobin).
Other uses include research models of disease, crops with improved traits, and rapid diagnostic tests that use related Cas enzymes (Cas12 and Cas13) to detect viral genetic material.
Limitations and ethics
- Off-target cuts at similar sequences, reduced by engineered high-fidelity Cas9 variants and careful guide design.
- Delivery: getting a large protein and RNA into the right cells in the body is difficult.
- Mosaicism: not every cell is edited the same way.
- Unintended large changes from double-strand breaks.
- Germline editing — changing embryos so the edit is inherited — raises serious ethical concerns. In 2018 a researcher’s announcement of gene-edited babies was widely condemned, and he was later imprisoned in China.
Key takeaways
- CRISPR–Cas9 comes from a bacterial immune system that stores viral sequences.
- A 20-nucleotide guide RNA targets DNA by base pairing; Cas9 first finds a PAM (NGG).
- The guide forms an R-loop; the seed region must match.
- HNH and RuvC domains hydrolyse phosphodiester bonds using Mg²⁺, leaving a double-strand break.
- Edits come from cell repair (NHEJ or HDR); base and prime editors change DNA without full breaks.
For the molecule being edited, see DNA structure and how DNA replicates.
Advertisement