Chapter 17

Molecular Biology: Replication, Transcription & Translation in Depth

College

Chapter 13 gave the central dogma as a cartoon. Here it becomes mechanism: the specific enzymes, error-correcting proofreaders, and RNA-processing machinery that make information transfer accurate to roughly one error in ten billion bases.

At a glance
Core ideaNamed enzymes make information transfer accurate to ~1 error in 10¹⁰ bases.
Key termOkazaki fragments — the short discontinuous pieces of the lagging strand.
You can…Explain leading vs lagging synthesis, splicing and alternative splicing.
Watch outFidelity is layered — proofreading and mismatch repair multiply, not add.
1 Theory

DNA replication: the enzyme choreography

Replication begins at an origin of replication, where helicase unwinds the double helix, single-strand binding proteins keep the strands apart, and topoisomerase relieves the supercoiling that unwinding creates ahead of the fork. Because DNA polymerase can only extend a strand 5′→3′, and the two template strands are antiparallel, the two new strands are made differently:

  • Leading strand — synthesised continuously in the same direction the fork opens, needing only one RNA primer (laid by primase) to start.
  • Lagging strand — synthesised discontinuously, away from the fork, as a series of short Okazaki fragments, each needing its own RNA primer.

DNA polymerase III does the bulk of synthesis; DNA polymerase I removes the RNA primers and fills the gaps with DNA; DNA ligase seals the remaining nicks into a continuous strand. Polymerase III also proofreads using a built-in 3′→5′ exonuclease that detects and excises mismatched bases as it goes — bringing the raw error rate of ~1 in 10⁵ down to ~1 in 10⁷. A separate mismatch repair system scans the finished DNA afterward, cutting a further ~100-fold, for a final fidelity near one error per 10⁹–10¹⁰ bases. In eukaryotes, the very ends of linear chromosomes cannot be fully copied by this machinery (the "end-replication problem"); the enzyme telomerase extends repetitive telomere sequences to protect genes near the chromosome tips from being eroded.

Transcription: from gene to processed mRNA

In eukaryotes, RNA polymerase II does not bind DNA directly; general transcription factors assemble at the promoter (often containing a TATA box) to form a pre-initiation complex before polymerase can begin. The resulting pre-mRNA is processed before it leaves the nucleus: a 5′ cap (7-methylguanosine) is added for stability and ribosome recognition; a poly-A tail is added at the 3′ end; and the spliceosome — built from small nuclear ribonucleoproteins (snRNPs) — recognises intron boundaries and excises the non-coding introns via a lariat intermediate, splicing the coding exons together. Alternative splicing allows a single gene to be spliced in different combinations, producing several distinct proteins from one gene — a major source of proteome diversity beyond raw gene count.

Translation: reading the code at the ribosome

The ribosome has three tRNA-binding sites — A (aminoacyl, incoming), P (peptidyl, growing chain) and E (exit). Initiation: the small subunit, aided by initiation factors, finds the start codon (AUG) using an initiator tRNA; the large subunit then joins. Elongation: an aminoacyl-tRNA — charged with its correct amino acid by a specific aminoacyl-tRNA synthetase — is delivered to the A site (with elongation factor EF-Tu); the ribosome's rRNA itself catalyses peptide-bond formation (it is a ribozyme); the ribosome then translocates one codon along (EF-G), shifting tRNAs from A→P→E. Termination occurs when a stop codon is reached and a release factor triggers polypeptide release. The wobble hypothesis explains why the code tolerates variation at a codon's third base: the tRNA anticodon's third-position pairing is looser, so one tRNA can often read several synonymous codons.

2 Explanation

Why the lagging strand has to be discontinuous

This is not an evolutionary accident but a direct geometric consequence of two hard constraints: DNA polymerase can only add nucleotides to a free 3′-OH, extending strands 5′→3′; and the fork itself only opens in one direction. On the leading-strand template, the 5′→3′ direction of synthesis happens to match the direction the fork is opening, so polymerase can simply chase the fork continuously. On the lagging-strand template, matching 5′→3′ synthesis would mean building away from the fork — so the cell instead waits for a short stretch of exposed template to appear, primes it, synthesises an Okazaki fragment back toward the previous fragment, and repeats. It looks inefficient, but it is the only way to obey the polymerase's fixed chemistry while still replicating both strands at the same fork simultaneously.

Fidelity is layered, not single-shot

No single step in replication is perfect — a raw error rate of 1 in 10⁵ from base-pairing chemistry alone would be catastrophic over a 3-billion-base genome. Life instead stacks independent error-correcting layers (polymerase selectivity, proofreading exonuclease, mismatch repair) multiplicatively, the same way redundant safety checks in engineering compound to near-total reliability.

3 Practical

Worked example: timing a bacterial genome replication

E. coli has a single circular chromosome of about 4.6 million base pairs. Replication starts at one origin and proceeds via two replication forks moving in opposite directions around the circle, each synthesising at roughly 1000 nucleotides per second.

  1. Split the task between forks. Two forks share the genome, so each fork must copy 4.6 million ÷ 2 = 2.3 million bp.
  2. Apply the rate. time = distance ÷ rate = 2 300 000 bp ÷ 1000 bp/s.
  3. Compute. = 2300 seconds38.3 minutes for full replication.
  4. Sanity-check against real biology. Fast-growing E. coli can actually divide every ~20 minutes — faster than one replication round takes! This is possible because cells under rapid growth start a new round of replication before the previous one finishes, running multiple forks in parallel on the same chromosome.
  5. Draw the lesson. A simple rate calculation reveals a real constraint of cell biology, and the "overlapping replication" trick that bacteria evolved to grow faster than their own replication machinery would otherwise allow.
4 Q&A

Test yourself

Q1 Explain, mechanistically, why the lagging strand must be synthesised as Okazaki fragments rather than continuously.

DNA polymerase extends strands only 5′→3′, and the two template strands run antiparallel. As the replication fork opens in one fixed direction, the leading-strand template is exposed in the same direction synthesis proceeds, so one continuous polymerase can chase the fork. The lagging-strand template, however, is exposed in the opposite direction to the required 5′→3′ synthesis — so polymerase must instead wait for each new stretch of single-stranded template, prime it, and synthesise a short fragment back toward the previous fragment. Ligase then joins the fragments after primer removal, giving one continuous strand made of stitched-together pieces.

Q2 Replication's raw base-pairing fidelity is only about 1 error in 10⁵. Explain the two further steps that bring this down to ~1 in 10¹⁰, and why layering them multiplies rather than adds their benefit.

First, proofreading: DNA polymerase III's built-in 3′→5′ exonuclease detects a mismatched base immediately after insertion (a wrong base pairs loosely, stalling the enzyme) and excises it, roughly improving fidelity 100-fold (to ~1 in 10⁷). Second, mismatch repair scans the newly made DNA afterward, recognises any remaining distortions in the helix, and replaces the incorrect stretch, improving fidelity by roughly another 100–1000-fold. Because each layer independently catches a fraction of the errors that got past the previous layer, the surviving error rate is the product of each layer's failure rate, not their sum — which is why stacking several imperfect checks yields a dramatically more reliable overall system.

Q3 Contrast prokaryotic and eukaryotic transcription initiation, and explain why eukaryotic pre-mRNA needs further processing before translation.

In prokaryotes, a single RNA polymerase can bind the promoter directly (with the help of a sigma factor) and begin transcription immediately; there is no nuclear membrane, so translation of the mRNA can begin even before transcription finishes. In eukaryotes, RNA polymerase II cannot bind the promoter alone — a set of general transcription factors must first assemble a pre-initiation complex. Because the nucleus is a separate compartment, the pre-mRNA must also be processed (5′ cap, splicing out introns, 3′ poly-A tail) and exported through nuclear pores before ribosomes in the cytoplasm can translate it — an extra layer of regulation and quality control unavailable to prokaryotes.

Q4 What is the wobble hypothesis, and why is it advantageous for translation?

The wobble hypothesis states that base-pairing between a codon's third position and a tRNA anticodon's corresponding position is less strict than at the first two positions, so a single tRNA can often correctly pair with several synonymous codons that differ only in that third base. This is advantageous because it means a cell does not need a full 61 distinct tRNAs for the 61 sense codons — around 40–45 suffice — and it makes translation more tolerant of silent mutations, since many third-position substitutions still recruit the same tRNA and yield the identical amino acid.

Q5 How does alternative splicing increase protein diversity beyond the number of genes in a genome, and why was this discovery important for understanding the human genome?

Alternative splicing lets the spliceosome include or exclude different combinations of exons from the same pre-mRNA transcript, producing multiple distinct mature mRNAs — and therefore multiple distinct proteins — from a single gene. This was important because the human genome turned out to contain far fewer protein-coding genes (~20,000) than early estimates assumed would be needed to build the full complexity of a human body; alternative splicing (along with post-translational modification) resolved the apparent gap, showing that combinatorial regulation, not raw gene count, is a major source of biological complexity.

Concept mind map

How the ideas connect

Every key idea in this chapter, branching from the core concept — use it to see the whole picture at a glance.

helicaseDNA polymeraseleading strandlagging strandOkazaki fragmentsRNA splicingribosome sitesReplication & Expression
Infographic

The process, step by step

Step 1UnwindingHelicase opens the helix; primase lays RNA primers.
Step 2ElongationDNA polymerase III extends 5-prime to 3-prime along each strand.
Step 3Lagging strandSynthesised discontinuously as Okazaki fragments, later joined by ligase.
Step 4TranscriptionRNA polymerase makes pre-mRNA; introns are spliced out, a cap and tail added.
Step 5TranslationRibosome A, P and E sites move tRNAs to build the polypeptide.
Solved examples

Worked problems, step by step

Follow each solution line by line, then try to reproduce it on paper before moving on.

Example 1E. coli genome is 4.6 Mb; polymerase adds 1000 nt/s from one origin bidirectionally. Time to replicate?

  1. Two forks share the work: 4.6e6 / 2 = 2.3e6 nt per fork
  2. 2.3e6 / 1000 = 2300 s

Example 2Why can DNA polymerase not start a strand on its own?

  1. It can only add to an existing 3-prime OH
  2. An RNA primer supplies that starting 3-prime end
Practice problem set

Now you try

Work each one out first, then tap to reveal the worked answer.

1Why must the lagging strand be discontinuous?
DNA polymerase works only 5-prime to 3-prime, so on the strand running the other way it synthesises short Okazaki fragments moving backward from the fork.
2What is the role of DNA ligase?
Ligase seals the nicks between Okazaki fragments, joining them into a continuous strand.
3What are the three main steps of mRNA processing in eukaryotes?
Adding a 5-prime cap, splicing out introns, and adding a poly-A tail.
4What happens at the ribosome A, P and E sites?
A site accepts the incoming charged tRNA, P site holds the growing chain, and E site releases the empty tRNA.
5How does proofreading improve replication fidelity?
DNA polymerase 3-prime to 5-prime exonuclease activity removes mismatched bases so they can be replaced correctly.
6Why is transcription in eukaryotes separated from translation?
Transcription occurs in the nucleus and mRNA must be processed and exported before ribosomes in the cytoplasm translate it.