Path 4: RNA World (Replicator-First Hypothesis)

Rationale: Life began with self-replicating RNA molecules that both stored genetic information and catalyzed reactions. Initially, a purely RNA-based biology (no DNA or proteins) existed, which later evolved to use proteins and DNA. RNA can act as a gene and an enzyme. The discovery of ribozymes (like self-splicing introns and the RNA-based ribosome active site) supports this. A single type of molecule carrying information and function could simplify early life. Laboratory evolution shows RNA can catalyze its own replication under certain conditions. The RNA World hypothesis is widely accepted as a stepping stone because modern biochemistry still has vestiges of an RNA era (e.g. ATP, NAD, etc. are ribonucleotide-based).

Prerequisites: Prebiotic synthesis of ribonucleotides (a notoriously difficult step: ribose sugar, nucleotide base, and phosphate must assemble in the right way), mechanisms for RNA chain polymerization without enzymes, and theoretical understanding of how an “RNA replicase” ribozyme could emerge by chance.

Dependencies: A pure RNA world scenario might need input from Path 1 (to accumulate nucleotides in a soup) or Path 7 (clays to catalyze RNA formation). Eventually it likely merges with Path 6 (packaging in protocells) because naked RNA strands would be diluted or degraded.

Signs of Progress: Successful synthesis of RNA’s building blocks under plausible conditions (e.g. Sutherland’s 2009 study showing formation of activated nucleotides in a model prebiotic reaction). Creation of an RNA molecule that can copy another RNA (partial success: researchers evolved an RNA polymerase ribozyme that extends RNAs up to 95 bases). A key milestone was the 2009 experiment by Lincoln and Joyce demonstrating two RNA enzymes that catalyze each other’s synthesis, achieving self-sustained replication (albeit in a controlled setup). Continued improvements in such systems, or the emergence of an RNA replicator in a simulation, would strongly support this path.

Base Camp BC4.1: RNA Structure and Catalysis

Scope: Understand RNA as a molecule that can both store information (as sequence of nucleotides) and fold into functional 3D shapes (ribozymes). Cover the chemistry of RNA (ribose sugar, phosphate backbone, four bases) and examples of catalytic RNAs in modern biology.

Stepping-Stones: (a) Review nucleotide structure: know ribose vs deoxyribose (2’–OH group in RNA is crucial for both catalysis and why RNA is less stable than DNA). Recognize the bases A, U, G, C, and the Watson–Crick pairing (A:U, G:C) that lets RNA carry info like DNA. (b) Examine how single-stranded RNA folds: intramolecular base pairing can create helices, loops, bulges, etc. The concept of secondary structure (hairpins, stem-loops) and tertiary interactions (pseudoknots, coaxial stacking). (c) Study key ribozymes: the group I and II introns (RNAs that splice themselves out of transcripts – proving RNA can cut and ligate RNA), RNase P (an RNA enzyme in bacteria that cleaves tRNA precursors), and the peptidyl transferase center of the ribosome (which is an RNA-based catalyst for protein synthesis). Understand the significance: the ribosome’s active site is all rRNA, meaning modern protein synthesis is fundamentally an RNA enzymatic activity. (d) Investigate how ribozymes achieve catalysis: general strategies include acid-base catalysis (often using metal ions coordinated by RNA functional groups) and bringing reactants into proximity/orientation. Example: the hammerhead ribozyme cleaves itself by forming a specific loop that positions a 2’–OH for nucleophilic attack on the adjacent phosphate. (e) Connect structure to evolution: why was RNA likely the first genetic polymer? Because it can replicate (through base pairing) and also perform biochemistry – here emphasize that only RNA (and not DNA or peptides alone) can do both in one molecule, which addresses the “which came first” problem by suggesting RNA did.

Base Camp BC4.2: Prebiotic Nucleotide Synthesis (the “Nucleoside Problem”)

Scope: Tackle one of the hardest questions for the RNA World: how did the building blocks of RNA (ribose sugar, nucleotide bases, phosphate) form and combine on the early Earth? Survey both classical and contemporary solutions proposed by chemists.

Stepping-Stones: (a) Understand the challenge: forming ribose (a 5-carbon sugar) from presumed prebiotic feedstocks (likely formaldehyde) via the formose reaction yields a messy mixture of sugars, not just ribose, and ribose is relatively unstable (degrades to brown tars). Also, attaching a base to ribose to form a nucleoside is chemically difficult (requires activating the base or sugar, and usually gives low yields). Summarize early pessimism: Orgel famously highlighted this nucleoside problem in the 1990s as a strike against RNA-first. (b) Investigate the breakthrough by Sutherland (2009): a different approach where bases and ribose form together from common precursors (like cyanamide, glycolaldehyde, glyceraldehyde, phosphate) – this yielded activated pyrimidine ribonucleotides (like cytosine and uracil ribonucleotides) in surprisingly high yield. Work through the scheme: formation of 2-aminooxazole as a key intermediate that eventually leads to β-ribocytidine-2’,3’-cyclic phosphate, for example. This showed that ribose and base need not be made separately then combined; they can be assembled in an orchestrated route. (c) Learn other partial solutions: e.g. Stanley Miller’s experiments with HCN produced adenine (5 HCN → adenine). Borate minerals (borax) can stabilize ribose in solution, mitigating the degradation problem – know the significance of that in Benner’s work. (d) Examine alternative nucleic acids: maybe the first genetic polymer wasn’t RNA exactly, but a simpler analog like PNA (peptide nucleic acid) or TNA (threose nucleic acid) which are easier to form? Some propose that an earlier polymer with similar information properties might have served until RNA took over. Evaluate pros and cons of that idea. (e) Summarize current status: Are we now confident all four RNA bases and ribose could form prebiotically? We have routes for pyrimidines (Sutherland’s), some for purines (perhaps via formamide chemistry or HCN polymers, and indeed adenine’s formation is classic). Emphasize that phosphate availability is also something Sutherland’s chemistry cleverly circumvented by using phosphate as a catalyst early which then becomes incorporated. (f) Discuss chirality: ribose has multiple stereoisomers; prebiotic routes yield racemic mixtures. Life uses D-ribose. Consider how homochirality might emerge (e.g. crystallization-induced resolution, or chiral templates).

Base Camp BC4.3: RNA Replication and In Vitro Evolution

Scope: Explore how RNA molecules can replicate or be evolved in the lab, to infer how the first self-replicating RNA might have arisen. This includes classic experiments like Spiegelman’s Qβ replication, as well as modern in vitro selection of ribozymes.

Stepping-Stones: (a) Recall Spiegelman’s 1967 experiment: he took a replicating RNA virus (Qβ phage RNA, 4500 bases) and replicated it in a test tube with Qβ replicase enzyme, repeatedly transferring samples to fresh solution. After many transfers, the RNA evolved into a much shorter (~200 base) form dubbed “Spiegelman’s Monster” that replicated faster (due to being pared down to essentials). Understand the implication: replication + selection ⇒ smaller, more efficient replicators dominate; but note that was enzyme-driven, not prebiotic. (b) Learn about in vitro evolution (SELEX – Systematic Evolution of Ligands by EXponential enrichment): how researchers start with a random pool of RNA (like 10¹⁵ different sequences), apply a selection pressure (e.g. binding to ATP or catalyzing a reaction), isolate the winners, amplify them (usually by RT-PCR), mutate and repeat. This method has yielded artificial ribozymes: e.g. an RNA ligase ribozyme (Bartel’s lab, 1993) that can join RNAs, or a ribozyme that adds nucleotides (Johnston et al. 2001) albeit with limited efficiency. List a few achievements: the best RNA polymerase ribozyme currently (Horning & Joyce, 2016) can copy RNAs up to 95 bases with some accuracy – still not copying itself fully, but getting close. (c) Investigate the RNA replication cycle: if one had an RNA that can polymerize others, you still need strand separation (to avoid just making double-strands) – perhaps temperature cycles (like a day-night or hydrothermal oscillation) could do PCR-like denaturation. Also, initial replication might have been non-enzymatic (just template directed chemistry, then eventually RNA took over catalysis). (d) Study non-enzymatic replication attempts: e.g. see how activated nucleotides can polymerize on a template without enzymes – progress has been made (e.g. Powner & Sutherland achieved copying short templates using chemically activated nucleotides and helper molecules). Understand its limits (template sequences matter, copying stalls at certain lengths, etc.). (e) Consider the error threshold (Eigen’s paradox): early replicators (if short, error rates aren’t huge issue; but as they grow, fidelity needs to improve or information can’t be maintained). This motivates thinking about whether multiple short cooperating RNAs could collectively encode a system rather than one long RNA. (f) Recount one or two experiments demonstrating Darwinian evolution in RNA systems: e.g. Lincoln and Joyce (2009) made two ribozymes that catalyze each other’s synthesis – a cross-replicating system that undergoes exponential growth until resources deplete. They even observed competition when variants were present – an RNA quasispecies. This is a big proof-of-concept that RNA-based life is not a fantasy: the core is achievable.

Base Camp BC4.4: Transition to DNA/Protein World

Scope: Learn how the hypothesized RNA World might have given way to the modern DNA/RNA/protein world. Why did DNA and proteins take over, and how could that transition occur?

Stepping-Stones: (a) Recognize the limitations of RNA: while versatile, RNA is chemically fragile (the 2’–OH makes it prone to hydrolysis) and its catalytic repertoire, though broad, is less chemically diverse than protein enzymes (which have 20 amino acid side chains to play with). Over time, division of labor likely became advantageous: DNA (more stable, double-stranded) for information storage, proteins (more flexible catalysis) for biochemistry, with RNA mostly in an intermediary (mRNA, tRNA, rRNA) role. (b) Investigate the origin of translation (the RNA-protein partnership): this is the hardest step in origin-of-life research according to many. The ribosome is an RNA machine, but it’s assisted by proteins today – how could an RNA-only translation system arise? Consider models: perhaps at first, short amino acid chains (peptides) were synthesized by ribozymes (there are ribozymes now that can form peptide bonds in vitro without a full ribosome). Possibly amino acids originally served as cofactors (some ribozymes bind amino acids to expand catalytic function – a hint that RNA could recruit amino acids before there were coded proteins). Over time, specific RNAs might have started templating amino acids in order – that’s essentially the genetic code’s origin. (c) Study the “RNA-peptide coevolution” hypothesis: small proteins (or modified peptides) could strengthen ribozymes or give new functions, and RNAs that coded those useful peptides would be favored – forging a link between sequence and function that starts rudimentary coding. (d) Consider the genetic code: it’s universal (with minor deviations) – implying a single origin. There are various theories (stereochemical theory – certain codons have chemical affinity for their amino acids; coevolution theory – code structure reflects biosynthetic pathways; frozen accident theory by Crick – largely chance but became frozen once established). One should know the patterns (e.g. similar amino acids have similar codons) and possible reasons (error minimization: code seems to minimize impact of mutations). (e) Look at experiments on ribozymes that can add amino acids to growing peptide (there are artificial ribozymes that do a primitive form of translation – e.g. ribozyme that catalyzes peptide bond formation). Also, tRNA – these small adaptor RNAs are likely relics of the RNA World (they’re highly conserved in structure). The fact that the ribosome is a ribozyme strongly suggests that translation originated in an RNA context. (f) Examine origin of DNA: DNA is basically a modified RNA (2’–H instead of 2’–OH, and thymine instead of uracil). One theory: DNA was invented by viruses (some virus in an RNA world started using DNA for its genome to evade host defenses or increase stability). Regardless, once cells could make DNA (via ribonucleotide reductase enzymes), DNA’s stability would cause an “genetic takeover.” Learn about ribonucleotide reductase (an ancient enzyme). (g) Summarize a plausible chronology: RNA world (RNAs do info & catalysis) → RNA-protein world (RNA still genes, but proteins increasingly handle catalysis; the genetic code emerges) → modern world (DNA genomes, RNA intermediates, protein enzymes). Emphasize any evidence: e.g. the fact that many cofactors (ATP, NADH, Acetyl-CoA, etc.) include ribonucleotides hints that early metabolism was run by RNAs or ribozymes that used those cofactors; proteins later took over but kept the RNA-derived cofactors.

(Through Base-Camps 4.1–4.4, one builds a comprehensive understanding of the RNA World hypothesis from both chemical and biological angles. This equips a researcher to conduct experiments like ribozyme selections, to critically analyze new findings (e.g. a newly found ribozyme or synthetic pathway), and to theorize about bridging stages like the origin of coding. It forms one of the core pillars of origin-of-life research expertise.)

Bibliography (Path 4)

← Back to Matterhorn – Origin of Life