Everything in the “Toward autonomous healing systems” path from here on — reading a genome, editing it, delivering the edit — depends on one idea: biology runs on sequence information, and that information flows in a specific direction. This module builds that idea carefully, because almost every misconception about genetics comes from getting the direction or the storage wrong.

What DNA stores, physically

Established DNA is a linear polymer of four building blocks — the bases adenine, thymine, guanine, and cytosine (A, T, G, C). The information is the order of those bases, exactly as the information in this sentence is the order of its letters. There is nothing else. No hidden layer, no analog magnitude — just a sequence drawn from a four-letter alphabet.

The double helix matters for one reason above all: the two strands are complementary. A always pairs with T, G always pairs with C, held together by the hydrogen bonds you met in M-Chem-04. Because each strand specifies the other, the molecule carries its own backup. Pull the two strands apart and each is a template for rebuilding its partner. That single structural fact — complementarity via weak forces — is what makes faithful copying possible, and it is why the weak-forces module was a prerequisite.

The genius of the structure, as Watson and Crick understatedly noted, is that the pairing rule immediately suggests a copying mechanism. Storage and replication are not two separate problems that biology had to solve; the geometry solves both at once.

From gene to protein: the two-step flow

Established A gene is a stretch of DNA sequence. Turning it into a working molecule takes two steps:

Transcription. The relevant DNA stretch is copied into a single strand of messenger RNA (mRNA). RNA uses the same alphabet with one substitution (uracil, U, replaces thymine). This is a faithful re-encoding, not a translation — the language is still nucleotides. Translation. A ribosome reads the mRNA three bases at a time. Each triplet — a codon — specifies one amino acid. The ribosome strings the amino acids together in the order the codons dictate, and the resulting chain folds into a protein.

The mapping from 64 possible codons to 20 amino acids (plus stop signals) is the genetic code, and it is very nearly universal across all life — one of the strongest pieces of evidence that everything alive shares a common ancestor. It also means an instruction written in one organism can, in principle, be read in another, a fact biotechnology exploits constantly.

A worked example: reading a codon

Take the DNA triplet GAG on the coding strand. Transcribed to mRNA it stays GAG. The ribosome reads GAG and inserts the amino acid glutamate. Now change a single base: GTG on the DNA becomes GUG in the mRNA, which codes for valine.

That one-letter change — glutamate to valine at position six of the beta-globin protein — is sickle-cell disease. A single base out of three billion, altering a single amino acid, deforms haemoglobin enough to distort red blood cells. Hold onto this example: it is the exact target of one of the first approved CRISPR therapies you will meet in M-Med-02, and it shows why reading and editing sequence at single-base resolution is not academic.

What the central dogma actually says

Established Francis Crick's 1958 “central dogma” is the most misquoted idea in biology. It is not the claim that information only ever goes DNA to RNA to protein. Crick was careful: his claim was about a prohibition. Once sequence information has been translated into protein, it cannot flow back out — not into nucleic acid, not into another protein.

He explicitly permitted transfers the simple arrow forbids. RNA can be copied back into DNA (reverse transcription — this is exactly how the retroviruses of M-Bio-02, including HIV, work, and it is how the mRNA in a cell can leave a permanent DNA trace). RNA can template RNA. DNA can, in unusual cases, template protein directly. What is forbidden is the reverse flow out of protein. And notice the deep connection to the previous modules: prions look like they violate this, propagating “information” protein-to-protein — but they do not transfer sequence, only shape, which is precisely why Prusiner's prion concept was so hard to accept and why the dogma survived it intact.

Why “the genome as blueprint” is the wrong metaphor

Frontier It is tempting to picture the genome as a blueprint of the organism. It is not. A blueprint contains a scaled picture of the finished thing; the genome contains no picture of a body anywhere. It is closer to a recipe or, more precisely, a list of parts and the rules for when to make them. The organism is not read off the genome; it is built by a process the genome only partially specifies, in constant interaction with the cellular and physical environment.

This distinction is not pedantry — it is the reason the next module exists. Sequencing a genome gives you the complete recipe text, and it turns out that having the recipe is very far from understanding the dish. That gap between sequence and meaning is where modern medicine actually gets hard.

Checkpoint

Crick's central dogma is often paraphrased as “DNA makes RNA makes protein.” Why is that paraphrase misleading about what he actually claimed?

Show answer

Crick's real claim was about what information transfers are forbidden, not which are common. He asserted that once sequence information has passed into protein, it cannot flow back out of protein into nucleic acid or into another protein. He explicitly allowed RNA to DNA (reverse transcription) and RNA to RNA. So the dogma is a statement about the one-way street out of protein, not the simple linear arrow the paraphrase suggests — which is why retroviruses (RNA to DNA) surprised the public but not the dogma.

Interactive · planned

An animated codon reader would slot in here: drag along a strand of mRNA and watch each three-letter codon pull in its matching amino acid, with the genetic-code table lighting up. It would turn the abstract “triplet code” into something the reader can operate.