Molecular information
What is the genetic code?
Sixty-four codons. Twenty amino-acid meanings. A molecular dictionary shaped by error and demand.
It looks like a table. But the table hides a geometry.
The code is not random
The genetic code translates three-letter nucleotide words—codons—into amino acids. Because there are more codons than amino-acid meanings, several words can mean the same thing.
The redundancy is organized. Synonymous codons cluster. A one-letter error often leaves the meaning unchanged; when it does not, it tends to substitute an amino acid with related chemistry.
Nearby molecular words tend to carry nearby chemical meanings.
A code is a compromise
Perfect error tolerance would be easy: give every codon the same meaning. It would also be useless. Proteins require a chemically varied vocabulary.
- Robustness Likely mutations and mistranslations should change molecular meaning as little as possible.
- Demand Proteins require amino acids in unequal amounts and with distinct chemical roles.
The standard code lies near the optimal trade-off between mistranslation robustness and amino-acid demand. This balance emerges when the code is compared with alternative codon–amino acid assignments.
A noisy information channel
Codons are symbols. Amino acids are meanings. Translation is a channel, and molecular recognition is noisy.
Rate-distortion theory asks how much information can be transmitted through a noisy channel at an acceptable distortion. The genetic code poses the same problem in molecular form: how much specificity is worth paying for, given the damage caused by errors and the value of a diverse chemical vocabulary?
Errors give the code a geometry
Two codons are neighbors when the molecular machinery is likely to confuse them. The resulting error graph defines a geometry on codon space.
Smooth codes assign similar meanings to neighboring vertices. Close to the coding transition, the first non-uniform patterns are therefore the smoothest modes of this graph.
This is a physical hypothesis about constraints on the architecture of a code. It is not a claim that the historical origin of the standard genetic code is settled.
Why this code?
Historical contingency, chemistry, biosynthetic history, and selection for error tolerance may all have contributed. They need not exclude one another.
The genetic code is a dictionary shaped by the errors through which it must be read.
Further
The experimental route from codons to the genetic code runs through Nirenberg, Crick and Brenner, and the classic experiments collected in Landmark Experiments in Biology.
- 2007 A model for the emergence of the genetic code as a transition in a noisy information channel · Journal of Theoretical Biology
- 2008 Rate-distortion scenario for the emergence and evolution of noisy molecular codes · Physical Review Letters
- 2010 A colorful origin for the genetic code · Physics of Life Reviews
- 2019 Protein: The physics of amorphous evolving matter · Reviews of Modern Physics
- 2026 Seo Y, Tlusty T & Jo J — The genetic code at the balance point of error and demand · PLOS Computational Biology 22, e1014613