Molecular information

What is the genetic code?

Sixty-four codons. Twenty amino-acid meanings. A molecular dictionary shaped by error and demand.

The 64 RNA codons arranged in a genetic code table and labeled by their amino-acid meanings.
The dictionary. Sixty-four RNA codons map to amino acids and termination signals.

It looks like a table. But the table hides a geometry.

The same genetic code drawn as a landscape whose height represents amino-acid polarity.
The same dictionary, lifted by chemistry. Height is amino-acid polarity. Nearby codons tend to encode the same, or chemically similar, amino acids.

The code is not random

The genetic code translates three-letter nucleotide words—codons—into amino acids. Because there are more codons than amino-acid meanings, several words can mean the same thing.

The redundancy is organized. Synonymous codons cluster. A one-letter error often leaves the meaning unchanged; when it does not, it tends to substitute an amino acid with related chemistry.

Nearby molecular words tend to carry nearby chemical meanings.

A code is a compromise

Perfect error tolerance would be easy: give every codon the same meaning. It would also be useless. Proteins require a chemically varied vocabulary.

  • Robustness Likely mutations and mistranslations should change molecular meaning as little as possible.
  • Demand Proteins require amino acids in unequal amounts and with distinct chemical roles.

The standard code lies near the optimal trade-off between mistranslation robustness and amino-acid demand. This balance emerges when the code is compared with alternative codon–amino acid assignments.

A noisy information channel

Codons are symbols. Amino acids are meanings. Translation is a channel, and molecular recognition is noisy.

Rate-distortion theory asks how much information can be transmitted through a noisy channel at an acceptable distortion. The genetic code poses the same problem in molecular form: how much specificity is worth paying for, given the damage caused by errors and the value of a diverse chemical vocabulary?

Errors give the code a geometry

Two codons are neighbors when the molecular machinery is likely to confuse them. The resulting error graph defines a geometry on codon space.

Smooth codes assign similar meanings to neighboring vertices. Close to the coding transition, the first non-uniform patterns are therefore the smoothest modes of this graph.

This is a physical hypothesis about constraints on the architecture of a code. It is not a claim that the historical origin of the standard genetic code is settled.

Why this code?

Historical contingency, chemistry, biosynthetic history, and selection for error tolerance may all have contributed. They need not exclude one another.

The genetic code is a dictionary shaped by the errors through which it must be read.

The experimental route from codons to the genetic code runs through Nirenberg, Crick and Brenner, and the classic experiments collected in Landmark Experiments in Biology.