A codon is a sequence of three nucleotides that functions as a unit of information in the genetic code. During translation, codons in messenger RNA (mRNA) specify which amino acid is added to a growing protein, or signal that synthesis should stop. The term also applies to the corresponding triplets in DNA. In the standard genetic code, 64 possible codons comprise 61 amino-acid-specifying codons and three termination signals. (genome.gov)
Composition and notation
RNA codons are written using four letters: A for adenine, U for uracil, G for guanine, and C for cytosine. Because each of three positions can contain any of these four bases, there are possible combinations. Order matters: GCA specifies alanine, whereas CGA specifies arginine in the standard code. Codon sequences are conventionally written in the 5′-to-3′ direction. (educationalgames.nobelprize.org)
DNA versions use T, representing thymine, instead of U. Thus, ATG in a DNA coding strand corresponds to AUG in mRNA. During transcription, the RNA sequence is copied from the complementary DNA template strand; it therefore matches the coding strand apart from the replacement of T with U. Codon tables generally use RNA notation, although sequence databases may present equivalent tables using DNA letters. (nigms.nih.gov)
Reading frames
A reading frame determines how a nucleotide sequence is divided into successive triplets. Within an ordinary protein-coding sequence, codons are read consecutively without intervening punctuation and do not overlap. Starting one nucleotide later produces different triplets and potentially a different amino acid sequence. A single nucleotide strand therefore has three possible reading frames. (ncbi.nlm.nih.gov)
For example, the RNA sequence AUGGCUUAA, read from its first nucleotide, divides into AUG–GCU–UAA: methionine, alanine, and a stop signal. Starting at the second nucleotide instead yields UGG–CUU, specifying tryptophan and leucine. These assignments illustrate why identifying the correct frame is essential when interpreting a protein-coding gene. A triplet’s meaning depends on its placement in the translated sequence, not simply on its occurrence somewhere in an RNA molecule. (ncbi.nlm.nih.gov)
Recognition during translation
The ribosome reads mRNA and assembles amino acids in the order specified by its codons. Transfer RNA (tRNA) provides the adaptor between nucleotide information and amino acids. Each tRNA contains an anticodon, a three-base sequence that pairs with an mRNA codon in an antiparallel orientation, while carrying an amino acid at another part of the molecule. (nigms.nih.gov)
Recognition does not always require strictly conventional pairing at all three positions. Wobble base pairing, principally involving the third codon position and the first anticodon position, permits some tRNAs to recognize several codons. Consequently, a translation system does not require a separate tRNA type for every amino-acid-specifying codon. Francis Crick formulated the wobble hypothesis in 1966 to explain this relationship between codon recognition and the redundancy of the code. (pubmed.ncbi.nlm.nih.gov)
Start, stop, and synonymous codons
The usual start codon is AUG, which also specifies methionine at internal positions. Its role as an initiation signal depends on the surrounding sequence and the translation machinery. Alternative initiation codons occur in some organisms and organelles; for example, some bacterial coding sequences begin with GUG or UUG. Their initiation function must be distinguished from their ordinary amino acid assignments within a protein. (ncbi.nlm.nih.gov)
The standard stop codons are UAA, UAG, and UGA. During normal termination, these are recognized by protein release factors, rather than amino-acid-carrying tRNAs. Recognition triggers release of the newly synthesized chain from its attachment to tRNA. In eukaryotes, the release factor eRF1 recognizes all three standard stop codons. (nature.com)
Most amino acids have several codons, a property called degeneracy or redundancy. Codons specifying the same amino acid are synonymous. For example, UUU and UUC both specify phenylalanine, while leucine has six codons. Methionine and tryptophan each have only one in the standard code: AUG and UGG, respectively. Redundancy does not normally mean ambiguity: each amino-acid-specifying codon has one assignment within a given code. (ncbi.nlm.nih.gov)
Variants of the genetic code
The standard code is widespread, but not universal. Alternative assignments occur in mitochondria and certain organisms. In the vertebrate mitochondrial code, UGA specifies tryptophan rather than termination, and AUA specifies methionine rather than isoleucine. Accurate interpretation of a coding sequence therefore requires the appropriate translation table for its organism and cellular compartment. NCBI records distinguish these codes through translation-table identifiers. (ncbi.nlm.nih.gov)
Mutations and codon choice
A mutation within a coding sequence can have several consequences. A synonymous substitution preserves the encoded amino acid; a missense substitution changes it; and a nonsense substitution converts an amino-acid-specifying codon into a premature stop signal. Insertions or deletions whose lengths are not multiples of three can produce a frameshift, changing the downstream grouping into codons and often introducing an early stop. (ncbi.nlm.nih.gov)
Synonymous codons are not necessarily functionally interchangeable. Experiments in budding yeast showed that changing synonymous codon composition can alter mRNA stability and translation elongation. Studies of engineered sequences in Escherichia coli likewise demonstrate interactions between codon composition, RNA structure, and protein production. These findings underpin codon optimization in gene expression research, while showing that codon effects depend on sequence and biological context. (pmc.ncbi.nlm.nih.gov)
Experimental identification
In 1961, Marshall Nirenberg and Heinrich Matthaei showed that synthetic RNA consisting of repeated uracil directed a cell-free bacterial system to produce polyphenylalanine, establishing a key connection between UUU and phenylalanine. Further experiments using synthetic RNA and defined triplets helped decipher the remaining assignments. Nirenberg, Har Gobind Khorana, and Robert Holley received the 1968 Nobel Prize in Physiology or Medicine for interpreting the genetic code and its role in protein synthesis. (njc.rockefeller.edu)