This article is an open-access article distributed under the terms and conditions of the Creative Commons Attribution license (
The 3x redundancy of the Genetic Code is usually explained as a necessity to increase the mutation-resistance of the genetic information. However recent bioinformatical observations indicate that the redundant Genetic Code contains more biological information than previously known and which is additional to the 64/20 definition of amino acids. It might define the physico-chemical and structural properties of amino acids, the codon boundaries, the amino acid co-locations (interactions) in the coded proteins and the free folding energy of mRNAs. This additional information, which seems to be necessary to determine the 3D structure of coding nucleic acids as well as the coded proteins, is known as the
Mapping between messages in nucleic acid and protein alphabet is a fascinating story, a story that still unfolding. It is about to understand the rules of information transfer between DNA and proteins. First of all it is not only a biochemical puzzle and much of the early methods for devising codes came from combinatorics, information theory. Four and 20 (number of bases and amino acids) seems to be magical numbers with amazingly many possible mathematical connections between them [
The existence of a
Another brilliant code was created by Francis Crick [
There were many theoreticians involved in the invention of the Genetic Code. Finally,
Why do we have 64 triplets for coding 20 amino acids - more than three times the number needed? Explaining away this excess became a major preoccupation of coding theorists. One intelligent explanation is that the redundancy confers a kind of error tolerance, in that many mutations convert between synonymous codons. When a mutation does alter an amino acid, the substitute is likely to have properties to those of the original. Alternatively the mutation is likely to be a stop codon which completely aborts the wrong translation. Another possible explanation is that the
Crick had, of course, his own explanation: forget any logical connection between codons and coded amino acids, it is just a “frozen accident” or by other words “that is what we got, lets we like it…”
Evolution often occurs in stages, one step is followed by a plateau before the next step is possible to take. It is true even for the development of scientific thinking and understanding. The 50-es and 60-es were very fertile for biology and the foundation of the recent molecular biology was laid. It took 30+ years to recognize, that the discovery of DNA was not the discovery of
The number of mRNA images in Nucleic Acid Structure Database (NDB) is very limited [
The functional significance of the mRNA secondary structure is not known, therefore wobble base replacement is a widely used and accepted method to eliminate the folding energy of mRNA and achieve higher translational efficiency [
Frame-shift is a major concern regarding the translation. Nirenberg’s Genetic Code seems not to give any protection against the possible occurrence of frame-shifts even if it gives some promises to reduce the catastrophic consequences of wrong codon readings. However a second look at the base composition of codons (64 as it is), or the usage-weighed variants from different Species Specific Codon Usage Tables [
This codon related periodic variation of GC content means that there is a periodic pattern of folding energy (dG) along the mRNA, which distinguishes the central codon base from the 1st and 3rd and forms a physico-chemical barrier or boundary between the codons. This is a statistical rule which doesn’t apply for every single codon, but still shows a general tendency that there is some potential protection against frame shifts in the Nirenberg’s
There is even another completely separate line of evidence suggesting that codon positions are different, and the central codon position has a very special role. There always has been an effort to connect codons to their coded amino acids. The wobble base lost its importance because of its interchangeability. Most scientific efforts focused on to find stereo-chemical compatibility (spatial fitting) between the atomic geometry defined by 2 or 3 nucleic acid bases and the corresponding geometry defined by the residue of the coded amino acid [
A common regularity in an arrangement of codons and amino acids provides a strong support for the evolutionary connection between mRNA and coded proteins.
A
I interpret these results as a clear-cut answer for the Woese vs. Crick dilemma: there
Although there can be seen a quite large correlation between amino acid codons and amino acid properties (an interesting finding as such), one has to be cautious in saying that this correlation has a functional significance. Such significance could eventually exist there (e.g. in determining the structure of mRNA), but more evidence is needed. Before that this correlation still can be a fruitful hypothesis for further research.
The strong connection between codon structure and physicochemical properties of coded amino acids (the existence of
The interaction between restriction enzymes (RE) and their recognition sequences (RS) are known to be very specific and fortunately numerous such interactions are visualized and available from PDB. A review of the seven available crystallographic studies [
There was an idea published in early 80-s [
We constructed a bioinformatics tool to collect data of co-locating amino acids from known protein structures, listed in the PDB, for statistical analyses [
1) Co-locating amino acids are physico-chemically compatible with each other, i.e. large and small, positive and negative, hydrophobe and hydrophobe amino acids are preferentially co-located with each other. The novelty of this observation is that physico-chemical rules apply already on residue level and do not necessarily need large, complex interfaces of interacting proteins [
2) Co-locating amino acids are preferentially coded by partially complementary codons, where the 1st and 3rd bases are complementary but the 2nd may but not necessarily are complementary to each other (
The propensities for the 20×20=400 possible amino acid pairs were monitored in 81 different protein structures with the SeqX tool. The tool detected co-locations when two amino acids were within 6 Å distance of each other (neighbors on the same strand were excluded). The total number of co-locations was 34,630. Eight different complementary codes were constructed for the codons (2 optimal and 6 suboptimal). In the two optimal codes, all three codon residues (123) were complementary (C) or reverse complementary (RC) to each other. In the suboptimal codes, only two of three codon residues were C or RC to each other (12, 13, 23), while the third was not necessarily complementary (X). (For example, Complementary Code RC_1X3 means that the first and third codon letters are always complementary, but not the second and the possible codons are read in reverse orientation. P/N (positive/negative) ratio indicates the proportion of co-locating amino acids coded by the defined codon complementary rules.
These observations lead us to conclude that there are significant additional functional and structural connections between codons and coded amino acids to that what was described by Nirenberg and is known as the
The 64/20 Genetic Code is redundant, mainly because the 3rd codon bases, in most codons, are interchangeable without any consequence on the sequence of the coded protein. The only information expected from the DNA to the protein syntheses is the coding of amino acids, because it is believed, that the only information necessary to correct protein folding is only the correct amino acid sequence itself.
These believe is based on
The annealing action of these (mainly protein) chaperons is described as providing supplemental information for systems that do not otherwise have a definite ground state. Extensive research into chaperonin assisted protein folding [
By other words there is some 3x excess of information before translation, and there seems to be a shortage of information after translation. It is logical to assume, that folding information is stored in the redundant codons, more concretely in the wobble bases. The literature is actually rather rich with observations connecting the wobble bases to some structural feature of the coded proteins [
The preferential coding of co-locating amino acids by partially complementary nucleic acids, (for example by 5′>ANG>3′/3′<TNC<5′ pattern) immediately suggests a role for the wobble bases. They are integrated parts of codons, defining amino acid co-locations. They are not randomly chosen, but logically selected: the wobble base of Xc codon (defining Xa amino acid) is defined by the first residue of codon Yc (which is coding Ya amino acid) if that two amino acids (Xa and Ya) are co-locating and
Protein structures contain many amino acid co-locations (immediate neighbors on the same chain are excluded). Suppose that preferential partial complementarity coding of amino acids is not a rarity, but it is a rule. In that case the signs of non-randomness of wobble base selection should be seen not only in a small subset of proteins but even in very large data sets, like the species specific codon usage frequency tables.
Statistical analyses of A, T, G, C frequencies at 1st, 2nd, and 3rd codon positions in 113 species specific Codon Usage Frequency Tables and 87 protein structures showed strongly significant internal correlation between the frequency of nucleic acid bases at different codon positions. This strong relationship made it possible to predict the frequency of all possible wobble bases in all the 64 codons in all the 113 species (P<1.3E-64, N=113) and all the 87 proteins (p<1.1E-28, n=87) [
These strong correlations wouldn’t be possible with random selection of wobble bases. Therefore we concluded, that synonymous codons are not interchangeable with each other without disturbing the internal order of bases in integrated codon systems like a native mRNA or a species specific Codon Usage Frequency Order.
There are more than observations provided by theoretical and computational biology which are indicating, that native, natural proteins, as well as their coding sequences, are much more than the sequential collection of their building blocks. They are an integrated, interconnected system.
1) Wobble base mutations are expected to be “silent” without any consequences for the biological functions or phenotypes. They are often not. “Silent” polymorphism or mutation affects a) substrate specificity [
2) It has recently become clear that the classical notion of the random nature of mutations does not hold for the distribution of mutations among genes: most collections of mutants contain more isolates with two or more mutations than predicted by the mutant frequency on the assumption of a random distribution of mutations. Excesses of multiples are seen in a wide range of organisms, including riboviruses, DNA viruses, prokaryotes, yeasts, and higher eukaryotic cell lines and tissues. In addition, such excesses are produced by DNA polymerases
3) A compensatory mutation occurs when the fitness loss caused by one mutation is remedied with a second mutation at a different site in the genome.
Often it occurs in the same gene, alters the protein sequence [
Uneven concentration of mutations of smaller distances and their compensatory character are further indications of the integration and interconnectivity of codons in the same gene and consequently, conservation of structurally critical amino acid connections (but the amino acids) in the coded proteins.
However it should be noticed, that the distribution of mutations in directed evolution experiments is so complex phenomenon that it cannot directly be used to support the proteomic code idea.
The preferential partial complementarity coding of co-locating amino acids (
Biological rules are, of course, always statistical rules, probabilities and tendencies. Nucleic acids as well as proteins have many possible configurations where one or a few are expected to dominate and define the main and characteristic configuration. Coding- and coded sequences might have their own range of more or less different folding potentials. However when a protein is generated on the surface of ribosome the coding- and coded sequences are very close to each other. This temporary intimate closeness is a possibility for coding sequences for transferring folding information to coded proteins, information that is additional to that these proteins already have in their amino acid sequences. Some ideas how it is possible are sketched in
The RNA-assisted protein folding is an interesting theory which derived recently [
The
The theory of Ikehara [
The primeval genetic code continued to develop toward a more complex SNS-type primitive genetic code (S: G or C) containing 16 codons and encoding 10 amino acids (L, P, H, Q, R, V, A, D, E, G) before the recent 64 codon/20 amino acid-type genetic code became established.
Furthermore, Ikehara concluded from the analysis of microbial genes that newly-born genes are products of nonstop frames (NSF) on antisense strands of microbial GC-rich genes [GC-NSF(antisense)] and from SNS repeating sequences [(SNS)n] similar to the GC-NSF(antisense).
The similarity between GNC/SNS-type primitive codons (which are expressed even from the reverse-complement strands as GC-rich non-stop genes) and the
The Nirenberg’s
1) Codon boundaries are physico-chemically defined to a certain degree, which theoretically should give some protection against frame-shifts.
2) Codon residues are not randomly assigned, but there is a connection between codon architecture and the physicochemical properties of the coded proteins.
3) Amino acids preferentially interact with their codons (studied in restrictions endonucleases).
4) Co-locating amino acids are preferentially coded by partially complementary codons which create inter-connectivity between structurally important amino acids.
5) Wobble bases are not randomly assigned at all, their frequency is statistically well predictable from the frequency of bases at other codon positions.
6) Wobble base redundancy makes it possible the development of codon integration in coding sequences which might be used for compensatory mutations. This is the second line of defense against mutations (after the known tolerance provided by the coding redundancy).
7) The internally inter-connected and integrated system of codons makes it possible that coding sequences provide a mold for structure forming of coded proteins and function as
It is concluded that the redundant
The author wishes to thank for the friendly attention and advices of Dr. G.L.G. Miklos (Secure Genetics Pty Ltd, Australia). The continuous support and help of Dr. P. S. Agutter (BMC,
Co-location of Codon-like Triplets and Amino Acids in RE-RS Complexes. Examples for co-locations of amino acids in restrictions endonucleases (RE) and codon-like triplets in restrictions enzyme recognition sites (RS). The name of enzyme, the name and position of nucleic acid bases and amino acids are indicated. Four amino acids located at overlapping-codon like base sequence in
Complementary codes vs. amino acid co-locations (modified from [
Comparison of 12 randomly selected protein and corresponding mRNA structures (modified from [
Comparison of the protein and mRNA structures (modified from [
RNA assisted protein loop formation. Translation begins with the attachment of the 5' end of a mRNA to the ribosome (A). Ribonucleotides are indicated by blue + and the 1st and 3rd bases in the codons by blue lines, while the 2nd base positions are left empty. A positively charged amino acid [(+) and red dots], for example arginine, remains attached to its codon. The mRNA forms a loop because the 1st and 3rd bases are locally complementary to each other in reverse orientation (B). The growing protein is indicated by red circles (o). When translation proceeds to an amino acid with especially high affinity to the mRNA-attached arginine, for example a negatively charged Glu or Asp [(−) and blue dot], the charge attraction removes the Arg from its mRNA binding site and the entire protein is released from the mRNA and completes a protein loop (C). The protein continues to grow toward the direction of its carboxy terminal (COOH). (Figure is reproduced from [
RNA-assisted (translational) protein folding. There are three reverse and complementary regions in a mRNA (blue line, A): a-a′, b-b′, c-c′, which fold the mRNA into a T-like shape. During the translation process the mRNA unfolds on the surface of the ribosome, but subsequently refolds, accompanied by its translated and lengthening peptide (red dotted line, B–F). The result of translation is a temporary ribonucleotide complex, which dissociates into two T-shape-like structures: the original mRNA and the properly folded protein product (G). The red circles indicate the specific, temporary attachment points between the RNA and protein (for example a basic amino acid) while the blue circles indicate amino acids with exceptionally high affinity for the attachment points (for example acidic amino acids); these capture the amino acids at the attachment point and dissociate the ribonucleoprotein complex. Transfer-RNAs are of course important participants in translation, but they are not included in this scenario. (Figure is reproduced from [
The Common Periodic Table of Codons and Amino Acids.
Effects of a single codon residue on the structure of the amino acids.
The 16 amino acids coded by symmetric codons were sorted into