We analyse here the definition of the
The concept of the gene was introduced before the onset of molecular biology, in the wake of the work of Mendel, Morgan and their successors in the early twentieth century (Mendel
Originally, before the molecular carriers of function were understood and the coding aspect came to the foreground, a gene had been conceived as a simultaneous unit of inheritance, mutation and function. The principles of mutation are easy to understand at the biochemical level. The basic type of mutation is the exchange of a single nucleotide in the DNA. A single nucleotide, however, is too small to count as a unit of function. Such mutations may affect one or several functions and play a role in
Looking at some, apparently rather authoritative example in the wake of the genome sequencing project, (Snyder and Gerstein
In the beginning of modern molecular biology, genetically identified functions could be related first to polypeptides and then to DNA (Hershey and Chase Definition of the gene: a functional polypeptide basis of a unit function. By genetic analysis, the gene is identified as a phenotypic function. An individual function is based on co-operating proteins or polypeptides; the latter represent, hence, the basic unit functions. At nucleic acid levels, the closest equivalent is the coding sequence for such a polypeptide, inserted into the mRNA. In the general case, such a coding sequence - gene equivalent - is fragmented in the DNA, which constitutes the genotype, basis of a specific phenotype
While this was an important step, it turned out to be too simple. The reason is that whereas some polypeptides, like pancreatic RNase, assume a function by themselves, in most cases a genetically determined function is based on a higher order complex of polypeptides. Furthermore, these polypeptides typically may interact with low Mr compounds as Heme, vitamins, metal ions, etc. This means that several polypeptides or genes have to co-operate to secure a function. At that point, Jacob and Monod coined the notion of an “operon” constituted by several, possibly cooperating genes. Other problems then emerged with the discovery of regulatory genes. As an example, let us consider the lac repressor gene.
The
In molecular biology of eukaryotes, some researchers found the situation to be still more intricate and complicated. In bacteria, transcription and translation are tightly linked in a single physical complex. In eukaryotes, in contrast, the DNA is stored in the nucleus, which is the site of transcription, whereas the polyribosomes, where translation takes place, are located in the cytoplasm and thus removed from the DNA. The mRNA becomes autonomous, thus, and new types of controls become possible at that level. An untranslated region (5′-side UTR) preceding the coding sequence in the mRNA is needed to avoid a functional overload of the initial bases of the mRNA string. For both, chemical and steric reasons, the initial bases of the mRNA string cannot at the same time recognise and interact with the ribosome and bear the initiation triplet. However, there is also a 3′-side UTR at the end of the mRNA chain that, in the case of some genes (e.g., the Prion mRNA), can include more nucleotides than the coding sequence itself. These untranslated regions, being contiguous and in
Obviously, we can go on with problems and difficulties: In particular, mRNA form ribonucleoprotein complexes (mRNPs) in eukaryotic cells. More precisely, specific proteins recognise and attach to specific sequence motifs along the mRNA chain. This happens not only in the UTRs, but inside the coding sequence itself, for instance as shown in the case of globin mRNAs (Dubochet et al.
Even more devastating for the original gene concept is the existence in eukaryotes of giant precursor RNA and its gradual processing (Scherrer and Darnell
Under these circumstances, how are we to deal with this situation where no single term is adequate to capture all types of information involved in the expression of a single genetic function? Clearly, we need to distinguish and isolate the essential units of the process of gene expression, from both the mechanistic and logical point of view. Necessarily, this process will require new concepts and terms; we shall boldly enter this path. In the end, we shall not only find ourselves equipped with precise definitions for gene expression in terms of Molecular Biology, but we shall also be able to devise and apply mathematical algorithms that can analyse gene storage and expression in terms of information processing. A short version of the proposal to be presented here has been published recently (Scherrer and Jost
Genetic function is carried out by proteins composed of folded polypeptides. Their amino acid sequences are read off in the process of translation from the coding sequence contained in the mRNA. The mRNA coding sequence is the elementary counterpart of the biological function, and therefore constitutes the natural starting point for a gene definition wishing to capture biological function. This leads to Benzer’s original definition of the gene in terms of molecular biology, meaning the uninterrupted nucleic acid stretch that, as already mentioned above, was called “cistron” (Benzer
More specifically, to implement this program, we have both,
Let us list some of the many steps involved in gene expression, roughly in their temporal order (in fact, it is important to take a comprehensive view here): chromatin modification and activation, transcription and formation of pre-mRNPs, processing (including splicing) and transport of the pre-mRNP, formation and export of the mRNP to the cytoplasm, activation (or, perhaps more accurately, de-repression) of mRNA and, finally, translation. The
Actually, the terminology needs to proliferate a little at this point. To express that a polycistronic pre-mRNA and/or a full domain transcript (FDT) can control in
A “poly pre-genon” then controls more than one gene in case of a polycistronic pre-mRNA (several “fragmented” coding sequences in a row) or contains the fragments of several genes to be created by differential splicing. A “mono pre-genon” occurs when a single gene is contained in a genomic domain, or at later steps of processing of a polycistronic or polygenic pre-mRNA. This mono pre-genon accompanies the pre-mRNA of an individual gene. When an mRNA is produced by alternative splicing, the remaining elements of its pre-genon form its genon. The distinct genon in the mRNA eventually formed includes all
In view of the distinction between
It is a consequence of our concept that there are at least as many genes and genons as distinct open reading frames (ORFs) encoded in the genome. This means that in the human genome there would exist about 500,000 genes, each controlled by its own genon and producing a specific polypeptide (Scherrer and Jost
The genon and its precursors act at transcriptional and post-transcriptional levels and expire with mRNA translation and its eventual degradation once they have fulfilled their function. Therefore, the translation step is the natural cut-off point for our analysis. A complete picture should include all aspects of the control of gene products, of their types as well as of their numbers, that is, RNA and protein degradation as well as biogenesis and the interplay and coordination of biosynthesis and degradation. However, we shall not treat here the post-translational programs governing gene expression, nor the catabolic side of protein homeostasis.
To prepare the subsequent discussion of the genon concept, we now shall discuss the types of information involved in gene expression and the various types of gene products. According to our conceptual strategy, gene expression is governed by the coding sequence and the genon. The genon concept implies on one side the program in
The products of gene expression can be of protein or RNA nature; these products may carry out some structural or enzymatic function, or may control gene expression in a mechanistic or regulative manner. This suggests a two-fold distinction, between protein and RNA genes (P- vs. R-genes for short), and between structural and controlling genes (s-genes vs. c-genes). As these distinctions are independent, we thus have sP- and sR-genes as well as cP- and cR-genes.
It has to be kept in mind, however, that some types of gene products may act simultaneously in several of these categories, for instance as sP and cR genes (e.g., the SRA protein gene involved, as an RNA, in differential splicing, Hube et al.
By definition, “protein-gene” implies that the corresponding gene function is carried out by a protein, constituted by one or several polypeptides.
As outlined above, the coding sequence is the mRNA equivalent of the gene, being defined by genetic analysis carried out at the level of the
The unit of a coding sequence is the
On the other hand, given amino acids and hence triplets are not equivalent within the polypeptide chain. Indeed, the same type of amino acid may assume different “functions” at the level of the secondary protein structure, in terms of hydrogen-bonding or ionic interaction, once the polypeptide chain is folded in the 3D space. This type of function may hence be projected back onto the corresponding genomic sequence. To single out a given triplet, its position within a coding sequence and/or exon should be labelled. One may hence conceive a notation for the position of a given triplet within a coding sequence. A possible and efficient description, compatible with alignment schemes in bioinformatics, is the following: chromosome / genomic domain / maximal open reading frame / exon / triplet position within exon. (One should note, however, that there is as yet no generally agreed and universally employed convention in bioinformatics for describing the position of a triplet in the genome of a species.)
Accordingly, a given triplet (formally or as a physico-chemical entity) can be followed from the genomic DNA to a collection of amino acids within polypeptides. This is an essential feature when handling the triplet and its information content by a mathematical approach.
By definition, structural protein genes contribute to cellular structure and function either directly or via enzymatic activities. They may constitute the building blocks of the nuclear and plasmatic membranes, the endoplasmic reticulum, the nuclear matrix and the cytoskeleton. As enzymes they govern the intermediary metabolism as well as protein, RNA or lipid biosynthesis and degradation. There are the proteins acting as the mechanistic and enzymatic carriers of the system of protein biosynthesis, which do not discriminate among specific types of DNA, pre-mRNA or mRNA. Among the latter are, for example, the RNA polymerases, the non gene-specific splicing factors, the non-specific transport factors as “exportin” or NLS (nuclear localisation signal) (Rodriguez et al.
Regulatory protein genes control gene expression from chromatin activation to transcription and translation; they may function as repressors and activators of transcription, or act at post-transcriptional levels by interaction with pre-mRNA and mRNA. Four sets of such regulatory proteins can be distinguished: (1) the non-histone type chromatin proteins, as the transcription factors (TFs; see Latchman
The proteins binding nucleic acids at DNA and RNA levels, the non-histone chromatin proteins, the pre-mRNP and cytoplasmic mRNP proteins, all three constitute distinct populations of proteins including several hundreds and, possibly, up to thousand members in animal cells. Since they seem to act on specific genomic domains, and on RNPs including specific types of mRNA, it follows that they must act in a pleiotropic manner, constituting
Among these cP-gene products two types should be distinguished: (1) those which act on specific individual genons regulating, hence, the expression of specific genes, and (2) those which control the expression of whole sets of genes or gene families. Among the latter are, for instance, some types of transcription and translation factors.
By definition, “RNA-gene” implies that the corresponding gene function is directly carried out by an RNA, in association or not with proteins.
The most important member of this class of RNA is the ribosomal RNA (rRNA), which serves as the scaffold of ribosomal subunits by organising the sequential alignment of ribosomal proteins (Scheer and Hock
Among the RNAs intervening in control of gene expression we have to distinguish those which handle many types of (pre-)mRNA without individual selection, in contrast to those which selectively recognise and control, in a sequence-specific manner,
The object of this chapter is to point out that large parts of the genome relate to other mechanisms than gene expression
The nucleic acids carrying the genome and the gene expression machinery must assume at least two basic functions: (1) contain the information relating to the genes and, (2) serve as the physical support for this information. Function (1) is all evident within the definition of gene and genon developed above, whereas the implications of function (2) are less clear.
First, the support of genetic information has to obey the necessities of various quite distinct functions as (1) long-term storage of genetic information, (2) its transmission from generation to generation, and (3) the intra-cellular mechanisms of gene expression including selective transport of the transcripts to the sites of translation (see Fig.
It is often forgotten that both, DNA and RNA, act a priori as the mechanical support of genetic information and have to adapt to stringent rules deriving from their own physico-chemical properties. Concerning information storage and regulation at DNA level, an important factor coming into play is, for instance, the quite high physical rigidity of the DNA double strand which does not allow free and random movements, in particular in the conditions of high viscosity in the cellular nuclei. There are limits to folding up of hetero- and euchromatin and, e.g., to rapid “flip-flop” movements of DNA loops assumed to operate according to some popular models (de Laat and Grosveld
As suggested above, the genomic DNA may have an architectural function organising both, overall nuclear as well as local chromatin organisation. Cavalier-Smith (
Surprisingly neglected by actual Molecular Biology is the fact that DNA and RNA have to operate in a 3D space; and passive “crystallisation” or interaction of macromolecules cannot possibly explain all of genomic and cellular 3D organisation. (DNA “knows” that there is iron and light in the world, but seems to have “forgotten” that its environment is a 3D space !). When genes are being expressed, their reconstitution from RNA fragments in course of splicing, as well as the physical transport of mRNA from sites of transcription to those of expression, have to be organised in the 3D space and necessitates a precise dynamic architecture in space and time. Within these mechanisms, relating to semi-static and dynamic nuclear architecture, the positions of exons and the sites of RNA-protein interactions within the transcripts obey certain rules, which must be compatible with the selectivity of RNA processing and its implementation in the 3D space. Furthermore, since it became obvious that the genome is distributed in specific, experimentally identifiable sectors of the nuclear space, assigning specific positions to chromosomes and genomic domains, the organisation of the DNA
That the nuclear DNA might carry information other than that related to the genetic code could be inferred for a long time on the basis of data pointing to its possible role in cellular structure. The C-value paradox (Cavalier-Smith
That DNA may have a structural role independent of its gene content is also demonstrated by the phenomenon of the “petit” mutants in yeast (Bernardi
In chromosomes also, there are DNA segments which relate to structure rather than gene content. The genome is subdivided into genomic domains. The definition of genomic domains may be based either on the organisation of DNA, chromatin and/or chromosomes; or on functional considerations, such as units of replication or transcription. As pointed out in the “Cascade Regulation Hypothesis” (CRH; Fig.
There is, thus, good reason to consider the interbands of polytene chromosomes as borders of genomic domains. All the more since some molecular biological and biophysical facts point to the same interpretation. Interband DNA has some qualities of insulators, as defined by molecular genetics (Gaszner and Felsenfeld
The higher order organisation of DNA into genomic domains is embedded into the super-organisation of chromatin and chromosomes, which divide the genome into individual segments. Phenotypically very similar animals of closely related species may have vastly different numbers of chromosomes. Indeed, the fusion of the 46 telomeric chromosomes of
The Unified Matrix Hypothesis (UMH) was an early attempt to give a logical interpretation to the, apparently, surplus DNA, lightly qualified as “junk” (Ohno
A straightforward illustration of this proposition was the phenomenon of The
On this basis, the proposition was made within the UMH (Fig.
The main conceptual implication was that
This is not the place to further develop this theory; suffice to say that in recent years more and more relevant data could be placed within the originally loose frame of the UMH. The recent reports about “kissing chromosomes”, showing that distant chromosomal sites must be linked physically, to allow the expression of specific genes within “3D gene regulation”, is a most eloquent illustration of this basic concept (Kioussis
Here we need just to point out that there exist basic functions of DNA that are only indirectly related to gene expression. The UMH indicates disjunction of the actual genome size, which varies vastly within the C-value correlation, in particular in its repetitive elements, from gene expression. As pointed out above, in the same group of species with vastly varying DNA content, the sequence complexity of the expressed genome may remain almost constant (Rosbash et al.
Although the overall architectural function of DNA seems dissociated from the specific mechanisms of protein biosynthesis, an architectural function in gene expression of the transcripts as well became more and more evident. The observations of an RNA-dependant nuclear matrix (De Conto et al.
One may propose that the DNA defines the overall nuclear architecture per se and, in particular, the euchromatic part of chromatin which is unfolded and DNase-sensitive. The directly DNA-dependant 3D network is more “static” than the dynamic RNA-dependent architecture. It is liable to modification, however, in the process of cell differentiation, when the relative parts of hetero- and euchromatin are modified. The concept of “Quantal Mitosis” (see “Formation of differentiation-specific
On the other hand, there is the transcript-dependant, dynamic nuclear architecture as a result of RNA transcription, processing and transport. It is encoded in the (pre-)mRNA and its (pre-)genons. However, in both cases - the non-transcribed as well as the transcribed genome - the architectural function turns one-dimensional DNA and RNA into 3D structures, into which the coding parts are inserted. This conceptual deduction seems liable to explain to some extent the 95% of “surplus” DNA in a logical manner.
Another type of genetic information fixed by evolution into the genome without being directly involved in gene expression may be related to mechanisms termed, possibly,
Applied to the genon concept, this means that in the nucleic acid backbone, within the
Thus, merely mechanistic criteria of the information carriers and their higher order complexes must be respected as solidity, flexibility and folding characteristics, adapted chemical stability (
A particularly interesting illustration of such phenomena is meiotic recombination and sister chromatid exchange which imply the formation of the synaptonemal complex as the physical basis of meiotic crossing over (Colaiacovo
As defined above, the genon represents a regulatory program superimposed and attached to a given coding sequence. It is materialized in
The implementation of the genon-program in
We will restrict here discussion to the
The
Within the Cascade of Regulation, specific gene expression in a given eukaryotic cell may be subdivided in (at least) the following steps: Organisation of the DNA in the 3D-space and formation of the DNA-dependant matrix (step 1 in Fig. Organisation of chromatin into chromosomal territories. Formation of differentiation-specific local chromatin networks and the DNA-derived nuclear matrix. Activation of chromatin domains for eventual transcription of individual transcriptional units contained in a domain (step 2 in Fig. The primary transcripts (step 3 in Fig. Synthesis of the FDT or individual primary pre-mRNA. Association of nuclear RNA-binding proteins to pre-mRNA forming the pre-mRNPs. Formation of the RNA-derived nuclear matrix by integration of the pre-mRNPs. Processing of pre-mRNPs (step 4 in Fig. Differential splicing and formation of the pre-mRNP including exons of a single coding sequence (step 5 in Fig. Final processing of pre-mRNPs (step 6 in Fig. Import of mRNA into the cytoplasm (step 7 in Fig. Formation of cytoplasmic inactive (ribosome-)free mRNP with concomitant replacement of the majority of nuclear (pre-)RNP-type proteins by cytoplasmic ones (step 8 in Fig. Activation of mRNA and polyribosome formation (step 9 in Fig. Replacement of mRNP proteins by translation factors, forming the translated mRNPs. Formation of polyribosomes by association of 40 S and 60 S (native) ribosomal subunits forming functional ribosomes. Translation of the coding sequence in mRNA (step 10 in Fig. Formation of the nascent primary polypeptide and secondary protein structure (the genon has expired).
In addition, at several steps of biochemical information processing RNA interference (RNAi) takes place, in the nucleus as well as in the cytoplasm, by physical elimination or temporary masking of mRNA sequences by siRNAs or miRNAs. Another important but not clearly localised mechanism of information processing is RNA editing, by which a coding sequence in an already present (pre-)mRNA can be modified (review in Koslowsky
Mechanisms of expression and regulation within the cascade operate mainly by association of regulatory proteins and of interfering RNAs, and by the action of the enzymes involved in the transcription and processing machinery, including control of RNA editing. The physical support of the carriers of information is the nuclear matrix and the cytoskeleton, as well as the endoplasmic reticulum for proteins to be exported.
The known biochemical steps of DNA and RNA activation, of RNA processing and transport occur within the “Cascade of Regulation” which stepwise reduces the information content of the genome to that of a single gene, ultimately. It shall be pointed out, however, that in terms of information processing, information is gained during this process, to the extent that uncertainty about the eventual selection of a given triplet in the DNA, to be expressed within a polypeptide, is gradually reduced. The potential information of the genome thus becomes effective.
The content in genomic information is currently evaluated in terms of what in the biological literature has been called “sequence complexity”, that is the length of non-repetitive DNA or RNA (Britten and Kohne
More precisely, there are three main reasons for stepwise regulation of gene expression: Noise As pointed out in the Cascade Regulation Hypothesis (CRH), published first in 1968 (Scherrer and Marcaud Effort There exist different search strategies that, in principle, could be employed for the selection. If one performs the selection in a single step, one needs to screen all the available elements to find the right one. The selection effort is then proportional to the number of items to be scanned, which is in case of the human genome, of the order 106 or 107. As explained above, this effort is far too large to be biologically realistic. The other extreme is search by binary alternatives. Here, in the first step, the set of items to be searched is divided into two classes of equal size, and one selects one of those. In the next step, that class is again divided into two classes and the process is repeated until after log Within gene expression, selection effort means mainly the number of regulative factors needed within the transgenon. To keep their number in the genome low - for obvious reasons - cP-regulators have to operate in Reaction speed Gene expression is a long and complex process. When physiological adaptation must be rapid, the necessary information may not possibly be called from the genome: gene information must be stored close to the place of action, in the extreme case in form of pre-proteins as, e.g., trypsinogen, turning into a functional enzyme upon a simple biochemical signal. We have introduced the term “peripheral memories” (see Fig.
The expressed part of the genome can be measured by modern micro-array techniques, which give numbers of genes represented in a given cell isolate, and from which the non-repetitive sequence length, in terms of (known) RNA-sequence, might be calculated. However, such data are at present not available in a comprehensive manner, in relation to the biochemical steps of the gene expression cascade. We have therefore to rely on the published data of sequence length (“sequence complexity”) measured by re-association kinetics in hybridisation assays which are expressed as Cot- (for DNA - Britten and Kohne
Early hybridisation data indicated that 10–20% of the nuclear DNA is transcribed in most species, even in highly specialised cells as the red blood cell, where 90% of the protein output is globin (Imaizumi-Scherrer et al.
In the following, we will discuss the individual steps of the cascade of regulation in view of the genon concept.
The first step of the regulation cascade involves the selection of the chromatin fraction to be eventually activated in a given cell. The zygotic genome is being subdivided into stem cell lines according to the mechanisms of (lineage) determination (review in Tiedemann et al.
In this step, facultative heterochromatin may be transformed into euchromatin; but not all euchromatic DNA is by necessity transcribed, eventually. The classical criterion for chromatin liable to be activated is its DNase sensitivity (Razin et al.
Present consensus assumes the intervention of transcription factors (TFs) and promoters, which might render the DNA liable for transcription. Recently, however, data supporting other types of interpretation appeared. Transcription factors, for instance GATA-protein binding sites, are spread all along a genomic domain of, e.g., the human or chicken globin domain (Cantor and Orkin
At this step of the regulation cascade, the holo-transgenon has to provide for the regulatory proteins which interact with the DNA at specific sites provided by the proto-genon in
In this regulative step, the information of the proto-genon of a genomic domain is reduced to that of the pre-genon, which is carried along by the RNA. The primary transcripts, which may include fragments of several genes, carry the
These processes are controlled by the factors constituting the holo-transgenon of a given nucleus: presence or absence of specific TFs and of factors involved later-on in gene-specific processing (splicing) and transport, decides the fate of a given transcript in time and space.
Factors carried over from the DNA. According to recent data, these include promotor binding factors (Auboeuf et al. The “classical” pre-mRNP (also called HnRNP) proteins of relatively basic charge (pI), the “histones” of the pre-mRNP; there are less than 50 components known (Dreyfuss et al. The acidic pre-mRNP proteins (Maundrell et al. The ambivalent prosomes (Schmid et al.
Present in the nuclear sap or in specific compartments (e.g., the so-called “speckles” Handwerger and Gall
As pointed out above, RNA processing and transport can be interrupted at several metabolic steps, resulting in the constitution of “peripheral memories” (see Fig.
Processing of pre-mRNA represents the major regulative process in gene expression. Indeed, transcribed gene fragments are either temporarily stored in the nucleus or degraded, or else selected for productive splicing and transport of mRNA to the nuclear membrane. Indeed, during this process, 90% of transcribed sequence is eliminated either transiently or permanently (Kiss
Overall RNA processing occurs in steps. Early data showed that there are discrete steps in terms of size of the transcripts, RNA turnover times and sequence complexity. For example, in avian erythroblasts, RNA of very high Mr (among them globin RNA of up to 33 Kb) with half-lives (
In positive correlation with these old findings on global RNA processing, recent in situ hybridisation data indicate that primary globin transcripts occupy diffuse, not clearly defined sites in the nucleoplasm, that a large part of the transcripts accumulate around the nucleoli when RNA processing is interrupted, whereas productively processed and exported globin (pre-)mRNA form two distinct processing centres (PCs). The highly unstable primary transcripts would hence end up in the PCs, where intermediary products of globin pre-RNA processing accumulate and transport to the cytoplasm starts (Fig.
A most important feature of RNA processing concerns the nuclear matrix. As outlined above (“Formation of the RNA-derived nuclear matrix by integration of the pre-mRNPs” section), the primary transcripts constitute the backbone of the RNA-derived nuclear matrix (see, Ioudinkova et al.
During processing, eventually a pre-mRNA containing the exons of a single gene is formed containing, hence, a unique pre-genon. The most decisive mechanism operating at this step is differential splicing (Blencowe
The system controlling this process is once more the holo-transgenon of a given nucleus, which is modified according to cellular differentiation, during embryogenesis as well as in terminal differentiation. Concerned are the ubiquitous or partly selective factors and enzymes involved in splicing, among them the U-type small RNAs, resp. their RNP complexes. Less well known are the factors which govern the putative gene-specific splicing.
In parallel with pre-mRNP processing, part of the RNA is degraded. There is elimination of introns and intergenic RNA as the basic mechanism of processing. However, there is also elimination of part of the exonic and other functional RNA as a selective process under the control of the transgenon; RNA interference may also play a role at this level (Matzke and Birchler
The final pre-mRNA is transformed into mRNA with its unique genon, ready to be exported to the cytoplasm. Accordingly, factors constituting a specific transgenon are by now associated with the mRNA. Final processing may be concomitant with export; e.g., the last intron of globin pre-mRNA is eliminated just prior to export. (In the general case, the nucleus does not contain mRNAs, and the cytoplasm no pre-mRNA). Though it is not clear at present if final processing entails by necessity export of the mRNA to the cytoplasm, nevertheless, a final selection step at this level has to be taken into consideration.
From the nuclear processing centres (PCs), mRNA is exported to distinct sites in the cytoplasm prior to being dispersed (see Fig.
It is possible, although actually not established, that import of mRNA operates in a gene-specific manner. At this crucial step of the cascade - in view of the threshold of the nuclear membrane - qualitative selection and hence reduction of gene- and genon-specific information might operate.
The machinery of mRNA import is concentrated in the nuclear pore complex (Maco et al.
The final mRNAs entering the cytoplasm carry their unique genons which are exposed to the cytoplasmic holo-transgenon, allowing them to pick up sets of factors corresponding to their individual genons, respective transgenons. This process results in an almost total exchange of mRNA associated proteins relative to those of the nuclear pre-mRNPs. A notable exception are the already mentioned factors involved in mRNA exportation which shuttle between both compartments. Furthermore, the prosomes are found on both, nuclear pre-mRNPs and cytoplasmic silent mRNPs.
The holo-transgenon as defined by proteomic analysis of silent mRNP complexes includes several hundred proteins, in their majority of rather acidic pI. The composition of factors in a given cellular compartment is in constant change in function of physiological adaptation, controlled by internal agents as well as by factors from the environment. The proteins directly attached to silent mRNAs act as genuine cytoplasmic repressors (Civelli et al.
The advent of RNA interference has given a new dimension to cytoplasmic repression (Jackson and Standart
Within the genon concept, mRNA activation is controlled by the factors available within the holo-transgenon of a given cytoplasm. There are, in competition, the selective repressive factors of the silent mRNP on the one side, and on the other the rather ubiquitous translation initiation and elongation factors associated to the translated mRNA. The existence of a third class of putative factors might be postulated on theoretical grounds; those selecting individual mRNAs to change their repressed or active status.
Many facts indicate that translation per se is a compulsory, constitutive mechanism. Translation factors are ubiquitous and present in relatively high concentrations in the cell sap, whereas the proteins associated to the repressed mRNP, as well as prosome subunits, are only found within the complexes and not in free form (Maundrell et al.
Once translation has started, little regulatory intervention occurs in steady state that might involve genon and transgenon. Translation initiation is more temperature-sensitive than elongation; in less than optimal physiological conditions, ribosomes run off (Chezzi et al.
During translation, the rules of the genetic code and the translation machinery prevail by selection of triplets and assembly of the polypeptides. In steady state, the genon is hence put to rest as far as the coding sequence is concerned. In contrast, the 5′-side and 3′-side UTRs are likely to play a role by interacting with regulating proteins and interfering RNAs. Interestingly, polyribosomes have a tendency to form circles (e.g., Christensen et al.
Once the polypeptide has formed the genon, by definition, expires and the factors of the protein world modulate the nascent polypeptide to assume secondary, tertiary and quaternary structure, which, eventually, will assume the genetic function based on one or a set of cooperating genes. These post-translational processes of gene expression and its control have to be most complex; we will here not enter these matters. However, it seems important to point out, that gene expression must obey homeostasis of protein biosynthesis and degradation. Mechanisms coordinating protein biosynthesis and catabolism must exist, by necessity.
The main operator in clearing misfolded, or otherwise defect polypeptides is the Ubiquitin-proteasome system (Coux et al.
Forming a molecular cylinder, the prosome has the capacity to interact bi-functionally at either end, as can be directly observed in stress fibres of the cytoskeleton (Arcangeletti et al.
For this discussion, we exclude all mechanisms directly related to constitutive and basic protein biosynthesis within the frame of the genetic code as, for instance, the ribosomes and the basic tRNA machinery.
The
By definition, the
Regulation of transcription, and hence of programs of differentiation and physiological change, is in part under the influence of cell-external signals (see the “Exo-cascade” formulated in Fig.
The genon is embedded in the pool of
The transgenon, carried by cP-genes and cR-genes, is built up by the normal mechanisms of gene expression and regulation, leading to the synthesis of DNA- and RNA-binding proteins, the synthesis of siRNA and miRNA within the frame of RNA interference and of all other types of cR-genes which might affect differential regulation of gene expression.
Proteins cover all types of RNA in the cell. In case of mRNA and pre-mRNA, it was shown at an early time by electron microscopy that proteins are aligned all along the RNA molecules (Dubochet et al.
RNP-type proteins bind in a RNA-sequence dependant manner. The poly(A)-binding proteins (PABPs), attached to the 3′-side tail (length: 50–200 A residues) of the mRNA protect about 12–20 A-residues at a time (Baer and Kornberg
Early proteomic studies of 20 years ago allowed people to estimate that there are several hundred (up to 1,000) acidic, non-histone proteins attached to DNA, and as many to pre-mRNA and FDTs (Maundrell and Scherrer
These observations indicate that there must be a “code” governing the interaction of a limited number of NABPs in chromatin and mRNPs which, in general, are specifically DNA- or RNA-binding proteins. Relatively new data have confirmed, however, the old observation that the same protein may bind both, DNA and RNA, as outlined above. This was originally observed for the large T-antigen of SV 40 and polyoma virus (Darlix et al.
In addition to the signals encoded in the oligomotifs of the primary RNA sequence, there are post-transcriptional modifications of the transcripts (review in Shatkin and Manley
The transgenons carried by cP-genes are constituted by the normal mechanisms of gene expression and regulation by protein biosynthesis.
The second mechanism - recently discovered - of transient or final repression of specific mRNAs is RNA interference. Si- and miRNAs might block mRNA upon import to the cytoplasm, or during translation when mRNA segments become accessible as pointed out above. The phenomenon of RNA interference is at present most actively investigated and no general conclusions seem possible as yet. Actually, little could be said with any chance of precision, beyond the general considerations outlined above (see “Discriminating RNA regulators: siRNA and miRNA” section).
It is, however, evident that RNA interference represents at the same time a highly gene-specific system of control, liable to recognise precise RNA targets. It is hence at the same time more efficient but also less sophisticated than the regulatory protein factors. Indeed, the latter are capable to integrate controls to a much higher extent. The si- and miRNAs may represent primitive slots operating in an on/off mode. But this system as well has to be managed upstream by protein factors, not only enzymatic system involved in its generation, but also mRNP proteins. Being single as well as double stranded, interfering RNA is a target for any type of (pre-)mRNA binding protein as well.
RNA interference is likely to have evolved prior to RNA-binding proteins, possible already in pre-biotic systems. RNA hybridisation is the most basic process of RNA stabilisation and neutralisation. Later, chemically more sophisticated systems of RNA-protein recognition and mutual stabilisation may have evolved, much before the tRNA based protein-coding revolution happened, opening the gate to life and evolution.
The proposition of the Genon concept is not only thought to redefine the gene in unambiguous terms and allow better comprehension of gene expression and regulation; the ultimate goal is to provide a scheme clear enough to allow us the application of mathematical methods in analysis of genes and genomes. Here again we have to separate the definition of the gene per se from the programs that guide their expression in time and space.
The restriction of the definition of a gene to the coding sequence, constituted by the assembly of coding triplets, considerably facilitates the development of algorithms in view of mathematical analysis; as we will see below, the approach to be taken seems quite straightforward. It is evident, however, that the gene as a function represents more than the coding sequence and its equivalent in terms of the nascent polypeptide. Chemical modifications and the formation of secondary, tertiary and quaternary protein structure are not exclusively encoded in the primary amino-acid sequence; external factors as well govern the assembly of the structures underlying the functions expressed within the phenotype. Therefore, additional programs must exist which control this process; some programs may entirely or largely be encoded in a given genome but in addition, factors from the ecosystem seem to play a major role in the final gene function.
In line with our general conceptual decision of taking translation as the cut-off point, we here only take into account pre-translational processes and restrict gene expression to the formation of the primary protein structure. Nevertheless, the analysis to be presented can in principle also be extended to post-translational events.
In addition to the gene per se just mentioned, our information theoretical analysis will be concerned with the program of gene expression, i.e., with the genon. Again, it is natural to begin with the program in
The scope of this task can be seen by an overview of the types of decision-making programs that will come into play, following the mechanisms of gene expression exposed above.
The first program in
Formulating algorithms of control one has to take into account, furthermore, the fact that some decisions are made at DNA level which are born out at pre-mRNA level only; indeed, some proteins binding to specific DNA sequences are carried over to the pre-mRNA. Most often not taken into consideration, this is an important basic mechanism, which makes possible the sequence-related assembly of proteins with high affinity for DNA, and hence binding specificity, which, once assembled, are carried over to the RNA in
The analysis of the
The first step of that classification distinguishes factors produced by the genome itself from factors provided by the environment. The genome produces DNA and RNA-binding proteins as well as the small RNAs involved in RNA processing and the recently discovered RNA interference (RNAi). External factors provided by the environment include mineral ions, chemical compounds not produced internally (as some vitamins), diverse sources of energy, light (as a source for photosynthesis or as a signal for circadian rhythms), gravity (providing for example a gradient for spatial diffusion according to weight), etc. In between these two types of factors are the ones produced by other cells in a multicellular organism, like hormones, cytokines, and other secondary cell messengers. Here, for simplicity, we shall concentrate on genome-dependent
On one hand, we have to take into account those factors that physically interact with the
In order to appreciate the logic of the formal analysis, it is advantageous to start with biological simplifications, and approximate biologically realistic scenarios only gradually. In this regard, one might hence start with a single genon in a given mRNA that has available all possible trans-factors occurring within the holo-transgenon of a given genome. The genon in the mRNA then only needs to select the appropriate
The general question to be asked in terms of information theory concerns the information content, at the various and subsequent levels of gene storage and expression, of a gene as a product as well as the result of the expression program that led to its eventual realisation. Standard analysis is concerned with the amount of information about the biochemical identity of a polypeptide contained in its coding sequence. That, however, takes such a polypeptide out of its cellular context. First of all, a polypeptide is not simply read off from a coding sequence stored somewhere in the DNA, but, as we have amply explained, it is the result of an intricate regulation process leading to the coding sequence at mRNA level prior to translation. This involves contributions from regulatory elements in
As we will see, different formalism will apply to the “forward” and the “backward” analysis in terms of input from the genome, or from the exo-system, the latter bearing essentially on the holo-transgenons (excluding input in the frame of evolution). For any such analysis, it is essential to specify what one assumes as known and what one wants to know.
A clear-cut illustration of this problem is the number of different polypeptides imaginable within the rules of the genetic code: there are 64 triplets (−2: the start and stop codons) coding for 20 amino acids. Assuming average length of a polypeptide of, say, 500 amino acids, the number of all combinatorial possibilities is astronomically large, much beyond any range that evolution could have possibly explored. There are essentially two ways out of this impasse: to assume rules of possible sequence correlations or else, to put into the game the proteome as derivable from the sequence analysis of genomes published. Practically, these approaches have their limits since, in both cases, our knowledge is approximate, at best. Therefore, we will have to resort to experimentally founded assumptions to carry out this analysis. Concerning the human genome, e.g., we may hence assume the existence of about 500.000 different polypeptides to be potentially expressed, and up to one million gene products altogether, counting sR and cR genes and taking into consideration RNAi.
Entering our analysis, for a polypeptide actually expressed in a cell we can ask about the sequence at DNA or RNA level that is coding for it; this is the classical application of information theory to molecular biology. It deals with the selection of a given gene and leads to the issue of the degeneracy of the genetic code. Another aspect of this question is the localization of that coding sequence in the DNA.
Our main interest here, however, concerns the opposite direction, that is, going forward from a (piece of ) coding sequence in the DNA to the polypeptides (or other functional products) that it will eventually get expressed in. For such a coding sequence at DNA level, we already know the amino acid that each triplet is coding for. Looking only at this sequence we do not know, however, whether, and if so, when, where, and in which quantity that sequence is expressed in the cell under consideration. Thus, there is some uncertainty here, and we shall be concerned with quantifying that uncertainty.
In order to perform this quantification according to the rules of information theory, we need to specify the options available. Thus, we need to list those polypeptides in which our sequence could possibly be expressed. (Of course, in a particular situation at hand, it may not get expressed at all; this is one of the possible options.) It is now important to realize that there are some choices to be made here; we have to agree about what prior knowledge we already wish to admit. If we do not wish to admit any prior knowledge, we need to consider all combinatorially possible amino acid sequences (up to some specified length). As already pointed out, this is a very large number. We may, therefore, wish to impose some restrictions, in order to reduce the number of options and to include only more realistic ones according to the given cellular condition. We could restrict ourselves to consider only the amino acid sequences that are biochemically possible in the sense that they can give rise to well folded proteins, or to polypeptides that have been identified in some cell and are listed in some data base. We could even assume more prior knowledge, namely that we consider only those polypeptides that occur in the proteome of the organism in question. Or, finally, we could restrict our considerations to the ensemble of polypeptides present in the cell at the time of investigation. In any case, whichever of those ensembles we choose, the uncertainty then consists in identifying which member of the ensemble in question is realized by the expressed coding sequence, and also in which quantity. If there were no regulatory mechanisms like alternative splicing, silencing, or other decisions on the expression pathway, the expressed product itself would be completely specified by the (fragmented) coding sequence at DNA level. Still, however, the number of expressed copies is not yet determined. Repression mechanisms at various stages of the expression pathway could result in no expression at all, whereas repeated transcription/translation or other multiplicative steps could result in multiple products. Finally, mechanisms like alternative splicing even make it impossible to predict the biochemical identity of the expressed product from the (fragmented) coding sequence alone.
In the sequel, we shall set up the information theoretic scheme to quantify these uncertainties and to assess the relative contributions of the coding part, the gene, and the regulatory part, the genon, in resolving these uncertainties. Thus, the total information, needed to specify the types and numbers of functional products produced from a giving coding region at DNA level, is a sum that will be decomposed in the parts attributed to the gene and the genon. Numerical estimates (to be presented elsewhere in detail) will show that the by far larger part is the one coming from the coding sequence, whereas the contribution of the genon is rather small. As genon and transgenon are rather complex, involving many binding sites in
For the purpose of applying information theory to gene expression, we should first discuss the concept of information itself. Our starting point will be the theory of Shannon. In that theory, a sender composes a message from the elements of a code agreed upon with the receiver. The receiver knows the probabilities
The formula (
In our applications to molecular biology, we shall be concerned with sequences (of nucleotides or amino acids). For such a sequence, we want to know its composition, that is, we want to know which element (nucleotide or amino acid , resp.) occurs at each position. This is the information we are after. For formalizing this, there exist two alternative approaches, and in this section, we want to discuss those. One approach consists in simply taking the set of all possible sequences under the given circumstances as an ensemble and then quantify how much information is needed to specify a particular sequence within this ensemble. The other approach looks at the individual positions in the sequence in turn and quantifies how much information is needed to specify which nucleotide or amino acid occurs at that particular position. When we do this for each position and take correlations between the various positions into account, we can again quantify the information needed to determine the composition of our sequence. We shall now describe these two approaches in more formal terms.
Suppose that we are given an ensemble of
This is the maximal possible value of observations of relative frequencies, restriction of the ensemble, or encoding of regularities, or physical considerations, where we have some kind of an energy function, in the terminology of statistical physics a Hamiltonian
In molecular biology, we are not working with arbitrary ensembles, but often with ensembles of sequences, and for such ensembles, there is an alternative approach to entropy. Let
Here, without further knowledge, all the
Since there are unequal distribution of the sequence correlations leading to the consideration of block entropies One should note, however, that the computation of block entropies is numerically feasible only for relative small values of the block length
Ensemble and sequence entropy represent two different ways of computing the same quantity, and they should therefore yield the same value. Estimates for these quantity, however, can be different, because they will employ different aspects. Thus, whereas in the case of uniform probabilities, the values (
The application of information theory to molecular biology has been controversial. To clarify the issue, the following point might be useful. Usually, information theory is applied to messages. A message contains information when before receiving it one does not know the sequence of symbols in the message, that is, once the message is known that previous uncertainty is reduced. Shannon’s information measure quantifies that reduction of uncertainty, that is, the difference in knowledge before and after receiving the message. This suggests that, likewise, a stretch of DNA contains information about polypeptides or phenotypic properties because knowing that DNA sequence allows one to deduce the composition of those polypeptides or those phenotypic properties. Of course, because of the intervention of other factors, the knowledge of the DNA does not lead to complete knowledge of the relevant polypeptides or phenotypes. The point is, however, that knowing the DNA reduces the uncertainty about those polypeptides or phenotypes, and this then leads to a quantification of the information contained in the DNA. The remaining uncertainty then is assigned to other factors, and the corresponding information can then also be quantified.
The point we are emphasizing here in order to avoid misconceptions about genetic information [see e.g. Stegmann (
In information theory, the message from the sender to the receiver has to pass through a channel, and the latter may not faithfully transmit everything emitted by the sender. The channel may introduce noise, that is, random distortions or modifications of the message. Also, there may be systematic effects decreasing the information content of the message. Different messages may be received as the same message. This is called redundancy. Redundancy can have the positive effect of error tolerance, in the context of a triplet coding for an amino acid meaning that certain mutations do not affect the amino acid coded for. Indeed, for the receiver, it does not matter which of those different messages have been chosen by the sender as long as the received message remains the same. Thus, the sender can make some errors as long as they do not change the message for the receiver.
In particular, the genetic code is redundant in the sense that the genome as the sender emits nucleotide triplets while the proteome as the receiver obtains amino acids, and several triplets of different chemical composition lead to the same amino acid.
The application of information theory to molecular biology, however, should go beyond the relationship between individual nucleotide triplets and amino acids. A nucleotide and an amino acid not only have their specific chemical identity, but they are also parts of sequences, the DNA sequence, or a polypeptide chain constituting (part of) a protein. As such, in addition to their chemical composition, they are characterized by their position within that specific sequence. Moreover, the relationship between such a triplet in a specific position and the amino acids coded for by that triplet is not a relationship between individual physical objects, insofar as in a given cell, the triplet is usually expressed several times, and in different polypeptides. Each amino acid produced from the triplet can be considered as a physical instantiation of this particular triplet, and of no other triplet. There are many chemically identical triplets in the DNA, but the given amino acid as a concrete physical object is derived from precisely one such triplet.
Considering it that way, however, falls short of understanding the expression process, and if that were all that information theory can contribute, its usefulness would be rather limited. While in principle we can follow a specific expression pathway and trace the origin of a given amino acid back to a single triplet at its location in the DNA, the formation of that amino acid requires additional ingredients along the expression pathway. Some ingredients come from the cis DNA region containing that triplet. For instance the nucleotide sequence encoded in a promoter region is also needed, and enhancer and repressor sequences affect the expression. Factors in trans, which are specific for the intra- and extracellular environment, also guide the expression. The point in time within the processing sequence also affects the outcome. Thus, the relationship between specific individual chemical units is superseded by processing information that does not implement itself physically in the final product. So, on one hand, when tracing the process back in time, we have a relationship between individual chemical substances determined by their locations within specific sequences, while on the other hand, when going forward in time, we have the combination of cis and trans ingredients determining in which and in how many numbers of polypeptides a given triplet is expressed.
We have quantified the types and numbers of polypeptides derived from a given coding region (genomic domain) by the second term in (
Within the total protein content of a cell, we consider the ensemble of amino acids derived from the given triplet in the DNA. Whereas the chemical structure of these amino acids is the same, their number, that is, the number of copies derived from the same triplet, may vary. In addition, due to differential regulatory effects on the expression pathway, for instance differential splicing, these amino acids may find themselves in structurally different polypeptide chains. The corresponding types we identify by the index
Before listing some possibilities for quantifying that information content, we recall a general observation from our above discussion of the entropy: When we have a collection of physical objects, we can either list them as such, or we can seek regularities, for example identify types represented by several individual objects, to achieve a more compact representation. In the sequel, we shall begin with the naive list and then proceed to a representation in terms of types Explicit description of all physically present polypeptides in the cell containing an amino acid derived from the triplet under consideration. When no further regularities are taken into account, this becomes The preceding used the class of all possible types of polypeptides. This class, however, is too large to distinguish between the different information contributions. For determining the contribution of the protogenon, we should use the class of all polypeptide chains that can be produced from the same coding sequence in the DNA, under a set of specified trans conditions. Likewise, at the level of the pre-mRNA, the possibilities are already more reduced, and the selection between them is now governed by the pregenon. At the level of the mRNA, it is then the genon that is responsible for selective gene expression. Since the same type of information theoretic analysis can be applied at each level, we shall now discuss the protogenon. The pregenon and the genon then can be handled analogously, by simply replacing the different coding regions in the DNA eventually contributing to the final product by those present in the unprocessed pre-mRNA or the unique one in the mRNA. So, we return to the ensemble of products that can be derived from a given coding region in the DNA. Each type
We can perform the same type of analysis for larger cis regions than triplets, for example for DNA domains containing fragments of coding sequences or ORFs. The information measures will differ when we have overlapping ORFs, that is, when one triplet belongs to several ORFs as in the case of alternative splicing or other forms of differential processing. In that case, the uncertainty about the products derived from the triplet needs to take the uncertainties about the final products about all those ORFs into account.
In particular, we can then compare the information provided by different DNA domains and thereby specify the information content of the protogenon. Let us consider a sequence
We are now in a position to assess the information content of the protogenon. Here, we take as
When we wish to analyze a specific transgenon and its information contribution, then, instead of adding some further cis elements to the original sequence
In any case, when
Before proceeding, let us briefly make the following remark: whereas here we have considered the situation for
There is a different, but somewhat coarser, method of estimating the information provided by the (proto-, pre-) genon. To see this, we consider a polypeptide and look again at the case of the protogenon; we shall ask about all the DNA sites that contributed to its formation, that is, both, the coding triplets and the ones from regulatory regions that guide the process leading to that polypeptide. In "The coding region" above, we have already studied the sequence informations for the polypeptide and the corresponding coding region in the DNA. By the same method, we can then also evaluate the sequence information of non-coding regulatory sites, both in cis and in trans, i.e., either present in the cisgenon or provided by factors from the transgenon. The former include stop codons, enhancer, promoter, repressor sites, introns that play a role in the expression pathway as binding sites for certain proteins, and the like. The relevant part of the holo-transgenon derives from the coding regions for transcription factors and all other proteins regulating or interfering with the expression pathway.
There is a problem with this approach, however. The reason is that many of the regulatory elements, while being specific to a certain degree, need not only affect the polypeptide under consideration, but also interact with the regulation of other polypeptides. Therefore, the simple sum over the sequence entropy of all contributing sites seems to overestimate their specific information content. Putting it another way, when we consider two different polypeptides, we are not allowed to simply add the corresponding sequence entropies because some of the factors may contribute to both of them, leading to an overestimate for the information needed for the two polypeptides. Nevertheless, this approach might be useful in deriving some upper bounds for the information needed to produce a polypeptide.
In this section, we want to investigate the information theoretic aspects of the genon, accompanying the potential gene on the expression pathway, from a different point of view. For that purpose, we shall analyze the relative contribution of cis signals and trans factors to the information needed to express a specific gene. The basic situation is that the cis region provides certain control signals, like enhancers at the DNA stage or binding sites for proteins forming RNP complexes at the RNA stage, whereas those binding factors then constitute the transgenon.
The contribution of the cis region with its combination of binding oligomotifs consists in a preselection of the possible binding elements at the particular site under consideration, out of all the proteins in the cell that can bind to DNA or RNA. We first consider one particular site
The basic case from which to start thinking about the genon is where the whole expression pathway is solely controlled by cis, in the sense that all necessary factors are provided by the program represented by the transgenon, and any specificity is entirely due to selection by cis of binding factors. In that case, all
Still, this needs to be expanded in two directions. First, a cis region can, and typically does, contain more than one protein binding site. When the binding properties of these sites are independent of each other, we can simply sum the expression given in (
In other situations, we need to modify this expression by taking correlations into account as in the previous sections. Second, there is an important combinatorial aspect because at one site, usually not a single polypeptide is binding, but some combination of such polypeptides that then biochemically form a quartenary protein complex. Furthermore, some other proteins facilitate or inhibit the binding of certain other ones. Therefore, instead of single proteins, we need to consider protein combinations, as in a language where instead of individual phonems or letters, one rather takes morphems or words as basic elements. The principle expressed in (
In summary, the process information content of a cis region is quantified by a reduction of possibilities. Therefore, it cannot be computed directly from the nucleotides forming the region, but rather depends on the proteome in the cell. This may seem paradoxical, namely that we cannot compute the information contribution of a DNA region by looking at the nucleotides, but rather need to compare the number of possible binding proteins with the smaller number of those actually capable of binding to that particular region. Of course, it is determined by the latter’s nucleotides which proteins can bind there and which ones can’t, but in order to do the computation we need to know which trans factors are the candidates.
Conversely, the information contribution coming from trans simply consists in the selection of those factors that actually bind to a given (proto/pre)genon, out of those possibilities allowed by the structure of the signals in the DNA domain as composed by its nucleotides. Thus, here the difference is between those that can possibly bind, given the concrete nucleotides, and those that are actually provided by the holo-transgenon of the given cell. Returning to (20), the uncertainty left after evaluating the information provided by cis is the term −∑
The crucial entropy (14)
The case of the genon seems different. First of all, in our computations of information, we have ignored an important aspect of the contribution of the genon. The genon not only decides what is produced among the alternatives provided by the coding sequence at DNA level, and in which quantities, but also at which place in the cell and at what time, within development and differentiation and the cell cycle, it is produced. In principle, one could also quantify this in information theoretic terms. For that, one would need to identify the spatial and temporal scale at which significant differences within the cell and its life occur. Another explanation for the apparent information loss concerning the genon can be offered in terms of Ashby’s law of requisite variety (Ashby
In this paper, we have developed a definition of the gene that conceptually separates the gene as a product, from the genetic information relating to the regulation of gene expression, the latter being defined within the genon concept (Scherrer and Jost
This emphasis distinguishes our approach both from DNA sequence based definitions in the wake of the human genome sequencing project that lost the functional aspect out of sight, and from more recent definitions that are motivated by the ENCODE project (ENCODE Project Consortium
To put it differently: Since there are two distinct aspects involved in the production of a collection of polypeptides from coding fragments in the DNA, namely translation of triplets into amino acids, and regulation of the assembly of those sequences of triplets from the initiation of transcription to the final mRNA prior to translation, two distinct concepts are needed. One is the gene that then becomes freed from all ballast and can again assume a pure role of a functional unit, and the other is the genon that guides and controls the assembly of the gene through the steps of the expression process.
Let us try to put our conceptual framework into perspective. Our information theoretical analysis is entirely sequential, as it is motivated by the principle of the cascade of regulation, and it integrates well a substantial body of biochemical knowledge and theoretical concepts accumulated about genome organization and gene expression (cf. Scherrer
In any case, it seems that a conceptual and information theoretical discussion of the gene has its natural point of termination at the stage just prior to translation when the coding information is read off from the mRNA, a limit adopted within this essay. After the sequence identity of a polypeptide has been determined, physical and biochemical processes take over to determine the shape in 3D of proteins as well as their spatial localization and co-localization within the cell. This then constitutes the basis of the metabolic functioning of the cell. It will be a fundamental task for the future to integrate the information-theoretic analysis developed here, which finds its natural place in the transcriptome, with a geometric approach concerning both, the proteome as well as the transcriptome.
The terms in glossary are italicised short amino acid sequence interacting with a nucleic acid in course of contiguous genomic element acting in gene controlling the expression of other genes theoretical model of eukaryotic gene regulation proposing stepwise reduction of the genomic information potential in course of RNA processing and transport network of filaments (some known to contain DNA since running in and out of the electron microscope fragment of a coding sequence in the DNA placed between full domain transcript, RNA resulting from the transcription of an entire genomic domain in the DNA; generally but not necessarily identical to here defined as the uninterrupted nucleic acid stretch of the coding sequence in the mRNA that corresponds to a polypeptide or another functional product; thus, in eukaryotes typically not yet present at DNA level, but assembled from gene fragments (exons) in course of RNA DNA domain containing fragments of one or several genes coordinated by program controlling the expression of a gene, superimposed onto and added to the coding sequence in ensemble of all ensemble of all factors that can respond to the non-coding stretch of DNA placed between exons in the genomic DNA (synonymous to intervening sequence) matrix attachment region where a DNA sequence is linked to the nuclear matrix and, hence, protected to DNase digestion messenger ribonucleic acid, carrying the coding sequence of a gene as well as specific signals guiding its expression (the nucleic acid binding protein nuclear body where the (highly amplified) ribosomal DNA is located and the ribosomal subunits synthesised oligonucleotide sequence, recognized by specific amino-acid motifs ( a sequence of genetic information temporally stored outside the genomic DNA in form of regulative interventions after transcription at the level of pre-mRNA and mRNA, according to the corresponding (pre-)genons; to be distinguished from precursor of primary transcript that is converted into primary transcript that is converted into ribosomal RNA by mechanism of cleavage of transcripts ( polypeptide and its coding sequence, equivalent of triplet-based coding sequence in signals at DNA level that control, via mechanism of transient or final repression of specific gene coding for a functional RNA ribonucleoprotein complex, i.e, complex of RNA and proteins (selective binding of proteins to mRNA is essential for regulation of the gene expression process) ribosomal RNA backbone, aligning the ribosomal proteins to form the 30S (18S rRNA) and 50S (28S rRNA) ribosomal subunits; has, furthermore, ribozyme function particular type of RNA gene contributing to cellular structure, either directly or via enzymatic activities ensemble of regulation at the level of the polyribosomes during translation of mRNA postulates that a large part of the non-coding DNA has an architectural function, providing a frame for the selective interaction of specific regions in the DNA, within or between chromosomes, as seen in untranslated region preceding or following the coding sequence in mRNA probability of an event or a message contingent upon the occurrence of another event or message uncertainty about a specific element to be chosen from an ensemble of elements with known probabilities uncertainty about the content of a message prior to its reception, on the basis of known probabilities for the various possible messages (see formula in text); expected information to be gained from receiving a message uncertainty about a specific sequence composed from symbols with known probabilities and correlations
We thank our colleagues who contributed by discussion over the years to the evolution of the ideas presented here, and in particular, Manfred Eigen and the participants of the Klosters Winter Seminar (1997–2007). The first author thanks the Max Planck Institute for Mathematics in the Sciences (Leipzig) for its hospitality and best working conditions. The excellent secretarial help of Antje Vandenberg is gratefully acknowledged. This work was supported by the French CNRS, the Universities Paris 6 and 7, and by bioMérieux SA.
This paper appeared only after the review (Scherrer and Jost
“Nascent”: we use this term in its strict logical meaning of “at birth”, or the final product when released from the site of formation.
The negative sign in front of the sum arises here to make the whole expression positive, because the
“Relative” here expresses the normalization ∑
When looking at finer details of the regulation process, however, that redundancy dissolves. For example, the splicing process depends on certain recognition sites in exonic regions for the formation of certain RNPs, and here, triplets that translate into the same amino acid can be functionally different. Also, even at the translation stage, the frequency of translation depends on the presence of the appropriate tRNAs, and the more frequent triplets might also have more tRNA partners and are therefore also more frequently translated. Thus, different frequencies of triplets coding for the same amino acid can make a functional difference in the cell.
Here, biochemically, one should think of oligonucleotides; for example,
For
In fact, in the investigations of B.L.Hao and his group, it was found (personal communication) that going beyond
Assuming, for simplicity, that then the coding region in the DNA is uniquely determined; in any case, even though that need not strictly hold, the number of possible coding regions for a given polypeptide chain is rather small, and therefore, there is little remaining uncertainty.
This is made more precise in the Shannon-MacMillan-Breiman theorem which tells us that the effective number of different sequences that need to be considered is