The environment has been playing an instrumental role in shaping and maintaining the morphological, physiological and biochemical diversities of prokaryotes. It has been debatable whether the whole-genome Guanine-Cytosine (GC) content levels of prokaryotic organisms are correlated with their optimal growth temperatures. Since the GC content is variable within a genome, we here focus on the correlation between the genic GC content levels and the temperature range conditions of prokaryotic organisms.
The GC content levels in the coding regions of four genes were consistently identified as correlated with the temperature range condition when the association analysis was applied to (
To further validate the identified correlation relationships, we examined to what extent the temperature range condition of an organism can be predicted based on the GC content levels in the coding regions of the selected genes. The 84.52% accuracy for the complete genomes, the 84.09% accuracy for the in-progress genomes, and 82.70% accuracy for the metagenomes, especially when being compared to the 50% accuracy rendered by random guessing, suggested that the temperature range condition of a prokaryotic organism can generally be predicted based on the GC content levels of the selected genomic regions.
The results rendered by various statistical tests and prediction tests indicated that the GC content levels of the coding/non-coding regions of certain genes are highly likely to be correlated with the temperature range conditions of prokaryotic organisms. Therefore, it is promising to carry out “reverse ecology” and to complete the ecological characterizations of prokaryotic organisms, i.e., to infer their temperature range conditions based on the GC content levels of certain genomic regions.
There are countless ways in which prokaryotes influence our daily life. For example, they mediate the chemical cycles that convert key elements of life into biologically accessible forms; they make certain nutrients/metals/vitamins available to their biological hosts; and they can also be used to breakdown hydrocarbons and treat crude oil leakages. On the other hand, the environment has undoubtedly left footprints on the morphological, physiological and biochemical diversities of prokaryotes during the evolution. The close interactions between prokaryotes and the environment, especially driven by horizontal gene transfer and homologous recombination, have made prokaryotes the most genetically diverse superkingdoms of life [
Temperature is one of the elements characterizing the ecological contexts of prokaryotic organisms. The National Center for Biotechnology Information (NCBI) Microbial Genome Project Database (
Guanine-Cytosine (GC) content is one of the genomic traits that have been hypothesized to be correlated with the temperature condition of an organism. Since the GC pair is bound by three hydrogen bonds while the adenine-thymine (AT) pair is bound by two hydrogen bonds, it has long been expected that organisms growing at higher temperature would have a higher proportion of GC than AT pairs. However, it remains ambiguous whether the whole-genome GC content level of an organism is correlated with its temperature condition. Analysis on 368 bacterial species seemed to have confirmed the existence of the positive correlation between the whole-genome GC content level and the optimal growth temperature of prokaryotes [
It is worthy of notice that the GC content can be substantially variable within the same genome. For instance, the GC content in coding regions is often higher than that of the whole genome [
The distribution of the GC content level of the non-coding region surrounding the menB gene (K01661: naphthoate synthase) for mesophilic (dashed, red) and thermophilic/hyperthermophilic organisms (solid, blue).
Among the 895 complete prokaryotic genomes available from the Kyoto Encyclopedia of Genes and Genomes (KEGG) (
Number of organisms under various temperature range, oxygen requirement, habitat and salinity conditions, where the number in the parentheses is the number of organisms for a specific combination of the temperature range condition and a condition of some other environmental factor.
| Condition | #Org | Habitat | Oxygen | Salinity |
|---|---|---|---|---|
| Psychrophilic | 14 | Aquatic(5) | Facultative(7) | Mesophilic(2) |
| Specialized(4) | Aerobic(3) | - | ||
| - | - | - | ||
|
|
||||
| Mesophilic | 722 | Host-associated(284) | Facultative(282) | Non-halophilic(149) |
| Multiple(242) | Aerobic(244) | Mesophilic(10) | ||
| Aquatic(111) | Anaerobic(99) | Moderate halophilic(9) | ||
|
|
||||
| Thermophilic | 57 | Specialized(36) | Anaerobic(26) | Non-halophilic(4) |
| Aquatic(10) | Aerobic(13) | Moderate halophilic(2) | ||
| Multiple(6) | Facultative(11) | Mesophilic(1) | ||
|
|
||||
| Hyperthermophilic | 36 | Specialized(21) | Anaerobic(22) | Non-halophilic(3) |
| Aquatic(12) | Aerobic(8) | Moderate halophilic(3) | ||
| - | - | Mesophilic(2) | ||
Among the 2,798,133 genes of the 815 organisms being studied, 1,210,908 (~43.3%) genes are assigned with KEGG Orthology (KO) IDs, and a total number of 6,026 KO groups are being covered. We considered genes with the same KO IDs as orthologous, so that they can be used to estimate the distribution of the GC content level surrounding each gene (corresponding to a unique KO ID) in various genomes. For each gene, we obtained the GC content levels of both the coding and surrounding non-coding regions, where the non-coding region starts from the end of the previous gene till the start of the next gene with the coding region of the current gene being excluded. As the coding and non-coding regions were treated independently, we used
Our association analysis consisted of two steps – statistical tests and prediction tests. The Kolmogorov-Smirnov (KS) statistical tests were first carried out to identify those genomic regions of which the distribution patterns of the GC content levels are different for organisms of different temperature range conditions but are irrelevant to the phylogenetic distribution of the organisms or the distribution of other environmental factors. The support vector machine (SVM)-based prediction tests were then carried out to determine whether the temperature range condition of a prokaryotic organism can be inferred from the GC content levels of the genomic regions selected via the statistical tests.
The GC content estimates for each genomic region can be organized into two groups – one consisting of those obtained from the 722 mesophilic organisms and the other consisting of those obtained from the 93 thermo/hyperthermo-philic organisms. The KS test was conducted to determine whether these GC content levels are distributed differently for these two groups of organisms of different temperature range conditions. Specifically, the GC content of a genomic region was considered as potentially correlated with the temperature range condition if its corresponding
The distribution of the mesophilic organisms in various phylogenetic groups are different from that of the thermo/hyperthermo-philic organisms. Observe from Fig.
The distribution of the mesophilic organisms (left) and the distribution of the thermo/hyperthermo-philic organisms (right) among various phylogenetic groups.
To examine whether the distribution of the GC content level is correlated with the phylogenetic origin of an organism and whether the temperature range condition and the phylogenetic origin are coupled in being correlated with the distribution of GC content levels, we should ideally perform the multi-variate ANOVA [
Various physiological features, including the gram stain, shape, arrangement, endospores, motility, salinity, oxygen requirement, habitat, organisms that the prokaryote is pathogenic in, and the related disease, are together with the temperature range conditions provided in NCBI’s Microbial Genome Project Database for each prokaryotic genome. Besides the temperature range condition, we were also interested in the features regarding the salinity, oxygen requirement, and habitat, because (
Observe from Table
The reported association analysis on the genomic and environmental traits might be prone to the annotation errors at both the genomic and environmental sides. At the genomic side, since the annotations in the KEGG database involve both computation-based and manual curation [
Comparisons of environmental annotations in the NCBI and GOLD databases for the 895 complete prokaryotic genomes. For temperature range, the number of annotated organisms for three dominant conditions (mesophilic, thermophilic and hyperthermophilic) are shown in blue, green and brown, respectively. For oxygen requirement, the number of annotated organisms for three dominant conditions (aerobic, anaerobic and facultative) are shown in blue, green and brown, respectively.
From the statistical point of view, the significance of the correlation relationships between the temperature range condition and the GC content levels of the selected genomic regions can be measured by the
For the prediction test, each organism was represented by a feature vector consisting of the GC content levels of the selected genomic regions. A cross validation procedure [
To test the generalizability of the identified correlation relationships, we also applied the SVM-based classifier to predicting the temperature range conditions of 17 mesophilic and 17 thermo/hyperthermo-philic in-progress genomes in NCBI’s Microbial Genome Project Database (see Table S-2 of the supplementary material), as well as 29 mesophilic and 29 thermo/hyperthermo-philic metagenomes in the Integrated Microbial Genomes (IMG/M) system [
We here discuss the results of the statistical and prediction tests.
A collection of 413 genomic regions (including the coding and non-coding regions of 80 genes, the coding regions of 197 genes, and the non-coding regions of 56 genes) were consistently detected in all the 200 bootstrap KS tests and were included into {
Top 10 KEGG pathways enriched by genes of {
| Pathway ID | Specific Level Description | General Level Description |
|---|---|---|
| ko02040 | Flagellar assembly | Cell Motility |
| ko02030 | Bacterial chemotaxis | Cell Motility |
| ko00471 | D-Glutamine and D-glutamate metabolism | Metabolism of Other Amino Acids |
| ko00643 | Styrene degradation | Xenobiotics Biodegradation and Metabolism |
| ko00720 | Reductive carboxylate cycle (CO2 fixation) | Energy Metabolism |
| ko00523 | Polyketide sugar unit biosynthesis | Biosynthesis of Polyketides and Nonribosomal Peptides |
| ko00130 | Ubiquinone and menaquinone biosynthesis | Metabolism of Cofactors and Vitamins |
| ko00521 | Streptomycin biosynthesis | Biosynthesis of Secondary Metabolites |
| ko00271 | Methionine metabolism | Amino Acid Metabolism |
| ko00730 | Thiamine metabolism | Metabolism of Cofactors and Vitamins |
It can been seem from Table
When based on the GC content levels of the genomic regions in {
There were four genomic regions in {
The mean and standard deviation (std) of the GC content level in the four consistently selected genomic regions for mesophilic and thermo/hyperthermo-philic organisms, the genes corresponding to these four genomic regions and their functional annotations.
| KO Group | Function Description | mean±std | |
|---|---|---|---|
|
|
|||
| mesophilic | thermo/hyperthermo-philic | ||
|
|
Adenosylhomocysteinase | 57.35 ± 9.13% | 48.61 ± 10.47% |
|
|
DNA repair and recombination proteins | 62.17 ± 11.85% | 47.90 ± 12.99% |
|
|
LAO/AO transport system kinase | 62.03 ± 9.96% | 49.10 ± 13.40% |
|
|
Poorly Characterized | 66.04 ± 10.26% | 45.43 ± 11.03% |
These four genomic regions correspond to the coding regions of K01251 (adenosylhomocysteinase), K03724 (DNA repair and recombination proteins), K07588 (LAO/AO transport system kinase), and K09122 (hypothetical protein), and are interpreted as temperature-correlated but irrelevant to the distribution in various oxygen requirement, salinity, habitat and phylogenetic groups. Some of these computation-based findings can be further supported by experiment-based findings. For instance, K01251 reflects a common adaptation mechanism at high temperatures, because all the residues forming the network of aromatic and hydrophobic contacts and the residues potentially involved in cofactor binding are fairly well conserved among the hyperthermophilic but not in the mesophilic organisms [
The phylogenetic trees of these four genes were built by using ClustalW2 [
Number of neighboring organisms with the same temperature range conditions in the phylogenetic trees built from the four genes and the 16S rRNA genes. The number in parentheses is the number of organisms with the corresponding gene present.
| Temperature | group 1 (412) | group 2 (288) | group 3 (223) | group 4 (107) | ||||
|---|---|---|---|---|---|---|---|---|
|
|
||||||||
| K01251 | 16S | K03724 | 16S | K07588 | 16S | K09122 | 16S | |
|
|
323 | 318 | 234 | 230 | 158 | 161 | 52 | 53 |
|
|
57 | 59 | 34 | 33 | 35 | 39 | 43 | 45 |
When based on these four genomic regions, the prediction accuracy, averaged on the 500 cross-validation experiments, was 84.52% for the complete genomes (84.46% for mesophiles, 85.09% for thermo/hyperthermo-philes), 84.09% for the in-progress genomes (83.05% for mesophiles, 85.13% for thermo/hyperthermo-philes), and 82.70% for the metagenomes (79.31% for mesophiles, 86.21% for thermo/hyperthermo-philes), respectively, all of which were much higher than the 50% accuracy rendered by random guessing. It should be noted that with less than 1% of the elements in {
When the KS tests were conducted on the mesophilic vs. thermo/hyperthermo-philic organisms based on the temperature range annotations in the GOLD database, eight genomic regions were consistently selected out of the tests on the complete list of organisms and on the partial lists of organisms with particular phylogenetic/oxygen requirement/habitat/salinity categories being excluded. These genomic regions include the coding regions of K01251 (adenosylhomocysteinase), K03724 (DNA repair and recombination proteins), K07588 (LAO/AO transport system kinase), K08289 (purine metabolism), and K09122 (hypothetical protein), and the non-coding regions of K00261 (glutamate dehydrogenase), K03809 (Trp repressor binding protein), and K07008 (currently unclassified). Note that the coding regions of K01251, K03724, K07588 and K09122 were the genomic regions consistently selected when the KS tests were based on the temperature range annotations in the NCBI’s Microbial Genome Project Database. That is, despite the 11.51% inconsistency in the temperature range annotations in the NCBI and GOLD databases, the four genomic regions were always identified as to possess GC content levels that are correlated with the temperature range conditions, suggesting their robustness to the possible annotation errors in the databases.
There were six genomes whose temperature range conditions were consistently mis-classified (in at least 80% of the cross validation experiments that they were used for testing), whether the prediction was based on the entire {
We have investigated the correlation relationships between the GC content levels in the coding and surrounding non-coding regions of individual genes and the temperature range conditions of prokaryotic organisms based on the statistical tests and prediction tests. Through the KS tests conducted on all the mesophilic and thermo/hyperthermo-philic organisms, we have identified 413 genomic regions (277 coding and 136 non-coding regions) whose GC content levels show different distribution patterns for organisms under different temperature range conditions and that can be considered as potentially correlated with the temperature range condition. When these 413 genomic regions were used to predict the temperature range condition of an organism, the prediction accuracy can reach 92.45% for the complete genomes and 92.05% for the in-progress genomes. Four of these 413 genomic regions corresponding to the coding regions of K01251 (adenosylhomocysteinase), K03724 (DNA repair and recombination proteins), K07588 (LAO/AO transport system kinase), and K09122 (hypothetical protein) were consistently selected when the KS tests were conducted on partial lists of organisms with particular phylogenetic/oxygen requirement/habitat/salinity conditions being excluded and/or when the environmental annotations in the NCBI’s Microbial Genome Project Database or in the GOLD database were used. When these four genomic regions were used to predict the temperature range condition of an organism, the prediction accuracy reached 84.52% for the complete genomes, 84.09% for the in-progress genomes, and 82.70% for the metagenomes. Considering that they only account for less than 1% of all the 413 genomic regions potentially correlated with the temperature range condition but can to a great extent retain the prediction accuracy, we may interpret these four genomic regions as the core of the temperature range-correlated. Our results have also demonstrated that the correlation relationships between the genomic and ecological traits can potentially facilitate reverse ecology.
Supplementary materials for the conditions of different environmental factors, identified genomic regions, in-progress and meta-genomes, and consistently misclassified genomes are available at:
HZ and HW have carried out the experiments and written the manuscript together.
The authors declare that they have no competing interests.
This article has been published as part of