This is an Open Access article distributed under the terms of the Creative Commons Attribution License (
Two human coronaviruses are known since the 1960s: HCoV-229E and HCoV-OC43. SARS-CoV was discovered in the early spring of 2003, followed by the identification of HCoV-NL63, the fourth member of the coronaviridae family that infects humans. In this study, we describe the genome structure and the transcription strategy of HCoV-NL63 by experimental analysis of the viral subgenomic mRNAs.
The genome of HCoV-NL63 has the following gene order: 1a-1b-S-ORF3-E-M-N. The GC content of the HCoV-NL63 genome is extremely low (34%) compared to other coronaviruses, and we therefore performed additional analysis of the nucleotide composition. Overall, the RNA genome is very low in C and high in U, and this is also reflected in the codon usage. Inspection of the nucleotide composition along the genome indicates that the C-count increases significantly in the last one-third of the genome at the expense of U and G. We document the production of subgenomic (sg) mRNAs coding for the S, ORF3, E, M and N proteins. We did not detect any additional sg mRNA. Furthermore, we sequenced the 5' end of all sg mRNAs, confirming the presence of an identical leader sequence in each sg mRNA. Northern blot analysis indicated that the expression level among the sg mRNAs differs significantly, with the sg mRNA encoding nucleocapsid (N) being the most abundant.
The presented data give insight into the viral evolution and mutational patterns in coronaviral genome. Furthermore our data show that HCoV-NL63 employs the discontinuous replication strategy with generation of subgenomic mRNAs during the (-) strand synthesis. Because HCoV-NL63 has a low pathogenicity and is able to grow easily in cell culture, this virus can be a powerful tool to study SARS coronavirus pathogenesis.
Until recently only two human coronaviruses were known – human coronavirus (HCoV) 229E and HCoV-OC43, representatives of the group 1 and 2 coronaviruses, respectively. Both were identified in 1960s and are generally considered as common cold viruses. An outbreak of severe acute respiratory syndrome (SARS) in the spring of 2003 led to the rapid identification of SARS-CoV [
HCoV-NL63 is a member of the coronaviridae family that clusters together with arteri-, toro- and roniviruses in the order of the nidovirales. Coronaviruses are enveloped viruses with a positive, single stranded RNA genome of approximately 27 to 32 kb. The 5' two-third of a coronavirus genome encodes a polyprotein that contains all enzymes necessary for RNA replication. The expression of the complete polyprotein requires a -1 ribosomal frameshift during translation that is triggered by a pseudoknot RNA structure [
Coronavirus replication is a complex, not yet fully understood mechanism [
In this study, we analyzed the genome structure of HCoV-NL63. First, we focus on the unusual nucleotide composition of the RNA genome. We describe in detail the bias in the nucleotide composition and its influence on the codon usage of this virus. We provide a possible mechanistic explanation for a shift in nucleotide bias at two-third of the HCoV-NL63 genome that is based on the RNA replication mechanism. Second, we describe in detail the different sg mRNAs generated during HCoV-NL63 replication and their relative abundance.
We described previously that the newly identified HCoV-NL63 virus has a typical coronavirus genome structure and gene order [
Nucleotide content of coronaviridae RNA genomes. We arranged the viruses based on their C-count, which ranges from 14% (HCoV-NL63) to 20% (SARS-CoV).
To investigate if all coding regions of HCoV-NL63 display a similarly strong preference for U and against C, we also plotted the nucleotide count for the individual genes and 5' and 3' non-coding regions (Figure
Nucleotide content of individual HCoV-NL63 genes and the 5'/3' untranslated regions (UTR).
We plotted the nucleotide distribution along the genome (Figure
Nucleotide distribution along the HCoV-NL63 genome. The change in the C- and G-count at two-third of the genome is statistically significant for all tested coronaviruses (HCoV-NL63, HCoV-229E, SARS-CoV, HCoV-OC43) with p < 0.01 for C-count and p < 0.05 for G-count in Mann-Whitney U test for two independent samples.
Recently, Grigoriev reported an interesting feature within coronaviral genomes that is visible when the cumulative GC-skew is plotted [
Cumulative GC-skew diagrams for several coronaviral RNA genomes. The vertical bar indicates the border between the 1a/1b and the structural genes.
The bias in the nucleotide count led us to compare the codon usage of HCoV-NL63 with that of human mRNA (Table
Nucleotide composition of the first, second and third codon positions in the HCoV-NL63 genome.
Codon usage of HCoV-NL63 compared with that of human genes
| Amino acid | Codon | Humana | HCoV-NL63 | 1ab (20190) | S (4071nt) | ORF3 (678nt) | E (234nt) | M (681nt) | N (1134nt) |
| Arg | CGA | 0.62b | 0.16 | 0.12 | 0.22 | 0.44 |
|
0.00 | 0.26 |
| CGC | 1.07 | 0.28 | 0.21 | 0.37 | 0.00 | 0.00 | 0.88 | 1.06 | |
| CGG |
|
0.06 | 0.04 | 0.15 | 0.00 | 0.00 | 0.00 | 0.00 | |
| CGU | 0.46 |
|
|
|
|
0.00 |
|
|
|
| AGA |
|
0.78 | 0.76 | 0.74 | 0.88 | 0.00 | 1.32 | 1.06 | |
| AGG |
|
0.47 | 0.49 | 0.22 | 0.00 |
|
0.00 | 1.32 | |
| Leu | CUA | 0.70 | 0.39 | 0.31 | 0.29 | 2.65 | 1.28 | 0.88 | 0.26 |
| CUC | 1.97 | 0.36 | 0.24 | 0.74 | 0.88 | 2.56 | 0.44 | 0.26 | |
| CUG |
|
0.21 | 0.18 | 0.37 | 0.44 | 1.28 | 0.00 | 0.00 | |
| CUU | 1.30 |
|
2.84 | 2.87 |
|
|
|
|
|
| UUA | 0.74 | 2.68 | 2.75 | 2.51 | 2.65 |
|
4.41 | 0.79 | |
| UUG | 1.28 |
|
|
|
3.54 | 2.56 | 2.20 | 2.38 | |
| Ser | UCA | 1.20 | 1.49 | 1.38 | 1.92 | 1.33 | 0.00 | 0.44 | 2.91 |
| UCC | 1.76 | 0.33 | 0.30 | 0.59 | 0.00 |
|
0.00 | 0.26 | |
| UCG | 0.45 | 0.08 | 0.07 | 0.00 | 0.44 | 0.00 | 0.00 | 0.26 | |
| UCU | 1.49 |
|
|
|
|
0.00 |
|
|
|
| AGC |
|
0.34 | 0.27 | 0.59 | 0.00 | 0.00 | 0.88 | 0.79 | |
| AGU | 1.21 | 2.47 | 2.51 | 2.36 |
|
|
3.08 | 2.12 | |
| Thr | ACA | 1.49 | 1.72 | 1.89 | 1.33 | 0.88 |
|
|
0.26 |
| ACC |
|
0.44 | 0.40 | 0.52 | 0.00 |
|
0.44 | 1.06 | |
| ACG | 0.62 | 0.19 | 0.13 | 0.37 | 0.44 | 0.00 | 0.88 | 0.00 | |
| ACU | 1.30 |
|
|
|
|
|
1.76 |
|
|
| Pro | CCA | 1.68 | 1.03 | 1.01 | 1.11 | 0.44 |
|
0.44 | 1.59 |
| CCC |
|
0.18 | 0.15 | 0.22 | 0.00 | 0.00 | 0.00 | 0.79 | |
| CCG | 0.70 | 0.10 | 0.07 | 0.22 | 0.00 | 0.00 | 0.44 | 0.00 | |
| CCU | 1.74 |
|
|
|
|
1.28 |
|
|
|
| Ala | GCA | 1.60 | 1.51 | 1.66 | 1.33 | 0.00 |
|
1.32 | 0.26 |
| GCC |
|
0.64 | 0.55 | 1.11 | 0.88 | 0.00 | 0.44 | 0.79 | |
| GCG | 0.75 | 0.14 | 0.13 | 0.22 | 0.00 | 0.00 | 0.00 | 0.26 | |
| GCU | 1.86 |
|
|
|
|
|
|
|
|
| Gly | GGA | 1.64 | 0.46 | 0.49 | 0.44 | 0.00 | 0.00 | 0.44 | 0.26 |
| GGC |
|
0.48 | 0.34 | 1.18 | 1.33 | 0.00 | 0.44 | 0.00 | |
| GGG | 1.65 | 0.14 | 0.13 | 0.07 | 0.00 | 0.00 | 0.44 | 0.53 | |
| GGU | 1.08 |
|
|
|
|
|
|
|
|
| Val | GUA | 0.71 | 1.04 | 1.19 | 0.59 | 0.00 | 1.28 | 0.88 | 0.79 |
| GUC | 1.46 | 0.89 | 0.77 | 0.96 | 2.65 | 2.56 | 1.76 | 0.79 | |
| GUG |
|
0.77 | 0.68 | 1.03 | 0.88 | 1.28 | 2.20 | 0.26 | |
| GUU | 1.10 |
|
|
|
|
|
|
|
|
| Lys | AAA | 2.40 |
|
|
1.33 |
|
|
1.32 | 3.70 |
| AAG |
|
2.25 | 2.29 |
|
1.77 | 0.00 |
|
|
|
| Asn | AAC |
|
1.27 | 1.05 | 2.28 | 0.44 | 0.00 | 0.88 | 2.38 |
| AAU | 1.67 |
|
|
|
|
|
|
|
|
| Gln | CAA | 1.2 |
|
|
|
|
|
0.44 | 1.59 |
| CAG |
|
1.03 | 0.76 | 1.69 | 0.00 | 0.00 |
|
|
|
| His | CAC |
|
0.39 | 0.39 | 0.52 | 0.44 | 0.00 | 0.00 | 0.27 |
| CAU | 1.07 |
|
|
|
|
|
|
|
|
| Glu | GAA | 2.89 |
|
|
|
|
|
|
2.12 |
| GAG |
|
1.12 | 1.17 | 0.66 | 0.44 | 0.00 |
|
|
|
| Asp | GAC |
|
1.19 | 1.2 | 1.03 |
|
1.28 |
|
0.79 |
| GAU | 2.2 |
|
|
|
1.33 |
|
0.88 |
|
|
| Tyr | UAC |
|
0.89 | 0.82 | 1.03 | 1.77 | 1.28 | 1.32 |
|
| UAU | 1.21 |
|
|
|
|
|
|
0.79 | |
| Cys | UGC |
|
0.30 | 0.31 | 0.29 | 0.44 | 0.00 | 0.00 |
|
| UGU | 1.03 |
|
|
|
|
|
|
0.00 | |
| Phe | UUC |
|
0.68 | 0.58 | 0.74 | 1.77 | 2.56 | 0.88 | 1.06 |
| UUU | 1.72 |
|
|
|
|
|
|
|
|
| Ile | AUA | 0.73 | 1.30 | 1.29 | 1.62 | 1.33 | 2.56 | 0.88 | 0.26 |
| AUC |
|
0.33 | 0.30 | 0.44 | 0.00 | 0.00 | 1.32 | 0.26 | |
| AUU | 1.58 |
|
|
|
|
|
|
|
a data obtained from GenBank Release 142.0 [28].
b all values represent the percentage of a specified codon.
c the highest value for each codon group is typed bold.
The 5' end of HCoV-NL63 genome RNA contains the L sequence of 72 nucleotides that ends with the L TRS element. This TRS has a high similarity to short sequences that are located in front of each open reading frame (S-ORF3-E-M-N) [
Inspection of sg mRNA junctions indicated that they are indeed composed of the part of the HCoV-NL63 genome that is directly downstream of a particular body TRS, with its 5' end derived from the leader sequence. Apparently, strand transfer occurred on the 5' end of the body TRS, as indicated in Figure
Body-leader junctions of all HCoV-NL63 sg mRNAs. Shown on top is the leader (L) sequence and below the specific sequences upstream of the structural genes. The fusion of 5' L sequences to 3' sg RNA is indicated by the boxes. Sequence homology between the strands near the junction is marked by asterisks, the conserved AACUAAA TRS core is highlighted in gray.
To determine whether the predicted sg mRNAs encoding the S-ORF3-E-M-N proteins are produced in virus-infected cells, we performed Northern blot analysis on total cellular RNA (Figure
The left panel shows the Northern blot analysis of HCoV-NL63 RNA in infected LLC-MK2 cells. RNA of HCoV-NL63 (NL63 lane) was compared with RNA of MHV strain A59 (MHV lane). Non-infected LLC-MK2 cells are included as a negative control (control lane). MHV RNA bands represent the complete genome (1) and sg mRNAs 2a (2), S (3), 17.8 (4), 13.1 and E (5), M (6), N (7). HCoV-NL63 RNA includes the complete genome (1) and sg mRNAs for S (2), ORF3 (3), E (4), M (5) and N (6). The right panel shows the MHV and HCoV-NL63 genome organization and the HCoV-NL63 sg-mRNAs.
To determine the expression level of each subgenomic RNA, we measured the intensity of the signals. When plotted as a function of the genome position (Figure
Expression levels of the HCoV-NL63 genomic and sg mRNAs.
We analyzed the nucleotide composition of the HCoV-NL63 genomic (+) RNA, which was found to exhibit a typical coronavirus pattern with an abundance of U (39 %) and shortage of G (20%) and C (14%). In fact, HCoV-NL63 has the most pronounced nucleotide bias among the coronaviridae.
There is a significant fluctuation in the nucleotide count among the HCoV-NL63 genes. For instance, ORF3 and M appear as extreme U-rich and A-poor islands. It is possible that the unique nucleotide composition of some structural genes reflects their evolutionary origin, perhaps suggesting that some of these functions were acquired recently from another viral or cellular origin by gene transfer. These properties mimic the pathogenicity islands of prokaryotic genomes [
Inspection of the nucleotide composition along the genome indicates a bi-phasic pattern. The 5' two-third of the genome encoding the 1ab polyprotein has a stable nucleotide count with the typical U>A>G>C order, but rather striking differences are observed in the 3' one-third of the genome that encodes the structural proteins (Figure
We show that U-counts reach the highest values and C-counts the lowest values at the third position of the HCoV-NL63 codons (Figure
Inspection of the viral genome sequence led us to predict that the 1ab polyprotein is expressed from the genomic RNA and the 3' structural proteins and ORF3 from 5 distinct sg mRNAs. This was confirmed experimentally. We observed that sg mRNAs are more abundant when the corresponding TRS is located closer to the 3' end of the genome. The exception is formed by the E sg mRNA, which is relatively underexpressed. This may correlate with the low expression level of this protein. The general trend of increased gene expression along the genome has been reported previously for other coronaviruses [
The core sequence AACUAAA is conserved in the L TRS and all body TRSs, except for the E gene that has a single mismatch AACUA
The nucleotide content of different Coronaviridae family members was assessed using BioEdit software. The nucleotide distribution was determined using a Microsoft Excel datasheet (300 nucleotide (nt) window and 10-nt step). Codon usage was assessed using DNA 2.0 software. Data was processed in Microsoft Excel datasheet and all statistical analysis was performed with SPSS 11.5.0 software. The level of significance of the nucleotide bias was established for 300-nt non-overlapping windows with the non-parametric Mann-Whitney U test for two independent samples. Cumulative GC-skew graphs were generated as described previously [
HCoV-NL63 RNA was obtained from virus-infected LLC-MK2 cells (2 × 107) after 6 days of culture (virus passage 7). Mouse Hepatitis Virus (MHV) RNA was obtained by infecting 2 × 107 LR7 cells with MHV strain A59. The medium was removed and the cells were dissolved in 15 ml TRIzol® and RNA was isolated according to the standard TRIzol® procedure. RNA was subsequently precipitated with 0.8 volume of isopropanol, dried and dissolved in 50 μl H2O. Integrity of the RNA was analyzed by electrophoresis on a non-denaturating 0.8% agarose gel. RNA was stored at -150°C.
The cDNA used for sequencing and probe construction was made by MMLV-RT on viral RNA with 1 μg of random hexamer DNA primers in 10 mM Tris pH 8.3, 50 mM KCl, 0.1% Triton-X100, 6 mM of MgCl2 and 50 μM of each dNTPs at 37°C for 1 hour. The single stranded cDNA product was made into double-stranded DNA in a standard PCR reaction with 1.25 U of Taq polymerase (Perkin-Elmer) per reaction with appropriate primers (see below).
Gel electrophoresis of viral RNA was performed on a 1% agarose gel with 7% of formaldehyde at 100 Volt in 1×MOPS buffer (40 mM MOPS, 10 mM sodium acetate, pH 7.0). Transfer onto a positively charged nylon membrane (Boehringer Mannheim) was done overnight by means of capillary force. RNA was linked to the membrane in a UV crosslinker (Stratagene). For generation of the HCoV-NL63 probe, the RT-PCR product was further amplified with 5' primer N5PCR1 (CTG TTA CTT TGG CTT TAA AGA ACT TAG G) and 3' primer N3PCR1 (CTC ACT ATC AAA GAA TAA CGC AGC CTG). Similarly, the MHV probe was amplified with 5' primer MHV_UTR-B5' (GAT GAA GTA GAT AAT GTA AGC GT) and 3' primer MHV_UTR-B3' (TGC CAC AAC CTT CTC TAT CTG TTA T). Labeling of the probes was done in a standard PCR reaction with specific 3' primers (N3PCR1 and MHV_UTR-B3') in presence of [α-32P]dCTP. Prehybridization and hybridization was done in ULTRAhyb buffer (Ambion) at 50°C for 1 and 12 hours, respectively. The membrane was then washed at room temperature with low-stringency buffer (2×SSC, 0.2% SDS) and at 50°C in high stringency buffer (0.1×SSC, 0.2% SDS). Images were obtained using the STORM 860 phosphorimager (Amersham Biosciences) and data analysis was performed with the ImageQuant software package. The size of sg mRNA fragments of HCoV-NL63 were estimated from their migration on the Northern blot using the sg mRNA of MHV as size marker.
The L/body TRS junctions were PCR-amplified from an HCoV-NL63 cDNA bank. We performed 35 cycle PCR with the 5' L primer (L5 – TAA AGA ATT TTT CTA TCT ATA GAT AG) and gene specific 3' primers (S gene – SL3' – ACT ACG GTG ATT ACC AAC ATC AAT ATA; ORF3 – 4L3' – CAA GCA ACA CGA CCT CTA GCA GTA AG; E gene – EL3' – TAT TTG CAT ATA ATC TTG GTA AGC; M gene – ML3' – GAC CCA GTC CAC ATT AAA ATT GAC A; N gene – 3-163-F15 – ATT ACC TAG GTA CTG GAC CT). The PCR products were analyzed by electrophoresis on a 0.8% agarose gel and products of discrete size were used for sequencing using the BigDye terminator kit (ABI) and ABI Prism 377 sequencer (Perkin Elmer). Sequence analysis was performed by Sequence Navigator and AutoAssembler 2.1 software.
The complete genome sequence of HCoV-NL63 [
The authors declare that they have no competing interests.
KP carried out the viral RNA isolation, RT-PCR, sequencing of sg mRNAs, Northern blot evaluation and all computer analysis done in this study; MFJ carried out the full genome sequencing; all authors participated in writing the manuscript. Lv/dH and BB are the principal investigators
We thank Berend Jan Bosch and Peter Rottier for providing MHV infected cells and Alexander Nabatov and Barbara van Schaik for technical support.