Re-use of this article is permitted in accordance with the Creative Commons Deed, Attribution 2.5, which does not permit commercial exploitation.
Most human cytomegalovirus (HCMV) genes are highly conserved in sequence among strains, but some exhibit a substantial degree of variation. Two of these genes are UL146, which encodes a CXC chemokine, and UL139, which is predicted to encode a membrane glycoprotein. The sequences of these genes were determined from a collection of 184 HCMV samples obtained from Africa, Australia, Asia, Europe, and North America. UL146 is hypervariable throughout, whereas variation in UL139 is concentrated in a sequence encoding a potentially highly glycosylated region. The UL146 sequences fell into 14 genotypes, as did all previously reported sequences. The UL139 sequences grouped into 8 genotypes, and all previously reported sequences fell into a subset of these. There were minor differences among continents in genotypic frequencies for UL146 and UL139, but no clear geographical separation, and identical nucleotide sequences were represented among communities distant from each other. The frequent detection of multiple genotypes indicated that mixed infections are common. For both genes, the degree of divergence was sufficient to preclude reliable sequence alignments between genotypes in the most variable regions, and the mode of evolution involved in generating the genotypes could not be discerned. Within genotypes, constraint appears to have been the predominant mode, and positive selection was detected marginally at best. No evidence was found for linkage disequilibrium. The emerging scenario is that the HCMV genotypes developed in early human populations (or even earlier), becoming established via founder or bottleneck effects, and have spread, recombined and mixed worldwide in more recent times.
Human cytomegalovirus (HCMV; family
Most genes are highly conserved in sequence between HCMV strains, but a number of genes predicted to encode membrane-associated or secreted proteins are characterized by a striking degree of variability, as revealed by examination of individual genes [reviewed in
One of the most variable HCMV genes is UL146, which encodes a chemokine designated vCXC-1. This gene is variable throughout its length [
The most variable gene in the vicinity of UL146 is UL139, which is located 5.2 kbp distant and is predicted to encode a type I membrane glycoprotein. Variability is concentrated in a region of the ectodomain [
The aims of the present study were to investigate whether additional UL146 genotypes exist in a large number of clinical samples obtained from a wide range of locations and clinical settings, and to define the range of UL139 genotypes in these samples. Ancillary interests were to examine the relative frequencies and geographical distribution of genotypes, to assess whether infections with more than one HCMV strain are common, and to investigate the evolution of UL146 and UL139.
A collection of 184DNAsamples was derived from 179 anonymized clinical samples obtained in various geographical locations in accordance with local ethical guidelines, plus 5 commonly used laboratory strains (Davis, Merlin, TB40/E, Toledo and Towne). Details of the 171 samples in the collection that yielded sequence data are available on request, and include the age, sex, and pathology of the patient, the clinical source of the sample, and the UL146 and UL139 genotypes determined. The samples numbered 18 from Australia, 10 from Hong Kong, 6 from Germany, 13 from England, 18 from The Gambia, 24 from Hungary, 7 from Italy, 6 from The Netherlands, 41 from Scotland, 5 from the USA, 8 from Wales, and 15 from South Africa. A minority of strains (40) had been passaged in human fibroblast cell culture, either as routine diagnostic specimens or as laboratory strains. DNA was extracted by standard methods from body tissues, urine, saliva or infected cells. The South African samples were obtained from the saliva of mothers (10 of whom tested HIV-negative) attending rural clinics in KwaZulu/Natal [
UL146 and UL139 were amplified separately by single round or nested PCR, using primers in conserved regions (
Primers Used for PCR and Sequencing
| Gene | Primer | Sequence(5′–3′) | Genome location |
|---|---|---|---|
| UL146 | AB4 | TAGACACTACGTCGTAAATG | 180494–180513 |
| UL146 | A162 | TGTAGAATTAGTCTAGATTCCTGA | 181524–181501 |
| UL146 | UL146–4A | GCTTGCGCGTTAGGATTGAGACAC | 180571–180594 |
| UL146 | UL146–3A | ATACCGGATATTACGAATT | 181341–181323 |
| UL139 | AB1 | GTCATTGTGAAAGTGACGTCTCAG | 186389–186412 |
| UL139 | AB2 | ATCTACTGTAAACCCTCTGCTCTG | 187148–187125 |
| UL139 | UL140–11A | GCGGCATTGGTGTACGCGTG | 186553–186572 |
| UL139 | UL140–3A | GTGGAAATTTTTACGTCATT | 187077–187058 |
With reference to RefSeq accession NC_006273.2 (HCMV strain Merlin).
For the single (and first) round, 1µl ofDNAwas added to the PCR reaction mixture, which consisted of 40µl of water, 5µl of buffer, 1µl of10µM dNTPs, 1µl of each the two primers (10µM) and 1µl (1U) of DNA polymerase (Advantage 2, BD Clontech, Basingstoke, UK). The conditions for amplification were 95°Cfor 2 min followed by 35 cycles of 95°Cfor 2 min, 60°Cfor 30 sec and 68°Cfor 1 min. Second round PCR utilized 1µl of first round PCR products as template amplified under the same conditions. PCR reactions were set up in a dedicated, PCR product–free room. Approximately one–third of the samples were tested on three separate occasions to assess reproducibility.
PCR products were separated by agarose gel electrophoresis. Appropriate DNA fragments were excised, purified using a Geneclean turbo kit (Q Biogene, Cambridge, UK), and eluted using 100µl of nucleasefree water. The single round or second round primers were used for direct sequencing.
In some cases, including those where direct sequencing indicated the presence of more than one genotype of UL146 or UL139, fragments were cloned using a pGEM–T kit (Promega, Southampton, UK). Following ligation and transformation into chemically competent
Sequence chromatograms were viewed using Editview (Applied Biosystems) and analyzed using Pregap4 and Gap4 [
Sample origin was divided into four regions (Africa, Asia, Europe, and Australia) for assessment of the geographical distribution of genotypes. Chi-square tests were used to assess the significance of variability of genotype frequencies among regions. Yates' correction for continuity was applied to chi–square tests in cases where the expected values fell below 5. Similarly, Chisquare tests with Yates' correction were applied to 60 samples where single genotypes were detected for both UL146 and UL139, in order to test for linkage disequilibrium. Samples containing mixed infections were excluded from this analysis.
The UL146 and UL139 genotypes in 184 samples were investigated by PCR and sequencing using primers in conserved regions. UL146 was amplified from 159 samples and sequences were determined from 134, and UL139 was amplified from 168 samples and sequences determined from 131.Atotal of 13 samples failed to yield products from either gene. Since some samples contained more than one virus strain, totals of 182 UL146 sequences and 183 UL139 sequences were obtained. Alignment and phylogenetic analyses involved the 350 UL146 sequences and 300 UL139 sequences derived from the present study or reported by others in the literature [
The UL146 coding sequences range in length from 114 to 126 codons, and phylogenetic analyses indicated that all fall into the 14 genotypes defined previously and designated G1–G14 [
UL146 Diversity
| Alignment lenth |
Diversity | ||||||
|---|---|---|---|---|---|---|---|
|
|
|
||||||
| Genotype | Samples | Frequences(%) | DNA | Protein | DNA |
Protein |
dN/dS |
| G1 | 34 | 9.71 | 345 | 115 | 0.011 | 0.026 | 1.19 |
| G2 | 25 | 7.14 | 360 | 120 | 0.002 | 0.005 | 1.48 |
| G3 | 10 | 2.86 | 375 | 125 | 0.010 | 0.016 | 0.50 |
| G4 | 8 | 2.29 | 369 | 123 | 0.004 | 0.006 | 0.26 |
| G5 | 16 | 4.57 | 348 | 116 | 0.007 | 0.012 | 0.71 |
| G6 | 2 | 0.57 | 351 | 117 | 0.029 | 0.051 | ND |
| G7 | 57 | 16.3 | 354 | 118 | 0.011 | 0.015 | 1.29 |
| G8 | 22 | 6.29 | 342 | 114 | 0.006 | 0.008 | 0.38 |
| G9 | 49 | 14 | 351 | 117 | 0.017 | 0.032 | 0.94 |
| G10 | 12 | 3.43 | 291 | 97 | 0.003 | 0.005 | 0.30 |
| G11 | 19 | 5.43 | 339 | 113 | 0.007 | 0.015 | 0.50 |
| G12 | 43 | 12.3 | 354 | 118 | 0.016 | 0.018 | 0.45 |
| G13 | 47 | 13.4 | 357 | 119 | 0.007 | 0.015 | 1.58 |
| G14 | 6 | 1.71 | 354 | 118 | 0 | 0 | ND |
| All | 350 | 100 | 225 | 75 | 0.642 | 0.521 | 0.27 |
Gaps removed
Jukes–Cantor
Protein diversity p from MEGA4.0
dN/dS(omega) from PAML 3.15 under the single–rate model.ND, not determined
Five percent significance for positive selection
Calculated from a comparison of a single of a member of each genotype
Differences in overall genotypic frequencies were observed (
The UL139 coding sequences range in length from 124 to 148 codons, and phylogenetic analyses indicated that all fall into 8 genotypes designated G1–G8.
Phylogenetic analysis of UL139.A:Alignment (CLUSTAL W) of amino acid sequences representing the eight genotypes. Predicted signal peptide and transmembrane sequences are highlighted in gray. Completely conserved residues are indicated in the consensus row (con). Below this is the CCMV sequence, which is included to illustrate conservation of the SETTTGTSSNSS motif (underlined). The CCMV sequence [
The protein encoded by each HCMV UL139 genotype contains a putative signal peptide sequence and a transmembrane region. Variation is concentrated in the N–terminal portion of the protein. Amino acid sequence variation between genotypes is high (p=0.275), whereas within each genotype it is low (p=0.007–0.095 with a mean of 0.025) (
UL139 Diversity
| Alignment lenth |
Diversity | ||||||
|---|---|---|---|---|---|---|---|
|
|
|
||||||
| Genotype | Samples | Frequency(%) | DNA | Protein | DNA |
Protein |
dN/dS |
| G1 | 48 | 16 | 255 | 85 | 0.059 | 0.095 | 0.82 |
| G2 | 82 | 27.33 | 240 | 80 | 0.009 | 0.009 | 0.76 |
| G3 | 29 | 9.66 | 339 | 113 | 0.014 | 0.013 | 0.34 |
| G4 | 68 | 22.66 | 201 | 67 | 0.023 | 0.015 | 0.76 |
| G5 | 28 | 9.33 | 255 | 124 | 0.018 | 0.024 | 0.65 |
| G6 | 24 | 8 | 312 | 104 | 0.010 | 0.013 | 1.08 |
| G7 | 14 | 4.66 | 237 | 79 | 0.006 | 0.015 | 2.38 |
| G8 | 7 | 2.33 | 228 | 140 | 0.007 | 0.007 | 0.19 |
| All | 300 | 100 | 153 | 51 | 0.285 | 0.275 | 0.48 |
Gaps removed
Jukes–Cantor
Protein diversity p from MEGA4.0.
dN/dS(omega) from PAML 3.15 under the single–rate model
One persent significance for positive selection
Calculated from a comparision of a single member of each genotype
Differences in overall genotypic frequencies were observed (
In order to assess positive selection (i.e., for amino acid sequence diversity), the dN/dS ratio was calculated for each UL146 and UL139 genotype (
The sequence data derived in the present work were divided into four groups representing strains obtained from Africa, Asia, Australia, and Europe. Insufficient sample numbers were obtained from America to war–rant inclusion. Observation of frequencies initially suggested no significant differences in the distribution of UL146 and UL139 genotypes among continents (
Geographical Distribution of UL146 Genotypes
| Genotype | Africa | Asia | Europe | Australia |
|---|---|---|---|---|
| G1 | 4 | 2 | 7 | 0 |
| G2 | 2 | 1 | 7 | 2 |
| G3 | 4 | 0 | 2 | 0 |
| G4 | 1 | 0 | 4 | 2 |
| G5 | 3 | 0 | 2 | 1 |
| G6 | 0 | 1 | 0 | 0 |
| G7 | 5 | 6 | 21 | 1 |
| G8 | 2 | 0 | 5 | 0 |
| G9 | 6 | 2 | 11 | 3 |
| G10 | 0 | 0 | 8 | 0 |
| G11 | 0 | 0 | 3 | 0 |
| G12 | 5 | 1 | 16 | 0 |
| G13 | 11 | 0 | 17 | 0 |
| G14 | 2 | 0 | 1 | 0 |
| Total=177 | 45 | 13 | 104 | 15 |
Geographical Distribution of UL139 Genotypes
| Genotype | Africa | Asia | Europe | Australia |
|---|---|---|---|---|
| G1 | 8 | 2 | 16 | 2 |
| G2 | 9 | 1 | 23 | 5 |
| G3 | 1 | 4 | 7 | 0 |
| G4 | 8 | 5 | 23 | 5 |
| G5 | 10 | 3 | 14 | 0 |
| G6 | 0 | 1 | 11 | 1 |
| G7 | 3 | 0 | 5 | 5 |
| G8 | 2 | 2 | 1 | 1 |
| Totals=178 | 41 | 18 | 100 | 19 |
Identical nucleotide sequences were frequently obtained from geographically distant and presumably epidemiologically unrelated patients. For example, certain samples from The Gambia, Scotland, and Hungary contained identical UL146 G12 sequences. Also, UL139 G2, which was identified in 27% of samples, was represented by identical sequences from Hungary, The UK, and The Gambia.
Potential linkage disequilibrium was investigated in 60 strains for which single genotypes of both UL146 and UL139 were obtained. Of 112 possible genotype pairs, 41 were observed at least once (
Analysis of Linkage Disequilibrium
| UL139 genotype | ||||||||
|---|---|---|---|---|---|---|---|---|
|
|
||||||||
| UL146 genotype | G1 | G2 | G3 | G4 | G5 | G6 | G7 | G8 |
| G1 | 1 | 0 | 0 | 3 | 1 | 0 | 0 | 0 |
| G2 | 1 | 1 | 0 | 0 | 0 | 0 | 0 | 0 |
| G3 | 1 | 0 | 0 | 0 | 0 | 0 | 0 | 0 |
| G4 | 0 | 0 | 1 | 1 | 0 | 0 | 0 | 0 |
| G5 | 0 | 1 | 0 | 1 | 0 | 0 | 1 | 0 |
| G6 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 |
| G7 | 1 | 4 | 1 | 3 | 2 | 1 | 0 | 0 |
| G8 | 0 | 1 | 0 | 1 | 1 | 0 | 1 | 0 |
| G9 | 0 | 1 | 0 | 0 | 0 | 2 | 1 | 0 |
| G10 | 0 | 3 | 0 | 1 | 0 | 0 | 0 | 0 |
| G11 | 0 | 2 | 0 | 1 | 0 | 0 | 0 | 0 |
| G12 | 1 | 1 | 1 | 1 | 2 | 1 | 0 | 0 |
| G13 | 2 | 2 | 0 | 5 | 1 | 0 | 0 | 1 |
| G14 | 1 | 0 | 0 | 0 | 0 | 0 | 0 | 0 |
| Totals=60 | 8 | 16 | 3 | 17 | 7 | 4 | 3 | 2 |
Multiple genotypes in one or both genes were detected in at least 14% of samples upon first analysis (rising to 29% when repeat experiments were included), distributed among immunocompetent and immunocompromised individuals. More than one genotype was detected in 11% of European samples, 16% of Gambian samples, 47% of South African samples, and 10% of Hong Kong samples (rising to 24%, 33%, 60%, and 60%, respectively, when repeat experiments are included).
This study focused on the genotype definitions, frequencies, occurrence in mixed infections, geographical distribution and evolution (in terms of linkage disequilibrium and mode of selection) of two hypervariable HCMV genes, UL146 and UL139. Totals of 182 UL146 and 183 UL139 sequences were obtained from a large panel of clinical isolates collected from Africa (South Africa and The Gambia), Asia (Hong Kong), Australia and Europe (various countries). These were used in all analyses, and were supplemented by 168 previously published UL146 and 117 UL139 sequences in analyses of genotype definitions, frequencies and mode of selection.
The UL146 sequences fell into the 14 genotypes described previously [
The UL139 sequences grouped into eight genotypes.A recent analysis of 26 clinical samples [
A region of sequence identity (SETTTGTSSNSS in
Studies of HCMV genotype frequency, including the present one, are usually based on the use of conserved PCR primers, and face limitations as a result. Firstly, there is no guarantee that all genotypes will be detected, since primers are chosen on the basis of alignments of available sequences. Secondly, samples containing more than one strain yield mixed sequences,which when cloned are recovered approximately in proportion to their abundance (although stochastic processes may introduce bias during PCR). Therefore, the absence of a genotype from a particular sample cannot be assured. If anyUL146 or UL139 genotypes have escaped recognition, they may emerge from future studies involving different primers or from whole genome sequencing exercises.
As found in previous studies [reviewed in
The occurrence of mixed infections is being recognized increasingly as potentially significant to the biology of HCMV. This feature adds to the limitations inherent in studies of whether particular genotypes are associated with disease outcome; other features include the number, origin and pathological categorization of samples, the choice of gene, the absence of linkage disequilibrium, and host factors. In light of these limitations, our opinion is that robust evidence in favor of any association between genotype and pathology has proved elusive in the literature. Further work utilizing genotype–specific approaches is required to explore the true frequency of mixed infections, both to validate studies of this type and to determine whether mixed infections have geographical or biological correlates.
Similar to the conclusions drawn from a study on UL73 (encoding gN) [
The extensive divergence between genotypes and the consequent inability to produce reliable sequence alignments for both UL146 and UL139 in the hypervariable regions compromised assessments of the role of positive selection in generating the genotypes. In contrast, variation within genotypes is low, and identical nucleotide sequences were obtained from geographically distant individuals. The analysis suggests that constraint has been the predominant factor in evolution within genotypes, with positive selection detected only marginally. A previous study [
IJK was a recipient of a FEMS Research Fellowship and a FEMS–ESCMID Joint Fellowship, and KRA was a recipient of a DAAD Fellowship (German Academic Exchange Service). We thank Mark Schleiss for providing the virus from which one of the samples (a BAC) was generated. We are grateful to Duncan McGeoch for comments on a draft of the manuscript.