This is an open access article distributed under the terms of the Creative Commons Attribution License (
RACE sequencing of ENCODE regions shows that much of the human genome is represented in poly(A)+ RNA.
Recent studies of the mammalian transcriptome have revealed a large number of additional transcribed regions and extraordinary complexity in transcript diversity. However, there is still much uncertainty regarding precisely what portion of the genome is transcribed, the exact structures of these novel transcripts, and the levels of the transcripts produced.
We have interrogated the transcribed loci in 420 selected ENCyclopedia Of DNA Elements (ENCODE) regions using rapid amplification of cDNA ends (RACE) sequencing. We analyzed annotated known gene regions, but primarily we focused on novel transcriptionally active regions (TARs), which were previously identified by high-density oligonucleotide tiling arrays and on random regions that were not believed to be transcribed. We found RACE sequencing to be very sensitive and were able to detect low levels of transcripts in specific cell types that were not detectable by microarrays. We also observed many instances of sense-antisense transcripts; further analysis suggests that many of the antisense transcripts (but not all) may be artifacts generated from the reverse transcription reaction. Our results show that the majority of the novel TARs analyzed (60%) are connected to other novel TARs or known exons. Of previously unannotated random regions, 17% were shown to produce overlapping transcripts. Furthermore, it is estimated that 9% of the novel transcripts encode proteins.
We conclude that RACE sequencing is an efficient, sensitive, and highly accurate method for characterization of the transcriptome of specific cell/tissue types. Using this method, it appears that much of the genome is represented in polyA+ RNA. Moreover, a fraction of the novel RNAs can encode protein and are likely to be functional.
Recent studies [
In addition to the diversity of transcripts from known loci, it appears that much more of the human genome is transcribed than was previously appreciated. Probing of tiling arrays with cDNA probes has indicated that there are at least twice as many transcribed regions of the human genome than had previously been annotated [
The different cDNA and tiling array studies to analyze transcription have also revealed extensive antisense transcription in mammalian genomes [
These various studies have raised many more questions than have been answered. How much of the human genome produces transcripts that are present in the mRNA population? What is the nature of the transcripts produced by the novel transcribed regions? What fraction of novel transcribed regions is likely to be protein coding? What is the level of transcripts produced from the novel transcribed regions? Finally, how much antisense transcription occurs in human cells?
In an effort to address some of these questions and thereby better characterize the human genome and its gene annotation, we have systematically analyzed the transcribed loci in 420 selected portions of the ENCyclopedia Of DNA Elements (ENCODE) regions using 5'-RACE and 3'-RACE sequencing. The ENCODE regions are 44 regions that comprise 1% of the human genome and have been highly characterized with respect to transcripts and transcription factor binding [
We have studied the transcripts produced from annotated gene regions, novel TARs previously identified by high-density oligonucleotide tiling arrays, and regions that were not previously shown to be transcribed (nonTx regions) using 5'-RACE and 3'-RACE and DNA sequencing [
Summary of RACE sequencing using polyA+ and total RNA from human cell lines and tissue
| Experiment | Number of exon primers | Number of novel TAR primers | Number of nonTx primers | Number of sequence reads | Number of detected transcripts on the genome |
| 1: NB4 total RNA | 34 | 39 | 0 | 291 | 154 |
| 2: Hela polyA RNA | 0 | 59 | 0 | 273 | 112 |
| 3: placenta total RNA | 32 | 20 | 44 | 195 | 85 |
| 4: placenta polyA RNA | 0 | 96 | 96 | 591 | 147 |
nonTx, region not previously shown to be transcribed; RACE, rapid amplification of cDNA ends; TAR, transcriptionally active region.
In total, 420 regions were analyzed; primers to each strand were designed and subjected to 5'-RACE and 3'-RACE reactions for a total of 1,680 reactions. Approximately 80% of the reactions generated products that were detected by gel electrophoresis (see Additional data file 1 for examples); 25% of these reactions yielded heterogeneous products (smears). The entire PCR reaction was subjected to DNA sequence analysis, and approximately 40% of the sequence reads mapped to the expected locations of the genome and were therefore deemed as products derived for the intended locus (see Materials and methods, below, for details regarding mapping of RACE sequences to the genome and the fitness score assignment). The average length of these sequence reads is 516 base pairs (bp). As expected, primers designed in known exons gave the highest proportion of valid RACE products. This is followed by the primers designed to the novel TARs. The nonTx regions gave the fewest RACE products (Figure
Frequency of PCR products obtained from different genomic regions. Primers designed to the sense and antisense strands of exons, novel transcriptionally active regions (TARs) and nontranscribed regions were used to generate rapid amplification of cDNA ends (RACE) products. The frequency of PCR products obtained is indicated. nontx, region not previously shown to be transcribed.
We first analyzed the RACE sequences from eight known gene loci. For six of these loci we analyzed RNA from cells in which the gene was known to be expressed. For two genes, 5'-RACE and 3'-RACE reactions were performed using primers designed to the forward and reverse strand of each exon. For an additional four genes we analyzed a subset (1 to 8) of exons in the gene. As shown in Figure
Distribution of RACE product sequences in the
We also analyzed expression of two gene loci, namely
RACE sequencing can detect transcripts not previously detected by microarray analysis in NB4 cells.
To gain a better understanding of why the
The novel RNA isoforms from annotated genes were examined for their ability to produce novel protein isoforms. The 16 novel RNAs identified in this study can produce five novel protein isoforms.
Antisense transcription plays diverse and important biologic roles, and recent studies using reverse transcription based approaches have reported a large amount of antisense transcription in the human genome [
To investigate further whether the antisnese transcript may be an artifact due to reverse transcription, we employed a novel strategy, namely direct chemical labeling of RNA followed by strand-specific oligonucleotide tiling microarray analysis. As shown in Figure
In addition to examining annotated genome regions, we analyzed a large number of novel TARs by RACE sequencing in order to gain a better understanding of their structure, their connectivity to known genes, and whether they might encode proteins of significant length. In all, 856 RACE reactions were generated to 214 TARs of the ENCODE regions [
RACE products from novel TARs and nonTx regions.
The majority (85%) of the RACE sequences from the TARs and nonTX regions map contiguously (without introns) to the genomic sequence. Products from primers that lie close together on the genome often overlap one another or known exons, suggesting extensive transcription throughout the entire region. In addition, whereas the RACE sequences derived from known exons are mostly connected with known exons, the sequences from nonTx regions are rarely connected to others (Figure
Features of the RACE products.
Approximately 16% and 11% of the products produced from TARs and nonTx regions, respectively, produce transcripts that are spliced with consensus GT-AG, GC-AG, or AT-AC splice sequences (see Consensus splice site analyses [under Materials and methods, below]; Figure
In order to determine better whether the novel transcripts may be functional, we examined their ability to encode protein. The sequences of RACE products were analyzed with respect to whether they contain open reading frames (ORFs) and/or whether the potential protein coding sequences are homologous to those in the nonredundant protein database. For two spliced sequences and 25 unspliced sequences, potential ORFs were found that have at least 50 codons, and the predicted protein sequence was homologous to that of a known protein present in the nonredundant database with a BLASTX threshold score of 1 × e-9 [
One example of a potential protein coding transcript is shown in Figure
Example of a novel transcript detected by RACE sequencing.
We examined the expression level of novel transcript 5NGSP2F8 using real-time quantitative PCR. The 5NGSP2F8 expression level is more than 1,000-fold lower than that of the
Even though it is estimated that only 20,000 to 25,000 protein coding genes exist in the human genome, the transcriptome is quite complex and contains protein coding, nonprotein coding, alternatively spliced, and antisense genes [
In addition to many annotated exons, high-density oligonucleotide tiling arrays has identified a large number (8,958) of novel TARs located in both intronic regions and intergenic regions distal from previously annotated genes [
Although many of the novel RNAs do not have long ORFs, a subset of them do (about 9%). From our limited study we found 27 protein coding sequences that are not present in RefSeq but are likely to encode proteins based on the presence of a more than 50-codon ORF that is homologous to other proteins in GenBank. A small fraction (two out of 27) of these is spliced. Additional studies of the entire human genome are thus likely to expand the number of protein coding genes accordingly.
Complementary natural antisense transcripts exert control at many steps of gene expression in prokaryotes and higher eukaryotes from transcription to translation, including transcript initiation, elongation, mRNA processing, location, and stability [
RACE sequencing was able to uncover novel transcripts from nontranscribed regions where microarray experiments did not detect any transcription, indicating the RACE sequence is more sensitive. This is probably due to the fact that micorarray signals are dampened by cross-hybridization to short oligonucleotides on the array. This problem is especially acute for genes that have homologous pseudogenes and paralogs. RACE sequencing offers several other advantages relative to microarrays. Microarrays do not provide information about transcript structure, splicing patterns, or the ability of these regions to encode proteins. Only sequencing full-length cDNA can resolve these issues. The recent developments of massively parallel sequencing technology has the potential to expedite this process greatly [
As noted above, quantitative measurements of transcript expression reveals that two known genes (
Our study highlights the enormous complexity of the human transcriptome and the vast amount of RNA transcripts generated both from alternative splicing and protein coding and nonprotein coding RNAs. The ability of RNA to encode protein and to serve a structural and regulatory role makes it a diverse molecule for mediating many functions. The remarkable complexity of RNAs of the human transcriptome coupled with their diverse functions may therefore help explain the dramatic increase of complexity in higher eukaryotes and phenotypic variation [
The regions of our analysis are selected mainly from the chromosome 22 ENCODE region, with additional targets in chromosome 11 and 21 ENCODE regions. Except for a few regions for test purposes, we selected most of the exon and novel TAR primer regions from among those expressed (cell type specific) regions in known exons and novel TAR regions detected by transcriptional tiling array experiments. The nontranscribed primer regions are selected in a tiled manner from among those regions that are neither known exons nor novel TARs.
We designed four primers for each targeted region, which can be exons of known gene, TAR, or previously identified untranscribed regions. Two gene-specific primers (GSP1 and GSP2) and two nested GSPs (NGSP1 and NGSP2) on both plus and minus strand were selected for each targeted region using a modified Primer3 program. The primers are 23 to 28 nucleotides long, with GC content of 50% to 70% and with Tm (melting temperature) above 70°C (optimally 73°C to 74°C). Self-complementary primers that could form hairpin were avoided. We also voided complementarity between GSPs and UPM (universal primer A in the SMART RACE™ kit [Clontech, Mountain View, CA, USA]), particularly in their 3' Ends (UPM long: 5'-CTAATACGACTCACTATAGGGCAAGCAGTGGTATCAACGCAGAGT-3'; UPM short: 5'-CTAATACGACTCACTATAGGGC-3'). Complementarity between NGSPs and NUP (nested universal primer A), particularly in their 3' ends, was avoided (NUP: 5'-AAGCAGTGGTATCAACGCAGAGT-3'). The primers were mapped against the genome to ensure that they mapped to only one location (with identity <80% to other locations).
Human NB4 cell line total RNA, Hela S3 polyA+ RNA, placenta total RNA, and polyA+ RNA (Ambion, Austin, TX, USA) were used in cDNA amplification by SMART RACE™ kit (Clontech), in accordance with the manufacturer's instructions [
We first use the BLAT alignment tool [
Where parameters such as
For those BLAT matches with multiple blocks, the corresponding splice sites in the transcripts were further examined in the following way. A splice site is defined as a consensus one if and only if a 'GT-AG' (or 'GC-AG'/'AT-AC', which appear much less often) pattern can be observed within windows of eight nucleotides on the two ends of it. (For example, for a splice site starting at chromosome position
Normalized signal intensities from across tiling array experiments were extracted for those primer regions and correspondingly assigned to the transcripts. These signal intensities were correlated with different transcript characteristics such as splicing events in our analysis.
The 'valid' transcripts were also compared against RefSeq gene annotation [
We consider a RACE sequence (either a single block one or with consensus splice sites) a 'novel transcript' if it is not connected to any RefSeq genes. We consider it a 'novel isoform' of a known gene if it overlaps with a known gene and has at least 50 bp not covered by existing annotation. We then compared all such novel transcripts to the nonredundant database using BLASTX [
Human NB4 cell line total RNA or placenta polyA+ RNA (Ambion) were used to make 5'-RACE-Ready cDNA, as described above. Real-time quantitative PCR experiments were performed in quadruplication using LightCycler® 480 Probe Master or TaqMan® Universal PCR Master Mix according to the manufacturer's instructions on a LightCycler® 480 system (Roche Applied Science, Indianapolis, IN, USA). Human
Total RNA and cDNA from human NB4 cells was chemically labelled with biotin using ULS reagent from Kreatech (Amsterdam, The Netherlands) for total RNA and LabelIT reagent from Mirus Bio (Madison, WI, USA) for cDNA. Five micrograms of total RNA and cDNA per array hybridization was incubated with labeling reagent for 30 minutes at 85°C and 60 minutes at 37°C, respectively. Samples were then purified with Qiagen PCR purification columns (Qiagen, Valencia, CA, USA) and ethanol precipitation, respectively. Labelled samples were hybridized to Affymetrix (Santa Clara, CA, USA) ENCODE 1.0 oligonucleotide tiling microarrays. Each sample was hybridized in triplicate to both the forward-strand and reverse-strand version of the array, using the manufacturer's standard hybridization, staining, and washing protocols. The arrays were scanned on an Affymetrix 7G Plus GeneChip scanner, and the signal intensity data were processed using a sliding window of 101 bp.
bp, base pairs; ENCODE, ENCyclopedia Of DNA Elements; EST, expressed sequence tag; GSP, gene-specific primer; NGSP, nested gene-specific primer; nonTx region, region not previously shown to be transcribed; ORF, open reading frame; RACE, rapid amplification of cDNA ends; RT-PCR, reverse transcription polymerase chain reaction; SAGE, serial analysis of gene expression; TAR, transcriptionally active region; UPM, universal primer A.
Experiments were designed by JQW with suggestions from MS. Experiments were performed by JQW. Bioinformatics analyses were performed by JD; JR and ZZ helped with data analyses. AEU and GE contributed to direct labelling of total RNA. Experiments were performed in the laboratory of MS and SW. Bioinformatics analyses were performed in the laboratory of MG. All authors read and approved the final manuscript.
The following additional data are available with the online version of this paper. Additional data file
Shown are examples of RACE PCR products on an agarose gel.
Click here for file
The scores are computed using a subset (from one experiment) of our first set of RACE sequences, as described in the first row of Table
Click here for file
Further explanation of consensus splice site analyses : In order to decide the window size for the consensus splice site analysis, we considered a simplified model in which a nucleotide sequence of length
Click here for file
This file can be uploaded to the University of California at Santa Cruz Genome Brower to view all RACE products.
Click here for file
We thank Janine Mok and Dan Gelperin for critical reading of the manuscript and Jin Lian for NB4 total RNA. We thank Kenneth Nelson and Rajini Haraksingh for technical assistance. We acknowledge the members of the Snyder laboratory for help and support. JQ Wu is supported by NIH Ruth L Kirschstein National Research Service Award and an NIH training grant. M Snyder and M Gerstein are partially supported by the grants from the NIH.