This is an Open Access article distributed under the terms of the Creative Commons Attribution License (
Reverse transcription PCR (RT-PCR) is widely recognized to be the gold standard method for quantifying gene expression. Studies using RT-PCR technology as a discovery tool have historically been limited to relatively small gene sets compared to other gene expression platforms such as microarrays. We have recently shown that TaqMan® RT-PCR can be scaled up to profile expression for 192 genes in fixed paraffin-embedded (FPE) clinical study tumor specimens. This technology has also been used to develop and commercialize a widely used clinical test for breast cancer prognosis and prediction, the Onco
RNA was extracted from formalin fixed paraffin embedded (FPE) tissue, as old as 28 years, from 354 patients enrolled in NSABP C-01 and C-02 colon cancer studies. Multiplexed reverse transcription reactions were performed using a gene specific primer pool containing 761 unique primers. PCR was performed as independent TaqMan® reactions for each candidate gene. Hierarchal clustering demonstrates that genes expected to co-express form obvious, distinct and in certain cases very tightly correlated clusters, validating the reliability of this technical approach to biomarker discovery.
We have developed a high throughput, quantitatively precise multi-analyte gene expression platform for biomarker discovery that approaches low density DNA arrays in numbers of genes analyzed while maintaining the high specificity, sensitivity and reproducibility that are characteristics of RT-PCR. Biomarkers discovered using this approach can be transferred to a clinical reference laboratory setting without having to re-validate the assay on a second technology platform.
Over the last decade, many studies have applied gene expression analysis to identify biomarkers for prognostic and/or predictive information in relation to human disease [
With the development of automated liquid handling and DNA microarrays, high throughput screening using hundreds of samples and hundreds or thousands of genes has become routine in many laboratories. DNA microarrays offer the advantage of simultaneously assessing the relative expression level of thousands of genes with a relatively small amount of starting RNA. However, DNA microarray measurements are limited in dynamic range, specificity and reproducibility, leading to high false positive and false negative biomarker discovery rates. As currently configured, DNA microarray technology also requires high quality RNA. Alternatively, reverse-transcription polymerase chain reaction (RT-PCR) technology offers the advantages of high accuracy and reproducibility, and precise quantitation over a wide dynamic range. To overcome the issue of fragmented RNA in FPE tissue specimens, assays can be optimized for short amplicons so the RNA from FPE tissue can be successfully analyzed [
It has been suggested that there is a bottleneck in scaling up TaqMan® RT-PCR using archival FPE samples to analyze beyond 30 genes [
In this paper we report the use of this high throughput, highly parallel TaqMan® RT-PCR process to screen RNA extracted from colon cancer FPE clinical trial specimens (NSABP C01 and C02 studies). We focused on selecting candidate genes known to be involved in pathways related to colon cancer and genes from published expression profiling data sets relating to colon cancer prognosis and response to therapy. Our results indicate that this approach yields high quality expression data that can be used for simultaneous evaluation of hundreds of candidate genes in defined cohorts of patients to identify prognostic and predictive colon cancer biomarkers.
FPE RNA is largely fragmented, making it a poor substrate for poly-dT primed reverse transcription. Random priming in combination with poly-dT priming has proven to be somewhat more successful for generating cDNA from FPE RNA [
While measuring the realistic limit of multiplexed RT-priming in FPE RNA we did not want the data complicated by biological variability contributed by individual samples. We therefore created a sample consisting of pooled FPE RNA that represented a range of high and low expression values across the gene panel and reflected the quality of samples that would be used in the gene identification study. This was done by pooling 36 breast FPE RNA samples from tumor blocks over 10 years old. Two of the 8 primer sub-pools were selected at random to verify that the priming reaction remained consistent as the RNA template type changed from a high quality sample to a highly fragmented sample. The results, shown in Figure
Once we confirmed the high complexity priming reaction was working consistently in FPET RNA, we prepared a gene specific primer pool containing 761 unique reverse primers for the biomarker discovery study using NSABP C01 and C02 clinical trial specimens.
To be included in the final study analysis, samples had to pass pathology, clinical and laboratory data QC requirements. To meet pathology acceptance criteria, a minimum of 5% of the tissue present in each sample was required to be invasive cancer cells. All samples were dissected to enrich tumor tissue and minimize non-tumor elements.
RNA was extracted from 354 FPE tissues cut from the NSABP C01/C02 clinical study samples (archived from 1977–1983) [
Sample exclusion criteria
| Number excluded | Percent excluded | Remaining samples | |
| Samples extracted | N/A | 0 | 354 |
| Insufficient RNA | -38 | 10.7% | 316 |
| Incomplete sample data | -4 | 1.1% | 312 |
| Unsatisfactory qPCR | -21 | 5.9% | 291 |
| Pathologically ineligble | -10 | 2.8% | 281 |
| Clinically ineligble | -11 | 3.1% | 270 |
|
|
76.3% | 270 |
Samples were excluded on the basis of several different categories. Various acceptance criteria were applied to the dataset. The table shows how many samples were excluded from the final dataset based on each category. There were 354 samples extracted and the final evaluable data set was 270.
Using a semi-automated extraction process, two lab technicians extracted RNA from 48 samples per day. Paraffin was removed from the tissue manually with xylene, followed by ethanol washes. After a proteinase K digestion and phenol-chloroform purification, the remaining RNA extraction steps were performed on an automated liquid handler using a plate-based protocol. The purified RNA was then quantified using an automated RiboGreen fluorescence assay (Invitrogen, Carlsbad, CA.). The resulting files containing the RNA quantification data were automatically collected by the laboratory information management system (LIMS). The LIMS used those data to generate an RNA concentration normalization work-list, which was then executed by an automated liquid handler during the reverse transcription reaction assembly procedure. The reverse transcription reaction was completed using an MJ Research thermocycler (Bio-Rad, Hercules, CA). Finally, quantitative PCR reactions were assembled at a rate of 64 plates (32 patient samples) per day. All automated liquid handlers were obtained from Tecan (TECAN Schweiz AG, Männedorf, Switzerland). This system yielded a total of 24,576 real time quantitative PCR reactions to be performed per day using four ABI PRISM 7900 HT instruments (Applied Biosystems, Foster City, CA.).
Process reproducibility was monitored throughout the study. Two Tecan Genesis liquid handling robots were used to assemble plates for 335 patient samples over a period of five weeks, for a total of 670 assay plates (this includes samples run with an RNA concentration of less than 1 ng/assay well). Given the large scale of the project and the number of potential sources of variability, we monitored the process performance and reagent stability throughout the duration of the study. This was done by using a single RNA reference sample made by pooling RNA extracted from 80 FPE colon specimens (block age ranging from 3 years – 6 years). As described previously, a pooled FPE sample was selected as a reference RNA since it more representative of a typical study sample than commercially available RNA, yet sufficiently abundant to provide a stable reference baseline throughout the study period. This sample was analyzed 3 times on the same ABI 7900 HT instrument to set a baseline reference profile for the study genes and then was assayed at intervals throughout the duration of the study. Figure
Figure
The four ABI PRISM 7900 HT machines used for the study underwent qualification to meet internal performance specifications prior to use in clinical studies. Figure
Two 384 well plates were prepared per patient RNA specimen to screen all 761 candidate marker genes. Assays were randomly assigned between the two plates except for reference genes which were present on both assay plates. The average CT values for each plate pair were compared and the data are illustrated in Figure
Although the two assay plates for each patient RNA were run on the same ABI Prism 7900 instrument to minimize process variability, data from each sample had to be internally normalized to allow all study specimens to be compared without being confounded by relative variability in RNA quality, quantity or process variability. This was accomplished by subtracting the averaged expression values from 6 reference genes (CLTC, NEDD8, RPLPO, RPS13, UBB, UBC) from the expression values for each gene in each sample. This method has previously been shown to effectively compensate for variability associated with RNA degradation in FPE material of different ages and qualities [
A high degree of reproducibility was seen for these reference genes, as shown by the RPLPO example, in Figure
As an indication of the robustness of this high complexity assay we sought to identify co-expressed groups of genes, that are plausible based on known biological pathways. Unsupervised cluster analysis was performed on the final sample set using all 761 genes, resulting in several distinct clusters representing known biological pathways. Figures
After data quality control and reference normalization, approximately 19% of the cancer related genes tested in this study were found to have a significant (p < 0.05) correlation with Recurrence Free Interval (RFI) by univariate regression analysis [
The development of a clinically validated test that could determine the risk of recurrence or death from stage II/III colon cancer and the likelihood of benefit from standard chemotherapy regimens is highly desirable but complex. The process begins with biomarker discovery and ends with a clinical validation study with prospectively defined endpoints. Since the method by which an mRNA species is measured will have a profound effect on the success of such a validation study, it is important to characterize and maintain the assay's performance, particularly its reproducibility and quantitative precision. Consequently, we have adopted RT-PCR, the most robust gene expression method available for gene discovery studies. Although nearly 150 genes were found to be significantly related to RFI in this study, some markers will prove to be false positives and true markers will vary in how robustly they correlate with outcome. It is therefore important to evaluate the candidate genes identified here by conducting further independent studies to identify truly useful disease biomarkers. Only after consistent association with clinical outcomes in multiple independent studies should genes be considered for inclusion in an assay used to make clinical decisions. Employing a single technology consistently throughout biomarker discovery and into clinical testing has the advantage of reducing the time required to fully validate and commercialize a multi-gene clinical decision-making tool.
RT-PCR is often carried out using oligo-dT priming to generically reverse transcribe mRNA from the polyA tail. However, this technique is unsuccessful with degraded RNA, such as that extracted from FPE tissue. We have previously shown that RT-PCR using gene-specific priming can be successfully applied to FPE tissue as old as 30 years [
DNA microarrays are a popular technology for biomarker discovery because one can quickly examine the expression of hundreds or thousands of genes. RT-PCR has often been subsequently used to verify the results of microarray data, since it offers much higher sensitivity, specificity, reproducibility and a greater quantitative dynamic range. The present results demonstrate that RT-PCR can also be applied to highly parallel gene expression analysis if robotic processes and assay miniaturization are used. We were able to extract and quantify RNA and generate expression data for 761 unique assays for more than 300 patients in less than 5 weeks.
While screening these patient specimens, it was important to monitor potential sources of variability such as primer and probe stability and gene specific primer pool stability. We used an FPE colon RNA pool as a reference sample and generated a baseline CT value for each of the 761 assays. This FPE colon RNA pool reference sample was then included during reverse transcription with every patient sample batch and used to monitor process stability throughout the study. Over the 5 week period, variability within the reference sample remained low, indicating that all patient samples were being analyzed with a stable assay process. Analysis of one of the reference genes, RPLPO, which was assayed on both plates for every patient sample, also highlighted the internal consistency of the process throughout the study.
The robustness of this technology is evidenced by results from hierarchical clustering of all 761 genes which identified known pathways and gene group clusters that one would expect to be co-expressed. One of the largest was a "stromal response" gene group containing genes that are associated with wound healing and are thought to be representative of fibroblast activation, or the 'stromal response' within tumor stroma. Stromal response is becoming increasingly recognized as a marker of invasion and poor clinical outcome in several different classes of solid tumors [
The aim of this study was to identify gene biomarkers that predict recurrence-free interval in patients with Stage II and Stage III colon cancer. Approximately 19% of the 761 genes showed a significant (p < 0.05) association with RFI by univariate Cox proportional hazards regression analysis. It is highly unlikely that any one gene will be able to predict clinical outcome or response to therapy to the extent that it will be useful to oncologists. A successful diagnostic tool is much more likely to consist of a panel of genes and an algorithm weighting and combining each gene contribution into one value that defines the unique risk of recurrence and potential for therapeutic response in each patient. This concept is supported by the observation that several different biological pathways were shown to be associated with RFI in this study.
We have demonstrated that RT-PCR can be scaled to enable studies testing hundreds of candidate genes for biomarker discovery in hundreds of archival FPE cancer biopsy specimens. Analysis of the data from this study has shown it to be biologically plausible and consistent with known pathways and gene groups identified to be important in cancer. We are applying this technology to colon cancer with the aim of developing a predictive and prognostic clinical test for patients with this disease.
Archival colon tumor FPE tissue blocks were provided by the NSABP from both the C01 "A Clinical Trial To Evaluate Postoperative Immunotherapy And Postoperative Systemic Chemotherapy In The Management Of Resectable Colon Cancer" and C-02 "A Protocol To Evaluate The Postoperative Portal Vein Infusion Of 5-Fluorouracil And Heparin In Adenocarcinoma Of The Colon" clinical trials. Patients were enrolled in these trials between 1977–1983. Samples used in this study are representative of the general study populations for both trials.
Histotechnologists wore gloves at all times when handling tissue blocks. Before and after each block was sectioned, any debris was removed from the microtome with a disposable cotton swab or brush. All appliances (brush, forceps, knife, knife holder base) were wiped with an RNase Zap wipe followed by a soft cloth wetted with de-ionized water. A tissue floatation bath was used to help eliminate wrinkles and distortions in sections being mounted to glass microscope slides. When dissection was required to remove significant non-tumor elements, a representative H&E stained slide was used as a guide to mark tumor and non-tumor portions of three 10 micron unstained slides. Tumor tissue was then scraped away from the non tumor material and placed into an extraction tube. The tumor tissue from all 3 sections was placed into the same tube.
RNA was extracted from the tumor-enriched portion of three 10 micron sections per patient block. In order to scale sample throughput to 48 samples/batch, the RNA extraction procedure used a semi-automated method performed on a TECAN robotic liquid handler (TECAN Schweiz AG, Männedorf, Switzerland). The samples originated in individual 1.5 ml Eppendorf tubes and paraffin was removed by incubating with lab grade xylene for 5 minutes. Tissue was then pelleted by centrifugation at room temperature (+18°C to +25°C) for 5 minutes at approximately 14,000 RPM. The xylene was removed and the procedure repeated. The tissue pellet was washed by inverting several times with 200 proof ethyl alcohol and again pelleting by centrifugation at room temperature. The ethyl alcohol wash step was repeated. Immediately prior to adding Proteinase K, samples were inspected for residual alcohol; if any alcohol was visible, it was aspirated without disturbing the tissue pellet. Proteinase K digestion was performed using reagents from the MasterPure® Purification kit (Epicentre, Madison, WI). Samples were incubated at +65°C for 2 hr with Proteinase K. Protein and genomic DNA were removed by manual addition of an equal volume of acid-phenol: chloroform and the removal of the upper aqueous phase after centrifugation for 5 minutes at approximately 10,000 RPM. Samples tubes were transferred to a TECAN liquid handler where purification was completed using the mirVana™ RNA purification kit (Ambion, Austin, Texas) on a 96 well glass fiber filter plate. Purification was followed by DNase I treatment on the same 96 well filter plate.
RNA was quantified using the RiboGreen fluorescence method as described by the kit manufacturer (Invitrogen, Carlsbad, CA.).
Candidate genes were obtained from published gene expression profiling data relating to colon cancer prognosis and response to therapy, biological pathways known to be important in cancer [
The reference sequence for each gene included in the study was obtained from the NCBI Entrez website. TaqMan® RT-PCR primers and probes were designed using an automated in-house primer design module. The complete list of assay primers and probes is shown in Additional file
Reverse transcription was performed using the Omniscript kit (Valencia, CA) for RT-PCR. For initial investigations, reverse primers (each at 100 μM) were collected into sub-pools of 94–96 primers. An aliquot from each sub-pool was added together to create the final GSP pool (each primer at 100 nmol/L). For the clinical study, reverse primers (each at 1 mM) were first collected in sub-pools of 89–96 primers and tested for priming performance before being combined into a master 761 gene-specific primer pool (each primer at 1 μmol/L). For both the initial investigations and the clinical study, each primer in the RT reaction was at a concentration of 50 nmol. Our standard procedure was to add RNA to the RT reaction at 12.5 ng/ul which equates to 1 ng/well (cDNA) for quantitative PCR. In each case, the RT reaction was performed in a single tube with the GSP pool (ie up to 761 reverse primers). The resulting cDNA was distributed equally among the wells of a 384 well plate, and the appropriate forward and reverse primer and probe were added to each assay well.
For each patient sample, two 384-well plates were used. Assays for 7 potential reference genes were included on both plates, with all other gene assays randomly distributed in single assay wells. RT-PCR assays for three K-ras gene mutations and one BRAF gene mutation were included in the assay panel along with each corresponding wild type allele assay. Therefore, 757 normal gene alleles were assayed and 4 mutant genes were assayed, bringing the total number of unique assays to 761. TaqMan® RT-PCR was performed according to instructions of the manufacturer, using Applied Biosystems Prism (ABI) 7900 HT instruments. Reactions were performed in a 5 μl volume with cDNA equivalent to 1 ng total RNA. Final primer and probe concentrations were 0.9 μmol/L (primers) and 0.2 μmol/L (probe). For the K-ras mutation assays a blocker oligomer was added to the primer and probe pool at a final concentration of 3.6 μmol/L (mutant 1) 3.6 μmol/L (mutant 2) and 12.95 μmol/L (mutant 3). These blockers are added to inhibit amplification of the non-mutant allele, permitting specific amplification of the mutant allele. PCR cycling conditions were 95°C for 10 minutes for one cycle, 95°C for 20 seconds, and 60°C for 45 seconds for 40 cycles. A reference sample (pooled colon FPET RNA) was assayed throughout the study to ensure reagent and process stability. As a negative control, wells without any template were also assayed every two weeks to ensure that no exogenous nucleic acid contaminations occurred.
The unsupervised hierarchical clustering of genes was performed using 1-Pearson R as the distance measure for gene expression and the un-weighted pair-group average as the amalgamation method [
KMCL lead the study and provided scientific input for the manuscript. JYW and AC prepared reagents, extracted and processed the samples. CS, JRH and GY provided biostatistical input. JLS and AN lead the development of the robotic systems. JB provided scientific input for the manuscript. CK reviewed blocks and slides from NSABP colon protocols and selected the appropriate blocks for processing, and MTC provided scientific input for the manuscript and managed the project.
GY and CK have no competing interests to declare. KMCL, JYW, AC, CS, JRH, JLS, AN, JB and MTC were all full-time employees of Genomic Health at the time this work was done.
Click here for file
Click here for file
Click here for file
We are grateful to the NSABP for providing the FPE tissue samples. We kindly thank Andrew Dei Rossi, Debjani Dutta, Mylan Pho and John Morlan for their assistance with preparation of reagents and samples; Mei-Lan Liu for her helpful advice and suggestions; Robyn Loverro and Kenneth Hoyt for the development of robotics systems, Joel Robertson for IT support; Xitong Li for the primer design module; Jeanne Yue for her assistance in generating the figures and Melanie Finnigan, William Hiller, and Teresa Oeller from Division of Pathology, NSABP for tissue sectioning.