Conceived and designed the experiments: MJL DS AK BA. Performed the experiments: MJL MC SM. Analyzed the data: MC DS BA. Wrote the paper: MJL.
Micro RNAs (miRNAs) are a class of small, non-coding RNA species that play critical roles throughout cellular development and regulation. miRNA expression patterns taken from various tissue types often point to the cellular lineage of an individual tissue type, thereby being a more invariant hallmark of tissue type. Recent work has shown that these miRNA expression patterns can be used to classify tumor cells, and that this classification can be more accurate than the classification achieved by using messenger RNA gene expression patterns. One aspect of miRNA biogenesis that makes them particularly attractive as a biomarker is the fact that they are maintained in a protected state in serum and plasma, thus allowing the detection of miRNA expression patterns directly from serum. This study is focused on the evaluation of miRNA expression patterns in human serum for five types of human cancer, prostate, colon, ovarian, breast and lung, using a pan-human microRNA, high density microarray. This microarray platform enables the simultaneous analysis of all human microRNAs by either fluorescent or electrochemical signals, and can be easily redesigned to include newly identified miRNAs. We show that sufficient miRNAs are present in one milliliter of serum to detect miRNA expression patterns, without the need for amplification techniques. In addition, we are able to use these expression patterns to correctly discriminate between normal and cancer patient samples.
MicroRNAs (miRNA) are single-stranded RNA molecules of about 21–23 nucleotides in length, which function in the regulation of gene expression. miRNAs are expressed as part of primary transcripts in the form of hairpins with signals for dsRNA-specific nuclease cleavage by the ribonuclease Drosha in combination with an RNA-binding protein. After the precursor miRNA is released as an approximately 70 nt RNA, it is transported from the nucleus to the cytoplasm by Exportin-5, and then is cleaved by Dicer RNase III to form a double-stranded RNA. Dicer initiates the formation of the RNA-induced silencing complex (RISC), which is responsible for the gene silencing observed due to miRNA expression and RNA interference
MicroRNAs have been found in tissues and also in serum and plasma, and other body fluids, in a stable form that is protected from endogenous RNase activity (in association with RISC, either free in blood or in exosomes (endosome-derived organelles)). Studies by Lu et al
miRNA signatures from normal and cancerous tissues have been used to classify several types of cancer and may also allow clinicians to determine a treatment course based on the original tissue type. It may also be possible to use miRNA expression patterns as a biomarker to monitor the effect of therapy on cancer progression
We have determined that a sufficient quantity of miRNAs is present in less than 1 ml of human serum to produce a detectable signal on a microarray using fluorescence or electrochemical detection (ECD). Using a simple phenol/chloroform extraction protocol, we recovered approximately 1.3 µg of serum RNA from each 800 µl of serum (average of 18 samples; standard deviation 0.3). The resulting pattern of miRNA expression could be used to distinguish between cancer patients and normal donors.
The approximate size of the small RNAs recovered from plasma was determined by isolating large RNA fragments (low ethanol concentration) and small RNA fragments (high ethanol concentration) using the Invitrogen PureLink miRNA isolation kit, after acid phenol/chloroform extraction and precipitation. The two RNA size fractionations were labeled with biotin (Mirus) and hybridized to a microarray. Results (not shown) indicated that the vast majority of signal was from the small RNA fraction, which was similar to the signal from the un-fractionated sample.
DNA contamination of extracted serum nucleic acid was examined by comparison of an extracted sample that was split, and half treated with DNase I. After labeling both treated and untreated samples with biotin and hybridizing to different sectors of the same 4×2K array, very little difference could be seen between treated and untreated samples (r2 = 0.9 and 0.96 respectively for two replicates) indicating that little DNA contaminates samples (data not shown).
The sensitivity of our miRNA assay was determined by adding dilutions of a synthetic RNA oligonucleotide to our assay during serum extraction. We were able to detect approximately 4,000 copies of serum microRNAs per microliter of serum (
RNA miRNA analog oligonucleotides, at concentrations ranging from 0 to 40 million copies per microliter, were spiked into 400 ul of serum after the addition of RLT buffer. RNA was then extracted from the serum using phenol/chloroform extractions and an ethanol precipitation. Samples were then labeled and hybridized on a microarray. Vertical bars indicate array signal intensities for specific miRNA probes representing the wild type sequence (Wild) and probes with two internal mutations (mut) for (A) oar|miR-431 and (B) oar|miR-127. Scales for the 4,000 and 0 copies data points (boxed in left panels) are expanded in the right panels: (C) oar|miR-431, and (D) oar|miR-127.
We have also determined that data collected from the same serum samples after being frozen at −80°C for 1 week after the initial microRNA assays, was similar to the original data. MicroRNAs from aliquots of 2 serum samples from cancer patients (1 prostate and 1 colon) were extracted, labeled and hybridized to arrays, and after one freeze/thaw event, new aliquots were again extracted, labeled and hybridized to a second array. Data sets, from re-assayed prostate cancer sample 811 and colon sample 792, showed strong correlations when raw array data were compared (r2 = 0.94 and 0.96 respectively). This result indicates that the assay is reproducible and stable over time.
Several prostate, ovarian, colon, breast and lung cancer serum samples as well as normal male and female donor sera have been analyzed on the pan-miRNA microarray to ascertain serum miRNA profiles and to confirm the specificity of the profiles for different cancers and normal donors. Several data analysis methods have been tested to determine the most relevant method for the discrimination of cancer versus normal. For preliminary analysis, we log2-transformed serum miRNA probe signals from a normal donor and compared this data to log2-transformed probe data from a prostate cancer patient and from a prostate cancer cell line. Although both the prostate cancer serum sample and the prostate cancer cell line (22Rv1) sample showed up-regulation compared to a normal serum sample, they did not show much similarity to each other. At this point, we simply note a relative up-regulation of serum miRNAs in cancer as compared to serum from normal donors (
Log transformed normal donor serum miRNA signals (blue line) were compared to miRNA array signals from a prostate cancer cell line 22Rv (open squares) and from a prostate cancer patient (closed diamonds). In general, cancer and cell line miRNAs seem to be up-regulated when compared to normal donor serum miRNAs.
We first set out to define a minimal set of probes that would allow us to discriminate between prostate and normal serum samples. Signal from each miRNA probe was first background corrected using negative control probes. Subsequently, each miRNA probe was expressed as the natural log of the ratio between itself and the same probe in a normal human male serum sample. This gives a value of 0 for all the base serum sample probes and an up or down regulation with respect to that sample (normal) for all the other samples (normal and cancer).
After data set normalization, the natural log of the ratio of the signal for a specific probe over the same probe from the normal serum sample was taken. 15 miRNAs showed up-regulation in all stage 3 and 4 prostate cancer samples when compared to sera from normal male donors. These miRNAs are listed below each data set. Five stage 3 and 4 prostate cancer sera (Yellow), and 8 normal male donor sera (red) were analyzed. Vertical lines indicate plus or minus one standard deviation of the mean.
In
(A) Analysis of signal from normal serum; (B) Analysis of signal from 22Rv1 cell culture; and (C) signal from prostate cancer patient serum. Z-scores (blue lines) were determined by subtracting the signal at each probe by the mean of the test probes from the entire hybridization, and then, by dividing the resulting value by the standard deviation of the signal across test probes, across the entire hybridization.
Hierarchical clustering was used to group samples from different disease and normal states (
Cancer samples and normal donor samples (brackets) were clustered using a hierarchical clustering program to show sample-to-sample relationships. Sample labels include donor condition (cancer type or normal), sample lot number (last three digits), gender, and cancer stage (2 – 4, or 0 for normal). Labels marked with a or b indicate repeat testing of the same sample.
| CANCER PATIENT SERUM SAMPLES | |||||
| LOT NUMBER | GENDER | AGE | STAGE | CANCER TYPE | TREATMENT |
| BRH233781 | Female | 63 | 4 | OVARIAN CANCER | Carboplatin, Taxotere |
| BRH233782 | Female | 66 | 4 | OVARIAN CANCER | Carboplatin, Gemzar |
| BRH233787 | Female | 69 | 4 | NON-SMALL CELL LUNG | Zofr, Deca, Carb, Neul, Veps |
| BRH234151 | Female | 68 | 4 | SMALL CELL LUNG | Topotecan |
| BRH234153 | Male | 62 | 3 | SMALL CELL LUNG | Zometa |
| BRH237121 | Female | 69 | 1 | COLON CANCER | None |
| BRH237122 | Female | 67 | 2 | COLON CANCER | Ferrlecit |
| BRH237123 | Female | 85 | 2 | COLON CANCER | None |
| BRH233792 | Female | 63 | 2 | COLON CANCER | None |
| BRH237120 | Female | 87 | 3 | COLON CANCER | Zometa |
| BRH233796 | Male | 69 | 3 | COLON CANCER | 5FU |
| BRH233798 | Female | 76 | 3 | COLON CANCER | None, Pretreatment |
| BRH249639 | Male | 47 | 4 | COLON CANCER | 5FU |
| BRH234157 | Female | 80 | 4 | BREAST CANCER | Femara, Zometa |
| BRH234158 | Female | 44 | 4 | BREAST CANCER | Xeloda, Zometa |
| BRH233808 | Male | 65 | 4 | PROSTATE | Taxotere, Zometa |
| BRH233809 | Male | 62 | 4 | PROSTATE | Taxotere, Zometa |
| BRH233811 | Male | 72 | 3 | PROSTATE | Lupron, Zometa |
| BRH233812 | Male | 61 | 3 | PROSTATE | None |
| BRH233814 | Male | 59 | 3 | PROSTATE | Taxotere |
| BRH249616 | Male | 64 | 2 | PROSTATE | None |
| NORMAL DONORS | ||
| LOT NUMBER | GENDER | AGE |
| BRH233823 | Male | 40 |
| BRH233824 | Male | 61 |
| BRH233825 | Male | 44 |
| BRH233826 | Male | 42 |
| BRH233827 | Male | 49 |
| BRH233828 | Male | 50 |
| BRH233829 | Male | 41 |
| BRH237135 | Male | 31 |
| BRH237124 | Female | 56 |
| BRH237125 | Female | 22 |
| BRH237126 | Female | 22 |
| BRH237127 | Female | 69 |
| BRH237131 | Female | 48 |
| BRH237132 | Female | 47 |
| BRH237133 | Female | 32 |
To further explore the miRNAs responsible for the clustering, Heat maps were used to look for similarities between miRNA expression patterns within each sample. This method is most effective when rows and columns are ordered to allow these patterns to be easily identified. Clustering was thus used to give this ordering (by identifying miRNAs that have similar expression patterns, and arranging them in close proximity). This data was ported to the open source program, Cluster
The set of miRNAs used for analysis was chosen based on significance in at least 5 hybridizations. A probe-set was judged significant if the ratio of perfect-match (PM)/mismatch (MM) probes was greater than 1.5.For each hybridization in the analysis, only significant signal was used for the clustering. Signal for those miRNAs whose signal was judged significant by PM/MM ratios was Log 2 converted. Then the signal was median normalized over that hybridization and Average Linkage Clustering was performed using a Spearman Rank Correlation. Clustering was visualized using the program TreeView
For each miRNA mature region, two controls were written on the chip: a sense wild-type probe (s), an anti-sense wild-type (PM) form and a double mutant control (MM). This probe set was used to evaluate significance as well as intensity of each miRNA mature form. For each hybridization, the raw signal was extracted and probes were grouped by the miRNA that they were designed for. PM signal was log2 transformed, and Z-normalized. For the classifier and for the focused cancer versus normal hierarchical clustering, we used a simple heuristic to determine if a probe was indeed present. For each miRNA that was evaluated, we counted a signal significant if the signal for PM/MM >1.0. If so, then the Z-normalized log2 (signal) was used for that miRNA, otherwise the probe data for that particular probe was not used. Normalization was thus performed over anti-sense wild-type probes across the chip; however, only data from significant probes were used for clustering and classification.
Data Mining was performed using the WEKA package
We next extracted signal from each hybridization for just this selected set of miRNA's. In
Data Mining was performed using the WEKA package
miRNA microarray data from cancer patients and normal donor sera were coded to remove any indication of disease status and submitted for analysis. Using the subset of attributes described above and in
miRNA expression signatures have a potential role in the diagnosis, prognosis and therapy of human diseases, including cancer, heart disease, viral infections and inflammatory diseases
Several studies have detailed the miRNAs that are associated with cancers
Recently a series of studies has been performed on miRNAs present in serum
One issue that appears in these studies is the fact that the miRNA expression patterns seen in serum are not identical to those seen from miRNAs taken directly from cancer cell lines. In our study, we found that although both the prostate cancer serum sample and the prostate cancer cell line (22Rv1) sample showed up-regulation compared to normal serum sample, they did not show much similarity to each other. This seeming discrepancy could be taking place for a number of reasons. The most obvious of which is the possibility that samples taken from the cell lines themselves are not representative of what appears in the serum. We can speculate that the most obvious source of miRNAs that appear in the serum is a product of tumor cell lysis; however, it may also be possible that their appearance in the serum is the product of a form of active transport involving the formation of exosomes (36). This would confound a direct comparison between miRNA expression patters derived from cell-line and tumor with those that are serum-derived.
Information on the use of miRNAs as biomarkers is predominantly associated with studies on tissue samples or cancer cell lines. Distinct patterns of miRNA expression are able to distinguish between cell type and stage in various cancers. This bodes well for diagnostic and prognostic applications of miRNA profiles. It also indicates there is clearly a need to define the expression profiles of miRNAs in serum of cancer patients and compare these to profiles observed in the serum of individuals representing a range of diseased and healthy states. It is anticipated that miRNA profiles in serum have the potential to be early markers for cancer detection and will also play a role in the monitoring of disease status during chemotherapy.
In this study, we have determined that a sufficient quantity of miRNAs is present in one ml of human serum to produce a detectable signal on a microarray using fluorescence or electrochemical detection. At the simplest level, this study has shown that serum miRNAs are up-regulated in cancer patients as compared to normal donors. In a comparison of stages 3 and 4 prostate cancer sera and normal donor serum miRNA levels, we found that 15 miRNAs (miR-16, -92a, -103, -107, -197, -34b, -328, -485-3p, -486-5p, -92b, -574-3p, -636, -640, -766, -885-5p) were up-regulated in serum from prostate cancer patients compared to normal donor sera.
Heat Map and cluster analyses show that serum miRNA signatures can also be used to separate cancer patients and normal donors in most cases. Sixty-five miRNAs (see
We have also shown that serum miRNAs can be detected at a level similar to that reported for TaqMan PCR from serum, approximately 4,000 copies per ul
Our results, in general, agree in many cases with previously published studies. However, differences in our results from other studies could result from several factors: 1) Serum miRNA expression profiles do not directly correspond to tissue profiles. The low levels of miRNAs in serum are better suited to studies of up-regulation and not down-regulation. In addition, it is unlikely that there is a direct correspondence between tissue miRNA levels and serum miRNA levels due to the possible mechanisms of miRNA release into circulation (cell lysis or exosome release
An example of published data from two different miRNA expression profiling techniques that do not show strong agreement is illustrated in the following comparison. Schetter et al
This study must be confirmed with a larger and better-documented data set. We do not yet know the affects of gender, age and cancer treatment on miRNA levels in serum. Radiation and chemotherapies that result in remission of cancer should also result in a change in the serum miRNA profiles. Wong, et al.
Arrays for serum miRNA analysis were constructed with 547 human miRNA sequences obtained from the Sanger Database version 10.0 which appeared on 8/2/07. Because of limited space on the array, the probe list was modified to exclude a few newer miRNAs that had recently been added. The array includes miRNA probes for all studies referenced by this paper. Three probes were written for each miRNA: an anti-sensed wild-type version, a double-mutant control probe, and a sense control version (
Sequences for miRNA probes were taken from the Sanger database version 10.0 (released, 8/2/2007). We generally used only the predominantly expressed form for each miRNA precursor. This was done to save space on the array. The final list included all of the dominant miRNA forms from the studies referenced in this paper
22Rv1 human prostate cancer-derived cells were cultured in standard plastic tissue culture plates in RPMI medium 1640 (GIBCO) supplemented with 10% FBS and 1% penicillin-streptomycin at 37°C in a 5% CO2 incubator. Cells were harvested in Qiagen RLT buffer and extracted with phenol/chloroform as described below.
Human serum samples were purchased from Bioreclamation, Inc, Hicksville, NY and include: stages 2 to 4 prostate, stages 1 to 4 colon, stage 4 ovarian, stage 4 breast, and stages 3 and 4 lung cancer sera (
An aliquot of 400 µl of each serum sample was mixed with 500 µl lysis buffer (RLT, Qiagen, Valencia, CA) and 800 µl acid phenol: Chloroform (Ambion, Foster City, CA), vortexed for 30 seconds and centrifuged at 16000 rcf for 10 min at 25°C. The aqueous phase was extracted 2× with an equal volume of acid phenol:chloroform and centrifuged at 16000 rcf for 10 min at 25°C. The resulting aqueous phase was then precipitated with 0.1 vol 5 M NaCl, 2 µl precipitation enhancer (Mirus, Madison, WI), 2 µl GlycoBlue (Ambion) and 2.5 vol 100% ethanol at −20°C for at least 1 hr. After centrifugation at 4°C for 30 min, the pellets were washed 2× with 75% ethanol and then air-dried. Precipitated RNA was resuspended in 50 µl molecular grade water (Ambion) and quantified with a NanoDrop ND-1000 spectrophotometer (Thermo Scientific, Wilmington, DE).
Approximately one µg of isolated RNA was labeled with a Mirus miRNA Biotin labeling kit (MIR8450) following manufacturers directions. Briefly, 1 µg RNA was diluted to 86 µl with water and 10 µl of 10
Sectored array chambers (4 chambers per array) (CombiMatrix 4×2K arrays™) were each filled with 30 µl of Pre-Hybridization Solution (CombiMatrix Corp) and incubated for 10 min at 45°C. MicroRNA was mixed with 9 µl 20× SSPE (Ambion), 4.8 µl BSA at 50 mg/ml (Ambion), 3.6 µl deionized formamide (Sigma) and 7.5 µl of 10% SDS (Ambion) and heated to 95°C for 3 min. 30 µl of each sample were added to sectored hybridization chambers, sealed with aluminum tape, and incubated at 45°C for 16 hr with rotation. After hybridization, arrays were washed 2× with 2× SSC with 0.1% SDS at room temperature (RT) for 10 sec (CombiMatrix Corp), 2× with 2× SSC at RT for 10 sec and then washed 1× with 0.2× SSC at RT for 10 sec each.
Arrays were blocked with 5× PBS/Casein Blocking Buffer at RT for 10 min and then labeled with either Cy5 labeling solution for fluorescence scanning or HRP Biotin Labeling Solution (CombiMatrix) for ElectraSense reading (CombiMatrix) and incubated for 30 min at RT. Arrays were then washed 2× with Biotin Wash Solution (2× PBST) for 30 sec each at room temp and again washed 2× with 2× PBS followed by scanning for fluorescence, or washed 2× with TMB Rinse Solution (CombiMatrix), followed by one wash with TMB substrate (CombiMatrix) and scanning with an ElectraSense reader (CombiMatrix) after fresh TMB was added.
To determine the sensitivity of our assay, commercially purchased RNA miRNA analog oligonucleotides (IDT), at concentrations ranging from 0 to 40,000,000 copies per microliter, were spiked into 400 µl of normal human serum after the addition of RLT buffer. RNA was then extracted from the serum using acid phenol/chloroform extractions and an ethanol precipitation. Samples were then labeled with biotin and hybridized on a microarray as previously described.
The approximate size of the small RNAs recovered from serum was determined by isolating large RNA fragments (low ethanol concentration) and small RNA fragments (high ethanol concentration) using the Invitrogen PureLink miRNA isolation kit, after phenol/chloroform extraction and precipitation. The two RNA size fractionations were labeled with biotin (Mirus) and hybridized to a microarray as described above.
Purified nucleic acid from a serum sample was split into two aliquots (DNase I-treated and untreated). One aliquot was digested with DNase I (New England Biolabs, Ipswich, MA) for 30 min at 37°C, following manufacturer's protocol, and then heated to 85°C for 15 min. This sample was then precipitated with NaCl and ethanol and both DNase I-treated and untreated samples were labeled with a Mirus biotin-labeling kit as described above.
Array data accession numbers: GPL8686; GSE16512; GSM414832 - GSM414867.
We would like to thank Luisa Dugan and Keri McLerran for synthesis and QC of oligonucleotide microarrays.