Anchored ChromPET, a technique to capture and interrogate targeted sequences in the genome, has been developed to identify chromosomal aberrations and define breakpoints. Using this method, we could define the
Chromosomal translocations play a major role in several genetic diseases. Translocations between genes have the potential to constitutively express or repress genes and hence lead to different diseases. The Philadelphia chromosome (Ph) is a prime example of such a translocation, where a fusion gene is constitutively expressed and leads to a particular class of leukemia. There are other translocations that have been implicated in cancers and other genetic diseases, and more are being discovered every day. A method that can quickly and robustly characterize specific translocations and produce DNA-based disease-specific biomarkers will have both diagnostic and prognostic applications. A method that is not dependent on the growth of cells in culture will bring the power of cytogenetics to many more cancers.
The incidence of chronic myeloid leukemia (CML) is 1 to 2 per 100,000 and the disease constitutes 15 to 20% of adult leukemias. CML is characterized by the Ph, resulting from the t(9;22)(q34;q11) balanced reciprocal translocation. The translocation generates the BCR-ABL1 fusion protein with constitutive kinase activity and oncogenic activity. The breakpoints in the
Detection of Ph or
Real-time reverse transcription PCR (RT-PCR) is the most sensitive technique available for the detection of
Here we introduce a method for detecting and monitoring the
Reagents used were APex Heat-Labile Alkaline Phosphatase (Epicentre, Madison, WI, USA; AP49010), Biotin-16-UTP (Roche, Indianapolis, IN, USA; 11388908910), DNAZol reagent (Invitrogen, Carlsbad, CA, USA; 10503-027), Dynabeads M-280 streptavidin (Invitorgen; 112-05D), End-It DNA End Repair Kit (Epicentre; ER0720), human Cot-1 DNA (Invitrogen; 15279-011), MAXIscript Kit (Ambion, Austin, TX, USA; AM1312), MinElute Reaction Cleanup Kit (Qiagen, Valencia, CA, USA; 28204), pCR4-TOPO-TA vector (Invitrogen; K4575-01), QIAquick Gel Extraction Kit (Qiagen; 28704), QIAquick PCR Purification Kit (Qiagen; 28104), QuickExtract FFPE DNA Extraction Kit (Epicentre; QEF81805), QuickExtract FFPE RNA Extraction Kit (Epicentre; QFR82805), Quick Ligation Kit (NEB, Ipswich, MA, USA; M2200S), SuperScript III Reverse Transcriptase (Invitrogen; 18080-093), TaKaRa Ex Taq DNA Polymerase (Takara, Otsu, Shiga, Japan; TAK RR001A), Taq DNA Polymerase (Roche; 11146165001), TRIzol (Invitrogen; 15596-026), and TURBO DNase (Ambion; AM2238).
K562 cells (CCL-243) and KU812 cells (CRL-2099) were purchased from ATCC and cultured according to ATCC instructions.
Genomic DNA from peripheral blood mononuclear cells were kindly provided by Dr Brian Druker (Oregon Health and Science University). Ph+ or Ph- patient samples were obtained with informed consent and under the approval of the Oregon Health and Science University Institutional Review Board. Mononuclear cells were isolated by separation on a Ficoll gradient (GE Healthcare, Piscataway, NJ, USA), followed by purification of genomic DNA using the Dneasy Blood and Tissue kit (Qiagen).
PCR primers used for this study are in listed in Table S1 in Additional file
All chromPET libraries were constructed according to the protocol supplied by Illumina with minor modifications. Genomic DNA was extracted with DNAZol reagent and 2 μg of DNA was sheared by a Nebulizer for 5 minutes by compressed air at 32 to 35 psi. After purifying the sample with a QIAquick PCR purification kit, fragmented DNA was run in 2.0% agarose gel, and 0.5-kb fragments were excised from the gel and extracted with a QIAquick Gel Extraction Kit. The ends of DNA fragments were polished by an End-It DNA End Repair Kit and A-tail added to the 3' end by 0.25 units of Taq DNA polymerase. The Y-shaped adapter containing the bar-code was ligated to both ends of DNA fragments by a Quick Ligation Kit and purified again by 2.0% agarose gel electrophoresis and a QIAquick Gel Extraction Kit. Y-shaped adapter ligated DNA was amplified by PCR primer PE1.0 and 2.0 for 15 cycles and the amplified fragment was again purified by 2.0% agarose gel electrophoresis and a QIAquick Gel Extraction Kit. The sequences of adapters and primers are given in Table S1 in Additional file
We amplified 6.6 kb DNA containing the M-Bcr region from normal lung genomic DNA using PCR primer pair M-BCR-F1 and R1. Amplified DNA (2 μg) was sheared in a Nebulizer for 8 minutes by compressed air at 32 to 35 psi to obtain 0.3-kb fragments, overhanging ends blunted by 2 units of T4 DNA polymerase, the 5' end dephosphorylated by 1 μl of APex Heat-Labile Alkaline Phosphatase, and an A base overhang added to the 3' end by 0.25 units of Taq DNA polymerase. Following each step, the sample was cleaned up by a MinElute Reaction Cleanup Kit. The DNA was cloned into the pCR4-TOPO-TA vector and the resulting construct used to transform
We hybridized 500 ng of biotin-labeled unique single-stranded RNA from the bait to 500 ng of heat-denatured chromPET library in 26 μl of hybridization mixture (5× SSPE, 5× Denhardts', 5 mM EDTA, 0.1% SDS, 20 U SUPERase-In), including 2.5 μg of heat-denatured human Cot-1 DNA and salmon sperm DNA at 65°C for 3 days. RNA-DNA hybrid was captured on Dynabeads M-280 streptavidin that had been washed three times and resuspended in 200 μl of 1 M NaCl, 10 mM Tris-HCl (pH 7.5), 1 mM EDTA and 100 μg/ml salmon sperm DNA. RNA-DNA hybrid capture beads were washed with 0.5 ml of 1× SSC/0.1% SDS once for 15 minutes at 20°C and then with 0.5 ml of 0.1× SSC/0.1% SDS for 15 minutes at 65°C three times. The annealed DNA was eluted by 50 μl of 0.1 M NaOH, neutralized by 70 μl of 1 M tris-HCl (pH 7.5) and converted to double-stranded DNA by paired-end PCR primer PE1.0 and 2.0. DNA fragments were purified by 2.0% agarose gel electrophoresis and high-throughput sequencing was performed according to the manufacturer's protocol (Illumina).
To identify the sample for each individual chromPET in the multiplexed sequencing runs, we used a 4-bp barcode that was included in the sample-specific Y-primers and was appended to the 5' end of each sequence. Allowing a 1-bp mismatch (only in degenerate positions) the chromPET was assigned to one of the samples or left unassigned. The 38-bp PET reads obtained from the sequencer were mapped to the targeted regions using Novocraft Novoalign program (version 2.05) [
The algorithm for breakpoint detection is based on a voting procedure. We allow each junctional chromPET to vote on the location of the actual breakpoint (Figure S2 in Additional file
DNA and RNA from freshly prepared cell lines, formalin fixed cells, and culture medium were extracted with DNAzol, Trizol, QuickExtract FFPE DNA Extraction Kit, or QuickExtract FFPE RNA Extraction Kit according to the manufacturer's protocol.
The chromPET library was constructed according to the manufacturer's protocol with a slight modification. We used Y-shaped adapters that encoded the bar-code sequence immediately after the sequencing primer and before the insert to be sequenced (Figure
We multiplexed the bar-coded libraries from two leukemia cell lines, K562 and KU812, into one lane and that from three patient samples, PS1, PS2 and PS3, into another lane of the Illumina Genome Analyzer. We performed 38 cycles of paired end sequencing using the protocols provided by the manufacturer.
As shown in Tables
Sequencing and mapping numbers for cell lines out of 3,249,760 total reads
| Cell line | ||
|---|---|---|
|
|
||
| K562 | KU812 | |
| Barcoded reads | 161,365 | 1,468,876 |
| Mapped | ||
| First tag | 24,385 | 243,684 |
| Second tag | 25,310 | 246,861 |
| Percent mapped | ||
| First tag | 15% | 17% |
| Second tag | 16% | 17% |
| Mapped uniquely | ||
| First tag | 12,800 | 125,795 |
| Second tag | 13,321 | 122,665 |
| Total anchored chromPETs | 2,839 | 21,798 |
| Junctional chromPETs | 131 | 427 |
| Percent breakpoint | 4.6% | 2.0% |
The number of chromPETs sequenced, mapped, anchored to
Sequencing and mapping numbers for patient samples out of 592,785 total reads
| Cell line | |||
|---|---|---|---|
|
|
|||
| Patient sample 1 | Patient sample 2 | Patient sample 3 | |
| Barcoded reads | 89,316 | 258,239 | 37,538 |
| Mapped | |||
| First tag | 8,952 | 30,586 | 3,782 |
| Second tag | 8,861 | 32,275 | 3,966 |
| Percent mapped | |||
| First tag | 10.0% | 11.8% | 10.1% |
| Second tag | 9.9% | 12.5% | 10.6% |
| Mapped uniquely | |||
| First tag | 4,824 | 16,456 | 2,186 |
| Second tag | 4,828 | 17,248 | 2,232 |
| Total anchored chromPETs | 994 | 3,753 | 403 |
| Junctional chromPETs | 23 | 92 | 10 |
| Percent breakpoint | 2.3% | 2.5% | 2.5% |
Number of chromPETs sequenced, mapped, anchored to BCR and junctional for each sample for patient samples.
Using the criteria on identification of bar-codes described in the Materials and methods, the percentage of chromPETs assigned to each sample was approximately 5% for the K562 cell line and approximately 45% for the KU812 cell line. For the patient samples, the percentages were 15%, 45% and 6% for PS1, PS2 and PS3, respectively. The numbers point to a low efficiency of bar-coding for two of the samples (K562 and PS3), and more study is needed on how to choose uniformly efficient barcodes.
Using default mapping parameters (described in the Materials and methods), we obtained a large but variable number of chromPETs (Tables
We next devised an algorithm that utilizes the mapping coordinates of each end of a junctional chromPET together with the distribution of sizes of normal chromPETs to predict the most likely position for the breakpoint between the
Figure S3 in Additional file
Predicted and actual breakpoints from each sample
| Prediction | Break point | Actual | Difference (bp) | ||||
|---|---|---|---|---|---|---|---|
|
|
|
|
|||||
| Sample | M- |
|
M- |
|
M- |
|
|
| K562 | 110,194-110,207 | 27,762-27,909 | 110,191-110,192 | 27,878-27,879 | 3 | 0 | |
| KU812 | 110,241-110,242 | 63,843-63,853 | 110,299-110,300 | 63,929-63,930 | 57 | 76 | |
| 110,096-110,097 | 63,804-63,805 | 144 | 38 | ||||
| Patient 1 | 109,790-109,830 | 125,280-125,623 | 109,781-109,782 | 125,326-125,327 | 8 | 0 | |
| 109,670-109,671 | 149,445-149,446 | 119 | a23,822 | ||||
| Patient 2 | 109,702-109,867 | 102,484-102,653 | 109,834-109,835 | 102,524-102,525 | 0 | 0 | |
| 109,869-109,870 | 102,526-102,527 | 2 | 0 | ||||
Predicted and actual breakpoints for each sample. The absolute difference (in base pairs) between predicted breakpoint site and sequenced breakpoint site is shown in the last two columns. All M-bcr coordinates are relative to chr22:23,522,552 (start position of
The bioinformatics prediction of breakpoints in K562 cells (Table
In a similar fashion we predicted the
We next examined the ability of Anchored ChromPET to identify aberrant translocations in patient samples. To this end, we tested this approach on DNA from blasts in blood samples from Ph+ patients 1 and 2. As a negative control, we also tested this technique in Ph- patient 3. The predicted breakpoints for PS1 and PS2 are reported in Table
Based on these results, we designed primer sets, amplified the junctional fragments and confirmed the
A few M-bcr-anchored chromPETs were also linked to the
Because a clinical sample is not uniformly composed of malignant cells, we next evaluated the sensitivity of detection of the DNA-based biomarkers identified by Anchored ChromPET. A dilution series of K562 cells was created by combining them with HCT116 colon cancer cells without the
The most important benefit of Anchored ChromPET is the precise identification of the breakpoints on DNA, which allows for optimal design of PCR primers for a DNA-based biomarker of the translocation junction. It is well known that RNA is less stable than DNA because the 2'-OH group of a ribonucleotide is more reactive than the 2'-H of a deoxyribonucleotide, causing RNA to break more easily, and because RNAses are present on body surfaces and in body fluids. Formalin-fixed, paraffin-embedded (FFPE) tissue is one of the most commonly archived forms for clinical samples. DNA and RNA from FFPE samples are highly fragmented and, in general, the recovery efficiency of DNA is better than that of RNA. Therefore, we evaluated the sensitivity of detection of DNA- or RNA-based junctional biomarkers in samples extracted from formalin-fixed cells. After extraction of DNA or RNA from 10,000 cells, we measured the yield of DNA or RNA junctions by quantitative real-time PCR and normalized the result to the yield from 1,000 fresh cells. As shown in Figure
Finally, as cells die they release their DNA and RNA into the body fluids and the ideal biomarker will be stable in serum at body temperature. We therefore measured the amount of DNA or RNA biomarkers that survive in serum-containing cell culture medium at 37°C following the growth of K562 cells (Figure
Anchored ChromPET makes it possible to detect gene rearrangements in a targeted region in a short time and provides a personalized DNA-based biomarker for following a patient's disease. This technique has the advantages of both karyotyping and RT-PCR. Twenty-five to 30 metaphase cells are usually examined during karyotyping so that the sensitivity of detecting a Ph-positive cell is 3 to 4%. Interphase FISH can be applied to nondividing cells isolated from peripheral blood to detect the juxtaposition of
We also evaluated the sensitivity of detection of the PCR product spanning the chromosome junction for molecular follow-up of the disease (Figure
With G banding, approximately 400 to 800 bands per haploid set can be detected by a trained cytogeneticist. The haploid human genome occupies about 3 × 109 bp. Thus, the resolution of karyotyping is 5 Mb and the resolution of interphase FISH is 50 to 100 kb. The resolution of RT-PCR for detecting fusion transcripts is not comparable to that obtained here because the chimeric RNA merely indicates the two exons that are fused to each other, with the DNA breakpoints localized anywhere within the adjoining introns. In comparison, we identify the exact DNA junction at the base-pair level by Anchored ChromPET, suggesting that the sequencing-based approach gives the best resolution of the DNA junction.
Anchored ChromPET therefore provides a high-resolution digital karyotype with better sensitivity than comparable methods for detecting the DNA translocation. Note that there is no detectable signal saturation and so the sequencing step can be scaled up by sequencing more DNA to sample even rarer DNA fusion events. About 5 to 10% of CML patients are Ph-negative by karyotyping, but the
Nondividing cells isolated from peripheral blood, which cannot be used for karyotyping, can be used for Anchored ChromPET. There are reports in the literature of successful isolation of 0.5- to 1-kb DNA fragments from blood smears and formalin fixed paraffin embedded tissue. Therefore, Anchored ChromPET and subsequent PCR detection of junctional DNA can be especially useful for retrospective analysis of patient material for both identification of the translocation and detection of minimal residual disease.
How do we expect this technology to be used in the diagnosis and management of new cases of CML? Most patients present in the chronic phase of CML, characterized by leukocytosis with the presence of precursor cells of the myeloid lineage. There are normally between 4 × 109 and 1.1 × 1010 white blood cells in a liter of blood, but this number is significantly increased, with up to 10% blast cells and promyelocytes in the blood in chronic phase CML. In acute phase CML more than 70 to 80% of white blood cells in the peripheral blood can be blasts. RT-PCR seems to be the easiest and most sensitive molecular method for detection of the
A major advantage of Anchored ChromPET is that we do not have to grow the cells in culture and so the method is expected to find wide application in searching for specific translocations for solid cancers where it is difficult to grow all the cancer cells in culture. In addition, since the sensitivity of the method can be increased by sequencing more DNA fragments, we expect it to reliably detect translocations carried by even a small fraction of the cells in a sample. Finally, for translocations (unlike
Only future experiments will define whether the DNA fusion or the RNA fusion will be the better marker for minimal residual diseases or early recurrence. However, since the detection of the DNA fusion does not need reverse transcription and is not as susceptible to the factors that degrade RNA, we anticipate that the DNA fusion fragment may be a more sensitive biomarker than the RNA fusion fragment. We could easily detect the DNA junctional fragment in filtered cell culture medium, suggesting that DNA derived from dead cells survives in serum at 37°C for an extended period of time. In contrast, it is hard to detect the RNA fusion transcript in the same cell culture medium. This observation suggests that another potential advantage of using the DNA junctional fragment as a biomarker is that it may survive as free nucleic acid in body fluids like blood or even urine. This, again, is something that we are interested in testing in the future.
The decrease in sequencing achieved by anchoring, by sampling only the ends of the fragments and by multiplexing multiple samples in the same lane of a sequencer brings the costs of sequencing down considerably. In our estimate, considering the current state of sequencing capabilities and the small number of sequences necessary to identify the breakpoint, we can reliably multiplex up to ten samples in a single lane of the Illumina sequencer, making the sequencing costs much lower than those for whole genome sequencing for identifying cancer-specific recombination biomarkers.
Table
These results demonstrate that the predictions from our algorithm match reasonably well to the breakpoints verified by experimental methods. Our results also suggest that breakpoints could be predicted using even a small number of junctional chromPETs (K562 and PS1). However, we could not predict a consensus breakpoint from PS3 and could not identify a junctional fragment from this DNA using PCR. So even though junctional chromPETs were assigned to patient 3, these are most likely the result of contamination during chromPET library construction. The fact that the contamination did not lead to a false positive call points to the robustness of the approach.
Ligation of a special adapter to the ends of genomic DNA fragments, PCR cycles beginning with an exon of
Well-designed RNA baits useful for the capture of DNA fragments can be commercially synthesized [
Detection of both reciprocal translocations in KU812 and two patient samples allowed us to analyze what happens to the ends of the chromosomes after the break that initiates the translocation. Some DNA sequence is lost at the
In contrast, in KU812 cells and patient 1, some of the DNA at the
The detection of the
B-ALL: B-cell acute lymphoblastic leukemia; BP: base pair; CHROMPET: chromosomal paired end tag; CML: chronic myeloid leukemia; FFPE: formalin-fixed: paraffin-embedded; FISH: fluorescent
AD in partnership with the University of Virginia has founded a company to commercialize this technology.
All authors contributed to the conception of this project. YS developed Anchored ChromPET library preparations and validated predicted regions by PCR. AM designed a strategy of data analysis. AD devised and supervised the project. All authors contributed to the drafting of the manuscript.
Click here for file
We are grateful to Dr Brian Druker at Oregon Health and Science University for providing us with genomic DNA from peripheral blood mononuclear cells from three patients with CML. We thank members of the Dutta Lab and Dr Amir Jazaeri for helpful suggestions and Dr Michael Douvas for reading the manuscript. This work was supported by R01 CA60499 and CA89406.