Methylation of the cytosine is the most frequent epigenetic modification of DNA in mammalian cells. In humans, most of the methylated cytosines are found in CpG-rich sequences within tandem and interspersed repeats that make up to 45% of the human genome, being Alu repeats the most common family. Demethylation of Alu elements occurs in aging and cancer processes and has been associated with gene reactivation and genomic instability. By targeting the unmethylated SmaI site within the Alu sequence as a surrogate marker, we have quantified and identified unmethylated Alu elements on the genomic scale. Normal colon epithelial cells contain in average 25 486
Progress in large-scale sequencing projects is critical to identify and decipher gene organization and regulation in many species including human. Nevertheless, cumulated evidences indicate that the complexity of living organisms is not just a direct outcome of the number of coding sequences and that the presence of multiple regulatory mechanisms accounts for a significant part of biological complexity (
Silenced regions in mammals and other vertebrates are differentiated, although not exclusively, by the presence of DNA methylation (
Noteworthy, even though a vast number of CpG dinucleotides are provided by the collection of repetitive sequences in the human genome, this dinucleotide is greatly under-represented throughout the genome, but it can be found at close to its expected frequency in small genomic regions (200 bp to a few kb), known as CpG islands (
Cancer cells are characterized by the accumulation of both genetic and epigenetic changes. Widespread genomic hypomethylation is an early alteration in carcinogenesis and has been associated with genomic disruption and genetic instability (
Here we report two variants of a novel methodology to quantify and identify unmethylated Alu sequences. The CpG site within the consensus Alu sequence AACCCGGG is used as a surrogate reporter of methylation. Unmethylated sites are cut with the methylation-sensitive restriction endonuclease SmaI (CCCGGG) and an adaptor is ligated to the DNA ends. Quantification of UnMethylated Alus (QUMA) is performed by real-time amplification of the digested and adaptor-ligated DNA using an Alu consensus primer that anneals upstream of the SmaI site and an adaptor primer extended with the TT dinucleotide in its 3′ end ( Schematic diagram of the QUMA and AUMA methods. DNA is depicted by a solid line, Alu elements are represented by dashed boxes. The QUMA and AUMA recognition sites (AACCCGGG) are represented by dashed/gray boxes. CpGs at SmaI sites are shown as full circles when methylated and as open circles when unmethylated. The methylation-sensitive restriction endonuclease SmaI can only digest unmethylated targets, leaving blunt ends to which adaptors can be ligated. ( AUMA of normal (N)–tumor (T) pairs of two different patients performed using primer BAu-TT. A highly reproducible band patterning is observed among the four replicates. Representative bands showing gains (hypomethylations) and losses (hypermethylations) are marked with up and down arrowheads, respectively.
Application of QUMA and AUMA to a series of colorectal carcinomas and their paired normal mucosa has offered global estimates of unmethylation of Alu elements in normal and cancer cells and has revealed a large collection of unique sequences that undergo highly recurrent hypomethylation and hypermethylation in colorectal tumors.
Fifty colorectal carcinomas and their paired non-adjacent areas of normal colonic mucosa were included in this analysis. Samples were collected simultaneously as fresh specimens and snap-frozen within 2 h of removal and then stored at −80°C. All samples were obtained from the Ciutat Sanitària i Universitària de Bellvitge (Barcelona, Spain). The study protocol was approved by the Ethics Committee. Human colon cancer cell lines (HT29, SW480, HCT116, LoVo, DLD-1, CaCo-2 and LS174T) were obtained from the American Type Culture Collection (ATCC; Manassas, VA). KM12C and KM12SM cells were generously provided by A. Fabra. DNA from tumor–normal pairs was obtained by conventional organic extraction and ethanol precipitation. DNA purity and quality was checked in a 0.8% agarose gel electrophoresis. RNA from cell lines was obtained by phenol–chloroform extraction and ethanol precipitation, following standard procedures.
The distribution of SmaI sites, putative amplification hits, PCR homologies, CpG islands and repetitive elements was assessed using the human genome assembly 36.1 from NCBI. Data were obtained from the Repbase (
To calculate the proportion of unmethylated Alu elements at the genomic level, the number of AUMA hits identified in bioinformatic analysis were corrected according to the distribution of experimentally generated AUMA products performing Monte Carlo simulations. One thousand Monte Carlo simulations were performed using an Excel Add-in (available at
One microgram of DNA was digested with 20 U of the methylation sensitive restriction endonuclease SmaI (Roche Diagnostics GmbH, Mannheim, Germany) for 16 h at 30°C, leaving cleaved fragments with blunt ends (CCC/GGG). Adaptors were prepared incubating the oligonucleotides Blue (CCGAATTCGCAAAGCTCTGA) and the 5′ phosphorylated MCF oligonucleotide (TCAGAGCTTTGCGAAT) at 65°C for 2 min, and then cooling to room temperature for 30–60 min. One microgram of the digested DNA was ligated to 2 nmol of adaptor using T4 DNA ligase (New England Biolabs, Beverly, MA, USA). Subsequent digestion of the ligated products with the methylation insensitive restriction endonuclease XmaI (New England Biolabs) was performed to avoid amplifications from non-digested methylated Alu's. The products were purified using the GFX Kit (Amersham Biosciences, Buckinghamshire, UK) and eluted in 250 μl of sterile water.
Quantitative real-time PCR was performed using 1 ng (the equivalent of 333 genomes) of DNA in a LightCycler 480 real-time PCR system with Fast Start Master SYBR Green I kit (Roche). Mastermix was prepared to a final concentration of 3.5 mM MgCl2 and 1 μM of each primer. The downstream BAu-TT primer (constituted by the 3′ end of Blue primer, and the GGGTT sequence including the GGG 3′ side of the cut SmaI site and the Alu homologous TT dinucleotide, ATTCGCAAAGCTCTGAGGGTT) and the upstream primer was an Alu consensus sequence (CCGTCTCTACTAAAAATACA) (see Supplementary Data). Magnitudes were expressed as number of unmethylated Alus per haploid genome after DNA input normalization. The number of haploid genomes present in the test tube was determined in the same multiwell plate by quantification of Alu sequences irrespectively of the methylation state. A real-time PCR using Alu consensus primers upstream of the CCCGGG site was performed (see Supplementary Data) and the number of genomes was calculated against a standard curve constructed with a reference genomic DNA measured by UV spectrophotometry.
To determine the efficiency of the assay and to perform absolute quantification, an external Alu product generated by PCR from a DNA fragment containing an AluSx element was used as standard (Supplementary Methods). The number of copies of the external control were spectrophotometrically quantified and dilution curves were generated and treated as samples. Comparison of dilution curves before and after sample processing indicated that the mean recovery was 73%. DNA samples overdigested with the methylation insensitive XmaI endonuclease were spiked with different amounts of the external standard and processed. The sensitivity of the QUMA detection was 100 unmethylated Alu's per haploid genome (Supplementary Figure 1) using 1 ng of genomic DNA per PCR. A linear response was observed between 1000 and 100 000 unmethylated Alu's per haploid genome (Supplementary Data).
DNA digestion with SmaI enzyme and ligation to the linker was performed as described above for QUMA, except for the XmaI digestion that was skipped. The product was purified using the GFX Kit (Amersham Biosciences) and eluted in 250 μl of sterile water. Six different chimeric primers constituted by the 3′ end of the Blue primer sequence (ATTCGCAAAGCTCTGA), the cut SmaI site (GGG) and two, four or seven additional nucleotides homologous to the Alu consensus sequence were used to enrich for Alu sequences (see Supplementary Methods). Three primers were designed to amplify ‘upstream’ of the SmaI site (towards ALU promoter): BAu-TT, BAu-TTCA, BAu-TTCAAGC. Three other primers were designed to amplify ‘downstream’ (towards ALU poly-A): BAd-AG, BAd-AGGC, BAd-AGGCGGA. Letters after the dash correspond to the 3′ sequence of the primer (see Supplementary Data). Data reported here were obtained by using the BAu-TT primer.
In each PCR reaction only one primer was used at a time. Products were resolved on denaturing sequencing gels. Although bands can be visualized by silver staining of the gels, radioactive AUMA's were performed for normal–tumor comparisons. A more detailed description of the PCR and the visualization of the bands are given as Supplementary Data.
Only sharp bands that were reproducible and clearly distinguishable from the background were tagged and included in the analysis. Faint bands with inconsistent display due to small variations in gel electrophoresis resolution were not considered. Band reproducibility was assessed with the analysis of PCR duplicates of three independent sample digests from two different samples and PCR replicates from the same digest from four paired tumor–normal samples. AUMA fingerprints were visually checked for methylation differences between bands in the tumor with regard to its paired normal mucosa. Under these premises, a given band was scored according to three possible behaviors: hypomethylation (increased intensity in the tumor), hypermethylation (decreased intensity in the tumor) and no change (no substantial difference in intensity between normal and tumor samples) (
The origin and chromosomal distribution of sequences generated by AUMA was analyzed using procedures analogous to CGH. Briefly, an AUMA product obtained from a normal tissue DNA was purified using Jet quick PCR product purification kit (Genomed, Löhne, Germany) and labeled with SpectrumRed dUTP (Vysis, Downers Grove, IL, USA) using a Nick Translation kit (Vysis). Similarly, genomic DNA of the same normal sample was labeled with SpectrumGreen dUTP (Vysis) and both probes were cohybridized to metaphase chromosomes. Procedures and image analysis were performed as described (
Differential normal–tumor representation of AUMA at the genomic scale was performed by competitive hybridization of AUMA products to BAC arrays. AUMA products from two normal–tumor pairs were purified using Jet quick PCR product purification kit (Genomed, Löhne, Germany) and 1 μg was labeled with dCTP-Cy3 or dCTP-Cy5 (Amersham Biosciences, UK) by use of the Bioprime DNA Labeling System (Invitrogen, Carlsbad, CA, USA). Probes were hybridized to SpectralChip 2600 BAC arrays (Spectral Genomics, Houston, TX, USA) following the manufacturer's instructions. Arrays were scanned with a ScanArray 4000 (GSI Lumonics, Watertown, MA, USA) and processed with GenePix software (Axon Instruments, Union City, CA, USA). The resulting data were processed to filter out low-quality spots based on spot area and similarity of readings between the two replicates of each BAC. Data manipulation was performed using Excel spreadsheets. Because AUMA products are not evenly distributed along chromosomes, only BACs with intensities above the 10% of maximum intensity in at least one of the two channels were considered for ratio calculations. The pattern of chromosomal alterations in these two tumors was determined by conventional CGH as described (
DNA excised from gels was directly amplified with the same primer used in AUMA (BAu-TT) (Supplementary Figure 2). The amplified product was cloned into plasmid vectors using the pGEM-T easy vector System I cloning kit (Promega, Madison, WI, USA). Automated sequencing of multiple colonies was performed using the Big Dye Terminator v3.1 Cycle Sequencing kit (Applied Biosystems, Foster City, CA, USA) to ascertain the unique identity of the isolated band. Sequence homologies were searched for using the Blat engine (
Differential methylation observed in some AUMA tagged bands was confirmed by direct sequencing of bisulfite treated normal and tumor DNA as previously described (
Briefly, 6 × 106 cells were washed twice with PBS and cross-linked on the culture plate for 15 min at room temperature in the presence of 0.5% formaldehyde. Cross-linking reaction was stopped by adding 0.125 M glycine. All subsequent steps were carried out at 4°C. All buffers were pre-chilled and contained protease inhibitors (Complete Mini, Roche). Cells were washed twice with PBS and then scraped. Collected pellets were dissolved in 1 ml lysis buffer (1% SDS, 5 mM EDTA, 50 mM Tris pH 8) and were sonicated in a cold ethanol bath for 10 cycles at 100% amplitude using a UP50H sonicator (Hielscher, Teltow, Germany). Chromatin fragmentation was visualized in 1% agarose gel. Obtained fragments were in the 200–500 pb range. Soluble chromatin was obtained by centrifuging the sonicated samples at 14 000
Immunoprecipitation was carried out at 4°C by adding 5–10 μg of the desired antibody to 1 ml of chromatin. Chromatin–antibody complexes were immunoprecipitated with specific antibodies using a protein A/G 50% slurry (Upstate, Millipore, Billerica, MA, USA) and subsequently washed and eluted according to the manufacturer's instructions. Antibodies against acetylated H3 K9/K14 (Upstate), dimethylated H3 K79 and trimethylated H3 K9 (Abcam, Cambridge, UK) were used. Enrichment for a given chromatin modification was quantified as a fold enrichment over the input using quantitative real-time PCR (Roche). For every PCR, a standard curve was obtained to assess amplification efficiency. All quantifications were performed in duplicate.
The availability of the human genome map has allowed us to make a detailed estimation of the frequency and distribution of the sites targeted by our approaches on the genomic scale. A Perl routine was used to score all positions containing the target sequences in all chromosomes and was also applied to perform a virtual AUMA (see Material and Methods section). Some of the most important data derived from the bioinformatic analysis are shown in Relative distribution the Alu elements and sequence targets considered in bioinformatic and experimental QUMA and AUMA. Mb: number of megabases occupied by each type of element; elements: number of elements considered (‘Rest’ has been set arbitrarily to 50%); SmaI site: CCCGGG sequence; vQUMA hits: AACCCGGG (or GGGCCCTT) sites in Alu elements; vAUMA hits: AACCCGGG (or GGGCCCTT) sites; vAUMA ends: vAUMA hits considering only putative AUMA products of <1 kb (see Material and Methods section); AUMA: elements at each one of the two ends of actual AUMA products. Content and distribution of QUMA and AUMA hits in the human genome aGenome Mb represented by each type of element. Total number corresponds to the number of megabases analyzed for the presence of hits. Only assembled chromosome fragments were considered. bElements considered in the analysis as obtained from the Repbase and the Genome Browser Databases (see Material and Methods section). cNumber of occurrences of the sequence AACCCGGG (or CCCGGGTT) within each type of element. dNumber of AUMA hits present in virtual PCR products of up to 1000 bp. eHits of actual AUMA products. Only bands appearing in normal tissue were considered. Eighty-seven bands contributed two hits each (174 hits) and 27 bands contributed only one due to poor sequence or incomplete homology with the NCBI Build 36.1 of the human genome (hg18 assembly, March 2006). Twenty-three additional bands were detected mainly in tumor tissue and were not considered to perform calculations. fEstimated number of unmethylated sites using Monte Carlo simulations (Material and Methods section). gIn respect to the total number of AACCCGGG (or CCCGGGTT) hits.Sequence Mb Number of elements SmaI sites (CCCGGG) AACCCGGG hits Virtual AUMA hits AUMA hits Unmethylated hits Unmethylated hits (%) Total 3080.4 1 118 195 486 835 168 309 5498 201 14332 ± 2418 8.52 ± 1.4% Alu (S+J+Y) 227.3 1 091 110 198 201 155 226 5109 59 (29.3%) 4104 ± 688 2.64 ± 0.44% AluS 141.2 660 415 122 459 97 951 3382 45 (22.4%) 3028 ± 510 3.09 ± 0.51% AluJ 54.0 283 104 14 017 1235 38 2 (1.0%) 151 ± 25 12.25 ± 1.97% AluY 32.1 147 591 61 725 56 040 1689 12 (6.0%) 925 ± 156 1.65 ± 0.27% CpG islands 16.2 27 085 49 430 1673 63 55 (27.4%) 1501 ± 97 90.5 ± 5.79% Rest 2836.9 – 239 204 11 410 326 87 (43.3%) 8530 ± 1650 75.9 ± 14.63%
Because of the C to T mutational bias at CpG sites (
The representativity of AUMA was analyzed by a virtual bioinformatic assay of the human genome sequence. A total of 168 309 AACCCGGG (or CCCGGGTT) hits were identified throughout the genome, with 92.9% of all hits within Alu elements (
The QUMA approach was applied to quantify unmethylated Alu's in a series of 18 colorectal carcinomas and their paired normal colonic mucosa. An external DNA fragment containing an AluSx element was used as a standard (see Materials and Methods section and Supplementary Data) in order to make an absolute quantification of the number of unmethylated Alu's. Replicates and dilution curves of the samples and standard were performed to assess reproducibility, sensitivity and accuracy (Supplementary Figure 1). Results were normalized by assessment of the number of haploid genomes per test tube (see Material and Methods section). The average number of unmethylated Alu's per haploid genome was 25 486 ± 10 157 in normal mucosa, and 41 995 ± 17 187 in tumor samples ( Quantitation of unmethylated Alu's in 17 paired normal mucosa and colorectal carcinoma by QUMA. The values represent the estimated number of unmethylated Alu's per haploid genome. Most tumors exhibited a higher level of hypomethylation when compared with the respective normal.
Because QUMA products are fully contained within the Alu sequence, it is not possible to identify and position in the genome the unmethylated Alu elements. To achieve this it is necessary to amplify the targeted unmethylated Alu element together with an adjacent unique sequence. This was attained through the use of the second method, the Amplification of UnMethylated Alus (AUMA). AUMA also targets the unmethylated AACCCGGG sequence, as in QUMA, but in this case a single primer is used in the PCR (BAu-TT) (see
AUMA products generated using the BAu-TT primer produced highly reproducible fingerprints consisting of bands ranging from ∼100 to ∼2000 bp when resolved in high-resolution sequencing gels (
It should be noted that different fingerprints containing alternative representations may be obtained by AUMA just by using primers that either amplify from the SmaI site towards the Alu promoter (upstream Alu amplification) or towards the Alu poly-A tail (downstream Alu amplification). Also the stringency of the Alu selection may be increased by using longer primers containing additional nucleotides corresponding to the Alu consensus sequence (see Material and Methods section). An illustrative example of AUMA fingerprints generated with different Alu-upstream and Alu-downstream primers is shown in Supplementary Figure 2. All the data reported in this article regarding AUMA were obtained using the BAu-TT primer.
Competitive hybridization between AUMA products and genomic DNA on metaphase chromosomes yielded a characteristic hybridization pattern demonstrating the unequal distribution of AUMA products along the human genome ( (
To determine the identity of bands displayed by AUMA, 38 tagged bands were isolated and cloned. Multiple clones from each band were sequenced, resulting in a total of 49 different sequences due to the coincidence of more than one sequence in some bands. Characterized bands included bands displaying no changes in the normal–tumor comparisons and bands recurrently altered in the tumor. A selection of characterized AUMA bands aNucleotide position within the contig (strand +). NCBI Build 36.1 of the human genome. bThe whole sequence or a fragment of the sequence lays not further than 200 bp of a predicted CpG island. cAs compared to the paired normal tissue.Band ID Size (bp) % GC Chromosome map (Location Gene CpG island Repetitive elements in band ends (5′/3′) Methylation status in tumor Ai1 c3 509 55 17p11.2 (18206453–18206961) SHMT1 Yes Alu Sx/MIR Hypermethylated Aj2 c1 458 49 1q32.2 (206389082–206389539) MGC29875 Yes Alu Sq/None Hypomethylated Ao1 c4 365 50 19q13.32 (53550179–53550543) AK001784 No Alu Sx/MIRb Hypomethylated Ap1 c6 358 56 5q35.2 (175157321–175157674) CPLX2 Yes None/MIR Hypermethylated Aq3 c6 339 58 8p23.3 (2007343–2007682) MYOM2 No LTR/Alu Y Hypomethylated Ar3 c3 329 57 2q14.3 (127875178–127875506) AF370412 Yes None/MIRb Hypomethylated As3 c6 306 63 16p13.3 (3160477–3160782) None Yes None/None Hypermethylated Au4 c1 268 57 16p13.3 (3162099–3162366) None No tRNA/None Hypermethylated
To obtain a more representative collection of AUMA bands, 200 clones obtained from normal tissue AUMA products were sequenced. The analysis revealed 88 additional sequences. This resulted in a total of 137 different loci represented in AUMA (Supplementary Table 1). Most sequences obtained by random cloning were also flanked by two AACCCGGG sequences in opposite DNA strands. Nevertheless, in 27 sequences the AACCCGGG site was only present at one of the ends, with the other end showing high homology with the primer although it was not a perfect match. The presence of these sequences suggests that, in some instances, a single cut in the sequence may be enough to produce an amplifiable fragment. This is not considered an artifact since these bands still represent an unmethylated AACCCGGG site.
Of the 137 identified loci represented in AUMA, 114 were isolated from normal tissue DNA and 23 from tumor DNA. Half of the sequences contained an Alu sequence at one of the ends and two were flanked by two inverted Alu's. AUMA sequences isolated from tumor tissue and not present in normal tissue (this corresponds to a tumor-specific hypomethylation) showed a higher proportion of Alu elements (16 out of 23, 70%), and included one sequence flanked by two inverted Alu's. Globally, 78 unmethylated Alu elements were identified and positioned in the human genome map.
To study the genomic distribution of unmethylated sequences in normal colon mucosa, we only considered the 114 sequences obtained from normal tissue. This resulted in a total of 201 unmethylated hits characterized throughout the genome. The nature of the sequences represented in actual AUMA showed striking differences with the distribution expected from the virtual AUMA analysis. The methylation status of the sequence is likely to be the main (if not the only) source of these differences because the virtual AUMA did not consider this state. Therefore we can use these differences to estimate the degree of unmethylation of the Alu repeats. Only 29.4% of the AUMA ends consisted of Alu's, as compared with the expected 92.9% resulting from the bioinformatic analysis. The highest downrepresentation corresponded to the youngest AluY family, which was present in 6.0% of the AUMA ends, while it was expected to add up to 33.1% in virtual AUMA. AluS representation in actual and virtual AUMA was 22.3% and 60%, respectively. Interestingly, AluJ representations, in both actual and virtual AUMA, were closer (1.0% and 0.7%, respectively) (
To calculate the proportion and distribution of unmethylated Alu elements on a genomic scale, we performed Monte Carlo simulations taking into account the observed and expected distribution of hits in each Alu family and CpG islands and the rest of sequences (see Material and Methods section). We estimate that at least 4104 Alu elements are unmethylated or partially unmethylated in normal colonic mucosa. This corresponds to 2.64% of all Alu elements containing the target sequence AACCCGGG (
In order to test the usefulness of the method for the detection of new altered methylation targets, we applied AUMA to a series of 50 colorectal carcinomas and their paired normal mucosa. Two cases were excluded from the analysis due to recurring experimental failure of the normal or tumor tissue DNA. For the rest of 48 normal–tumor pairs, consistent and fully readable fingerprints were generated and evaluated for normal–tumor differential representation. A given case presented, on average, 107 ± 2.9 informative bands (range 98–110). The variation was due to polymorphic display or variable resolution power of gel electrophoresis.
In this study, only those bands showing clear intensity differences between normal and tumor tissue fingerprints (
Virtually, all tagged bands (109 out of 110) were found to be altered in at least one tumor when compared to its normal paired mucosa. AUMA tagged bands presented a wide distribution in the hypomethylation/hypermethylation rates (proportion of tumors showing differential display compared to the paired normal tissue) ( Distribution of hypermethylation and hypomethylation rates in the 110 AUMA tagged bands. Rates were obtained by comparison of the AUMA fingerprints obtained in 50 colorectal tumors as compared to their respective matched normal tissue.
In order to determine whether normal–tumor differences were limited to isolated independent loci or changes that might affect larger chromosomal regions, we compared the distribution of AUMA products generated from two paired normal and tumor tissues and hybridized to BAC arrays. Differential hybridization was observed in many BACs, suggesting that relatively large regions encompassing from several hundred Kbs to a few Mbs may undergo concurrent hypomethylation or hypermethylation. Telomeric regions of many chromosomes contained most of the differential display (
To confirm that the changes observed in AUMA fingerprints corresponded to actual changes in the methylation status of the sequence, eight different sequences obtained from AUMA fingerprints were analyzed in normal and tumor tissues by direct sequencing of sodium bisulfite-treated DNAs (
Next, we wondered if DNA methylation changes detected by AUMA may have any functional consequences. We chose one of the most recurrent hypomethylated AUMA sequences (Aq3) and performed an insightful epigenetic characterization of the region in a series of normal–tumor pairs and in colon cancer cell lines.
Aq3 band is recurrently hypomethylated in tumors according to AUMA fingerprints ( (
Next, we wondered whether the DNA methylation status of the AluYd3 element was associated with alternative chromatin states. We performed Chromatin ImmunoPrecipitation (ChIP) analysis of histone 3 (H3) modifications indicative of active chromatin: acetylation of lysines 9 and 14 (AcH3K9/K14), and dimethylation of lysine 79 (2mH3K79); and silent chromatin: trimethylation of lysine 9 (3mH3K9). These histone marks were compared between cell lines HCT116 and LoVo (with 100% and 30% methylation of the AluY element, respectively). The silencing mark 3mH3K9 was 3.5-fold higher in HCT116 cells compared to the LoVo cell line (
Full genome sequencing has provided precise maps of repetitive elements, and several studies have investigated their distribution and relationship with genome structure (
Beyond these few studies, the extension and nature of the epigenetic state of interspersed elements is largely unknown. Global estimates of DNA methylation in repetitive elements have been obtained by Southern blot analyses (
Here we report a systematic screening of unmethylated Alus as a tool to determine the extent of DNA hypomethylation, to identify specifically unmethylated elements and to detect epigenetic alterations in cancer cells. QUMA is a very simple and specific method and provides accurate relative estimates of the number of unmethylated elements. QUMA is specially appropriate for comparative studies, but also provides a raw quantitation of the number of unmethylated elements per haploid genome, outlining the extent of hypomethylated Alu's in normal and pathologic cells. QUMA analysis indicates that about 1 out of 6 Alu elements containing the AACCCGGG site are unmethylated, while in tumors, this figure nearly doubles in agreement with previous studies (
To date there is still a lack of proper methodologies allowing genome-wide screenings for recurrent hypomethylated regions that may have some impact on tumor biology. Even though QUMA and other methodologies (
Due to sequence degeneration, both QUMA and AUMA are more effective in screening for unmethylation in younger elements. This trend is more clearly seen in AUMA, with only 9% of the Alu elements of the old J subfamily containing the SmaI site retain the AA dinucleotide needed for their amplification, while this figure is 91 and 80% in the younger AluY and AluS subfamilies, respectively (
AUMA was designed to amplify DNA fragments containing the target sequence (AACCCGGG), which is present in Alu and other repetitive elements. Because a single primer was used for PCR amplification, the target sequence must appear in both strands of the DNA at relatively nearby positions. As expected, Alu elements, with more than one million copies per human genome (
It is worth noting that AUMA patterns are highly reproducible not only in replicates but also among different samples, which indicates that the unmethylated status of these repeats is tightly controlled, probably by the epigenetic status of nearby regions. This is strengthened by the confirmation that unmethylation extends many CpG sites beyond the SmaI cut site. Moreover, about 50% of the bands tagged in AUMA fingerprints exhibited variable display among normal tissues (data not shown), suggesting the usefulness of this technique to investigate epigenetic polymorphisms.
Alu's and other repetitive elements tend to be highly methylated in most somatic tissues (
Alu families showed striking differences in their methylation level. Most of the Alu elements characterized here are from the younger families AluS and AluY (74% and 22%, respectively). Nevertheless, this observation is mainly due to the depletion of CpG sites in older Alu elements. Hence, only 1 out 230 AluJ elements maintains the AUMA target site (AACCCGGG), while younger elements show higher rates of maintenance in accordance with their age (AluS: 1 out of 7; AluY: 2 out of 5) (
Cancer-related hypomethylation is well documented (
The application of AUMA to a series of colorectal carcinomas and their paired matched normal tissue has revealed a high rate of alterations. This indicates the plasticity of epigenetic control of the elements screened by AUMA in colorectal carcinogenesis. Although some bands show bidirectional changes (hypomethylations and hypermethylations), which have been also reported in other sequences (
As an example, we have investigated the Aq3 sequence, one of the most recurrent hypomethylations in this study. Aq3 band is flanked by two repeats, a LTR and an AluY, which map within an intron between exons 8 and 9 of the
Another application of AUMA is the detection of genomic regions that have been silenced in cancer. Interspersed elements are concentrated in gene-rich regions and due to the intended selection of unmethylated repetitive elements in AUMA, it appears reasonable to postulate that normally unmethylated sequences are likely to pinpoint active genomic regions. In this context, AUMA provides a large collection of genomic regions undergoing hypermethylation, which are readily seen as bands recurrently loss in the fingerprints. DNA methylation associated epigenetic silencing is probably one of the most prevalent mechanisms of tumor suppression inactivation in cancer (
In summary, QUMA and AUMA methodologies are a simple and novel approach to explore and gain insights into the functional significance of interspersed genomic elements and neighboring sequences. Due to its distinctive features (bias for unmethylated elements in gene-rich regions and detection of both hypomethylation and hypermethylation) we think that these techniques constitute a new and unique tool that should complement global determinations and high-resolution genome-wide scanning strategies. Beyond unmethylated repetitive elements, AUMA can be also used to detect recurrent epigenetic changes associated with tumorigenesis including gene epigenetic inactivation.
Supplementary data are available at NAR Online.
We thank Gemma Aiza for technical support and Jessica Halow for critical review of the manuscript. J.R. was a fellow of the Generalitat de Catalunya; E.V. was a fellow of the Fondo de Investigación Sanitaria (FIS). This work was supported by a grant from the Ministry of Education and Science (SAF2006/351) and the Consolider-Ingenio 2010 Program (CSD2006-49). Funding to pay the Open Access publication charges for this article was provided by the Institute of Predictive and Personalized Medicine of Cancer (IMPPC).