This is an Open Access article distributed under the terms of the Creative Commons Attribution License (
Mutations in the mismatch repair genes
Associations between genetic variants in
Heterozygous and homozygous changes were detected in 13 of 35 analyzed variants. Two variants showed a borderline association with colorectal cancer, whereas the remaining variants demonstrated no association. Furthermore, the genomic regions covering
None of the variants in
Colorectal cancer (CRC) is a common malignant disease in the western world. The lifetime risk is about 5% and rising [
A twin study has demonstrated that up to 35% of the CRCs can be explained by inherited susceptibility [
A part from the clearly pathogenic mutations genetic screening has also revealed numerous missense, silent and non-coding variants of unknown significance in
The present study describes a population based analysis of the frequency of variants of unknown significance in
The subjects were selected from the Danish Diet, Cancer and Health (DCH) study, which is an ongoing prospective follow-up study [
Characteristics of the cohort
| Cases (n = 380) | Sub-cohort (n = 770) | |
| Sex (Number (%)) | ||
| Men | 213 (56%) | 427 (55%) |
| Women | 167 (44%) | 343 (45%) |
| Age* (Years (sd)) | ||
| Men | 58.4 (4.1) | 56.9 (4.5) |
| Women | 58.5 (4.5) | 56.4 (4.4) |
| Time observed (Years (sd)) | ||
| Men | - | 6.6 (1.2) |
| Women | - | 6.6 (1.0) |
*Age at inclusion into the cohort given as mean (sd) in years
Another cohort named familiar CRC consisting of 285 CRC cases was also included in the study. The majority of the variants were initially identified in this cohort. Some variants were found only in one family whereas others were identified in several families in the cohort. The cohort consists of individuals from HNPCC families (based on the Amsterdam II criteria) or individuals from families not fulfilling the Amsterdam II criteria, but with a clear familial accumulation of CRC. The presence of other clearly pathogenic mutations in the
DNA was extracted from frozen leukocytes as described previously [
Initially, 47 variants in
| Id | Variant | Amino acid change | Identified in the present study | Phylogeny | Pathogenic status | References |
|
|
||||||
| 23 | c.307-29 C>A | Intronic | + | - | - | - |
| 24 | c.350 C>T | p.Thr117Met | - | Conserved | Pathogen | [11,50] |
| 25 | c.453+79 A>G | Intronic | + | - | - | - |
| 26 | c.545+43 C>G | Intronic | + | - | - | - |
| 27 | c.655 A>G | p.Val219Ile | + | Ile other species | Neutral | [11,40,51] |
| 28 | c.790+10 A>G | Intronic | + | - | - | - |
| 29 | c.884+39 G>A | Intronic | - | - | - | - |
| 30 | c.884+83_84 ins T | Intronic | - | - | - | - |
| 85 | c.1217 G>A | p.Ser406Asn | + | Non-conserved | Neutral | [11,41,52] |
| 31 | c.1379 A>C | p.Glu460Ala | - | Non-conserved | - | - |
| 86* | c.1558+11 G>A | Intronic | - | - | Neutral | [53] |
| 32 | c.1558+14 G>A | Intronic | + | - | Neutral | [40] |
| 37 | c.1668-19 A>G | Intronic | + | - | Neutral | [40] |
| 101 | c.1689 A>G | p.Ile563Met | - | Non-conserved | - | - |
| 33 | c.1732-2 A>T | Intronic | - | - | Pathogen | [54] |
| 34 | c.1852_1853 AA>GC | p.Lys618Ala | + | Non-conserved | Neutral/Pathogen | [9,11,45,46,48] |
| 100 | c.1942 C>T | p.Pro648Ser | - | Conserved | Pathogen | [9,45,55] |
| 35 | c.1959 G>T | p.Leu653Leu | + | - | Neutral | [42] |
| 36 | c.2152 C>T | p.His718Tyr | - | Conserved | Neutral | [11,56] |
|
|
||||||
| 89 | c.-118 T>C | promoter | + | - | - | [32,33] |
| 87 | c.131 C>T | p.Thr44Met | - | Conserved | Pathogen | [55] |
| 88 | c.134 C>T | p.Ala45Val | - | Val other species | Neutral | [55] |
| 39 | c.212-23 A>C | Intronic | - | - | - | - |
| 90* | c.287 G>A | p.Arg96His | - | Non-conserved | Neutral | [57,58] |
| 91* | c.329 A>G | p.Lys110Arg | - | Non-conserved | Neutral | [57] |
| 92* | c.380 A>G | p.Asn127Ser | - | Conserved | Neutral/Pathogen | [48,59] |
| 93 | c.560 T>G | p.Leu187Arg | - | Conserved | Neutral/Pathogen | [60] |
| 41 | c.965 G>A | p.Gly322Asp | + | Conserved | Neutral | [51,61] |
| 43 | c.1511-9 A>T | Intronic | + | - | - | - |
| 48 | c.1786_1788 del AAT | p.Asn596del | - | Non-conserved | Pathogen | [62,63] |
| 95* | c.2006-6 T>C/G | Intronic | - | - | Neutral | [64] |
| 96 | c.2062 A>G | p.Met688Val | - | Conserved | - | - |
| 97* | c.2139 G>C | p.Gly713Gly | - | - | Neutral | [57,65] |
| 50 | c.2500 G>A | p.Ala834Thr | - | Non-conserved | Neutral | [10,66] |
| 51 | c.2542 G>T | p.Ala848Ser | - | Conserved | - | - |
* identified in CRC families from other Caucasian populations than the Danish
The multiplex PCR primers were designed using Oligo 6 [
The multiplex PCR amplification was performed with 6–9 primer pairs per reaction. Twenty ng of genomic DNA with 0.5 μl Accuprime™ DNA polymerase (Invitrogen, Carlsbad, CA), 1× Accuprime™ Buffer I, and 0.08–1 μM primers was amplified in 25 μl volumes using the following PCR conditions: 1 cycle of 95°C for 10 min; 13 cycles of 95°C for 30 sec., 67°C (-1°C/cycle) for 45 sec. and 72°C for 45 sec.; 20 cycles of 95°C for 30 sec., 55°C for 45 sec. and 72°C for 45 sec.; 1 cycle of 72°C for 10 min. A subset of the multiplex PCR products were analysed on a 2100 Bioanalyzer (Agilent Technologies, Santa Clara, CA) to test the performance of the multiplex PCR. The multiplex PCR products from each sample were pooled for further analysis.
The pooled PCR products from each sample were treated with Exonuclease I and shrimp alkaline phosphatase and used for template in the SBE reaction (mini-sequencing) as described by Lindross et al. [
'Anti-tag' oligonucleotides complimentary to the 'tag' sequences of the SBE primers modified with NH2 groups in their 3' ends, and containing a 3'-spacer of 15 T residues were coupled covalently to CodeLink Activated Slides (Amersham Biosciences, Uppsala, Sweden) according to manufacturer instructions. The only exception was that the 'anti-tag' oligonucleotides were dissolved in 150 mM sodium carbonate buffer, pH 8.5 with 1 mM betaine at a concentration of 20 μM. The oligonucleotides were printed onto the slides using a VersArray ChipWriter (BioRad, Hercules, CA) with 3 Stealth Micro Spotting Pins (TeleChem International Inc., Sunnyvale, CA). Each slide consists of 75 sub-arrays with 13 × 12 spots in each subarray. The spots were 130–150 μM in diameter and the centre-to-centre distance between two spots was 200 μM. The 'anti-tag' oligonucleotides were printed in duplicates in each sub-array. The spot quality was tested after each series of microarray preparation using a Cy3 labelled random oligonucleotide hybridizing to all spots on the slide independent on the 'anti-tag' sequence. The printed slides were stored at room temperature until use.
The slides printed with 'anti-tags' were pre-heated to 42°C in a custom-made aluminium reaction rack with a re-usable silicon rubber grid placed on the slides to form 75 separate reaction chambers on each slide [
The signal detection was performed mainly as described by Lindroos et al. [
A Bonferroni corrected Fisher exact test was used to test for association between the genotypes of each SNP and the familiar CRC cohort and the cohort of sporadic CRC, respectively. Fifteen tests were made for single marker associations, in the cohort of familiar CRC, yielding a Bonferroni corrected significance level of 0.05/15 = 0.0033. Likewise the Bonferroni corrected
The power was calculated by applying the Fisher exact test to 10000 independent simulated cases with the given odds ratio and frequency of the disease causing genotype. In that case the power is the proportion of the simulations that reach a
Each variant was tested for deviation from Hardy-Weinberg proportions in the sub-cohort, using the exact test described by Wigginton et al. [
Linkage disequilibrium (LD) was investigated using HaploView (v. 3.31), [
In
Single base extension (SBE)-tag arrays is a well described method for analysing single nucleotide polymorphisms [
We have identified the frequency of 35 variants in
The characteristics of the cohort of sporadic CRC cases and the sub-cohort are shown in Table
Genotype frequencies of the variants in the analyzed cohorts
| Id | Variant | Sporadic cases | Familiar CRC cohort | Sub-cohort | ||||||
|
|
HoWt | He | HoMut | HoWt | He | HoMut | HoWt | He | HoMut | |
|
|
||||||||||
| 23 |
|
|
|
|
|
|
|
|
|
|
| 24 | c. 350 C>T | 1.000 | 0.000 | 0.000 | 0.997 | 0.003 | 0.000 | 1.000 | 0.000 | 0.000 |
| 25 |
|
|
|
|
|
|
|
|
|
|
| 26 |
|
|
|
|
1.000 | 0.000 | 0.000 | 1.000 | 0.000 | 0.000 |
| 27 |
|
|
|
|
|
|
|
|
|
|
| 28 |
|
|
|
|
|
|
|
|
|
|
| 29 | c.884+39 G>A | 1.000 | 0.000 | 0.000 | 0.997 | 0.003 | 0.000 | 1.000 | 0.000 | 0.000 |
| 30 | c.884+83_84 ins T | 1.000 | 0.000 | 0.000 | 1.000 | 0.000 | 0.000 | 1.000 | 0.000 | 0.000 |
| 85 |
|
1.000 | 0.000 | 0.000 |
|
|
|
|
|
|
| 31 | c.1379 A>C | 1.000 | 0.000 | 0.000 | 0.994 | 0.006 | 0.000 | 1.000 | 0.000 | 0.000 |
| 86* | c.1558+11 G>A | 1.000 | 0.000 | 0.000 | 1.000 | 0.000 | 0.000 | 1.000 | 0.000 | 0.000 |
| 32 |
|
|
|
|
|
|
|
|
|
|
| 37 |
|
|
|
|
|
|
|
|
|
|
| 101 | c.1689 A>G | 1.000 | 0.000 | 0.000 | 1.000 | 0.000 | 0.000 | 1.000 | 0.000 | 0.000 |
| 33 | c.1732-2 A>T | 1.000 | 0.000 | 0.000 | 0.997 | 0.003 | 0.000 | 1.000 | 0.000 | 0.000 |
| 34 |
|
|
|
|
|
|
|
|
|
|
| 100 | c.1942 C>T | 1.000 | 0.000 | 0.000 | 0.997 | 0.003 | 0.000 | 1.000 | 0.000 | 0.000 |
| 35 |
|
|
|
|
|
|
|
|
|
|
| 36 | c.2152 C>T | 1.000 | 0.000 | 0.000 | 0.994 | 0.006 | 0.000 | 1.000 | 0.000 | 0.000 |
|
|
||||||||||
| 89 |
|
|
|
|
na | na | na |
|
|
|
| 87 | c.131 C>T | 1.000 | 0.000 | 0.000 | 1.000 | 0.000 | 0.000 | 1.000 | 0.000 | 0.000 |
| 88 | c.134 C>T | 1.000 | 0.000 | 0.000 | 1.000 | 0.000 | 0.000 | 1.000 | 0.000 | 0.000 |
| 39 | c.212-23 A>C | 1.000 | 0.000 | 0.000 | 1.000 | 0.000 | 0.000 | 1.000 | 0.000 | 0.000 |
| 90* | c.287 G>A | 1.000 | 0.000 | 0.000 | 1.000 | 0.000 | 0.000 | 1.000 | 0.000 | 0.000 |
| 91* | c.329 A>G | 1.000 | 0.000 | 0.000 | 1.000 | 0.000 | 0.000 | 1.000 | 0.000 | 0.000 |
| 92* | c.380 A>G | 1.000 | 0.000 | 0.000 | 1.000 | 0.000 | 0.000 | 1.000 | 0.000 | 0.000 |
| 93 | c.560 T>G | 1.000 | 0.000 | 0.000 | 1.000 | 0.000 | 0.000 | 1.000 | 0.000 | 0.000 |
| 41 |
|
|
|
|
|
|
|
|
|
|
| 43 |
|
|
|
|
|
|
|
|
|
|
| 48 | c.1786_1788 del AAT | 1.000 | 0.000 | 0.000 | 1.000 | 0.000 | 0.000 | 1.000 | 0.000 | 0.000 |
| 95* | c.2006-6 T>C/G | 1.000 | 0.000 | 0.000 | 1.000 | 0.000 | 0.000 | 1.000 | 0.000 | 0.000 |
| 96 | c.2062 A>G | 1.000 | 0.000 | 0.000 | 1.000 | 0.000 | 0.000 | 1.000 | 0.000 | 0.000 |
| 97* | c.2139 G>C | 1.000 | 0.000 | 0.000 | 1.000 | 0.000 | 0.000 | 1.000 | 0.000 | 0.000 |
| 50 | c.2500 G>A | 1.000 | 0.000 | 0.000 | 1.000 | 0.000 | 0.000 | 1.000 | 0.000 | 0.000 |
| 51 | c.2542 G>T | 1.000 | 0.000 | 0.000 | 1.000 | 0.000 | 0.000 | 1.000 | 0.000 | 0.000 |
* identified in CRC families from other populations than the Danish
§ na: not analyzed
Variants that where polymorphic in sporadic CRC cohort or in the sub-cohort are in bold
| Id | Variant | Sporadic-Sub-cohort | Familiar-Sub-cohort | Sporadic-familiar | All |
|
|
|||||
| 23 | c.307-29 C>A | 0.7259 | 1.0000 | 0.6673 | 0.8632 |
| 25 | c.453+79 A>G | 0.7908 | - | - | - |
| 24 | c.350 C>T | - | 0.2988 | - | - |
| 26 | c.545+43 C>G | 0.3263 | 1.000 | 1.000 | 0.4741 |
| 27 | c.655 A>G | 0.7715 | 0.5619 | 0.4869 | 0.7489 |
| 28 | c.790+10 A>G | 1.0000 | - | - | - |
| 29 | C884+39G>A | - | 0.2943 | - | - |
| 85 | c.1217 G>A | 1.0000 | 0.5359 | 0.4858 | 0.4833 |
| 31 | c.1379 A>C | - | 0.1880 | - | - |
| 32 | c.1558+14 G>A | 1.0000 | 0.6063 | 0.7685 | 0.8722 |
| 37 |
|
0.5382 |
|
0.0329 | 0.0148 |
| 33 | c.1732-2 A>T | - | 0.3109 | - | - |
| 34 | c.1852_1853 AA>GC | 0.2408 | 0.5672 | 0.6760 | 0.3569 |
| 100 | c.1942 C>T | - | 0.4747 | ||
| 35 | c.1959 G>T | 0.2473 | 1.0000 | 0.3642 | 0.4493 |
| 36 | c.2152 C>T | - | 0.0840 | - | - |
|
|
|||||
| 89 |
|
|
- | - | - |
| 41 | c.965 G>A | 0.2574 | 0.0244 | 0.3609 | 0.0582 |
| 43 | c.1511-9 A>T | 0.0958 | 0.7314 | 0.0796 | 0.1617 |
The Bonferroni corrected
The Bonferroni corrected
Polymorphic variants with borderline significant
Twenty-two of the variants analysed in the present study were neither detected in the sporadic cases nor in the sub-cohort (Tables
Linkage disequilibrium (LD) analyses of the genomic regions covering
| Id | Variant | PMUT | SIFT | PolyPhen | Phylogeny | Activity in functional assays | References |
| MLH1 | |||||||
| 24 | p.Thr117Met | Neutral | Not tolerated | Possible damaging | Conserved | Aberrant | [11,12] |
| 27 | p.Val219Ile | Neutral | Tolerated | Benign | Ile other species | Normal | [11,12] |
| 85 | p.Ser406Asn | Neutral | Tolerated | Benign | Non-conserved | Normal | [11] |
|
|
|
Pathogen | Tolerated | Benign | Non-conserved | NA | - |
|
|
|
Neutral | Tolerated | Possible damaging | Non-conserved | NA | - |
| 34 | p.Lys618Ala | Pathogen | Not tolerated | Possible damaging | Non-conserved | Normal/aberrant | [9,11,45–47] |
| 100 | p.Pro648Ser | Pathogen | Not tolerated | Probably damaging | Conserved | Normal (aberrant protein stability) | [9,67] |
| 36 | p.His718Tyr | Neutral | Not tolerated | Probably damaging | Conserved | Normal | [11] |
| MSH2 | |||||||
| 89 | p.Thr44Met | Neutral | Not tolerated | Possible damaging | Conserved | NA | - |
| 87 | p.Ala45Val | Neutral | Tolerated | Benign | Val other species | NA | - |
| 90 | p.Arg96His | Pathogen | Tolerated | Probably damaging | Non-conserved | NA | - |
| 91 | p.Lys110Arg | Neutral | Tolerated | Benign | Non-conserved | NA | - |
| 92 | p.Asn127Ser | Neutral | Not tolerated | Probably damaging | Conserved | NA | - |
| 93 | p.Leu187Arg | Neutral | Not tolerated | Probably damaging | Conserved | NA | - |
| 41 | p.Gly322Asp | Neutral | Tolerated | Benign | Conserved | Slightly reduced | [8,44] |
|
|
|
Neutral | Not tolerated | Probably damaging | Conserved | NA | - |
| 50 | p.Ala834Thr | Pathogen | Tolerated | Possible damaging | Non-conserved | Normal | [10] |
|
|
|
Neutral | Not tolerated | Possible damaging | Conserved | NA | - |
NA, not available
Danish variants that have not been described previously are in bold
There was no overall concordance between the
The putative role of all 35 variants in pre-mRNA splicing was analyzed using SNAP [
Identification of missense, silent and non-coding variants in genes involved in hereditary diseases always raises the intriguing question whether these variants are the disease causing mutations in the family/families where they are identified. Alternatively, they may be common variants causing a slight increase in sporadic disease susceptibility in the general population or simple neutral variants that are not involved in disease development. Missense, silent and non-coding variants are identified frequently in MMR genes (e.g.
In the present study, we have used a case-cohort design to elucidate the possible association between 35 variants in
Thirteen variants were polymorphic in the present study. The majority of the polymorphic variants (i.e. 9/13) were silent or present in non-coding regions. This corresponds with common variants being more abundant in introns and other regions than in coding regions [
Linkage disequilibrium (LD) analysis showed high LD in the genomic regions covering
The age of the cohort of cases with sporadic CRC is relatively young (mean ~58 years) compared to the mean age of onset of colorectal CRC in the general Danish population (mean ~70 years) [
Using
Twenty two variants were not detected in the sporadic CRC cases and in the sub-cohort. Among those were the six variants originally identified more or less frequently in other populations (Table
In conclusion, high penetrance cancer susceptibility genes involved in hereditary syndromes have rarely emerged as definitive low-penetrance genes as a result of common variants increasing disease susceptibility [
More than half of the analysed variants in
The authors declare that they have no competing interests.
LLC: Participated in the design of the study, carried out the genotyping using SBE-tag arrays and coordinated and drafted the manuscript. BEM: Carried out the statistical analyses and assisted in drafting the manuscript. FPW: Participated in the design of the study and assisted in drafting the manuscript. CW: Carried out the statistical analyses and assisted in drafting the manuscript. KK: Participated in the design of the study and the set-up of the SBE-tag arrays. AT: Responsible for the Danish Diet, Cancer and Health study. AO: Participating in the Danish Diet, Cancer and Health study. A–CS: Designed the SBE-tag array analysis for genotyping of the variants, CLA: Participated in the design of the study and assisted in drafting the manuscript. TFØ: Participated in the design of the study and assisted in drafting the manuscript. All authors read and approved the final version of the manuscript
The pre-publication history for this paper can be accessed here:
We are especially grateful to Gitte Høj for skilful technical assistance. This study was supported by grants from The Danish Cancer Society, Dansk Kræftforsknings Fond, A.P. Møller og hustru Chastine Mc-Kinney Møllers Fond, Beckett-Fonden, Eva & Henry Frænkels Mindefond, Helga og Peter Kornings Fond and Th. Maigaards Eftf. Fru Lilly Benthine Lunds Fond.