OMT (O-methyltransferase) genes are involved in lignin biosynthesis, which relates to stover cell wall digestibility. Reduced lignin content is an important determinant of both forage quality and ethanol conversion efficiency of maize stover.
Variation in genomic sequences coding for
Several polymorphic sites for three O-methyltransferase loci were associated with stover cell wall digestibility. All three tested genes seem to be involved in controlling DNDF, in particular
Stover cell-wall digestibility has long been shown to be crucial for forage quality, and more recently this trait is getting more attention in relation to biofuel production. Research into bioethanol has grown significantly as a response to global warming and increasing prices of fossil fuels. The conversion of lignocellulosic biomass into fermentable sugars by the addition of enzymes has long been recognized as an alternative to the existing starch-based ethanol production [
Lignin is synthesized by the phenylpropanoid pathway [
Four brown midrib mutations in maize (
Several studies reported improvement in the digestibility of corn silage in ruminant feeding of
Although many studies investigated phenotypic effects in relation to down regulation of lignin genes for maize and other species [
In this study, variation in genomic sequences coding for
The 40 lines within Exp. 1 represent a broad range of Central European forage maize germplasm, and were extremes within a larger collection of >300 maize inbreds with respect to stover cell-wall digestibility (unpublished data). Thirty-five lines originated from the current breeding program of KWS Saat AG and five lines were from the public domain (AS01, AS02, AS03, AS39, and AS40, identical to F7, F2, EP1, F288, and F4, respectively).
The inbred lines were evaluates in Grucking (sandy loam) in 2002, 2003, and 2004, and in Bernburg (sandy loam) in 2003 and 2004. The experiments consisted of 49 entries in a 7 × 7 lattice design with two replications. Nine entries from this design were consisting of not sequenced checks (such as knock-out mutants of brown midrib loci with known high cell wall digestibility). Therefore, only 40 lines were analyzed for cell wall properties and sequenced. The single row plots were 0.75 m apart and 3 m long with a total of 20 plants. The ears were manually removed and the stover was chopped 50 days after flowering. Approximately 1 kg of stover was collected and dried at 40°C. The stover was ground to pass through a 1 mm sieve. Quality analyses were performed with near infrared reflectance spectroscopy (NIRS) based on previous calibrations on the data of 300 inbred lines (unpublished results). Four traits were analyzed: Water Soluble Carbohydrate (WSC), Organic Matter Digestibility (OMD), Neutral Detergent Fiber (NDF) and digestible Neutral Detergent Fiber (DNDF). DNDF was estimated by DNDF = 100 - (IVDMD - (100 - NDF))/NDF based on Goering and Van Soest, 1970 [
Thirty-four inbred lines were employed in Exp. 2, including public lines and ecotypes both from Europe and U.S., covering substantial variation in forage quality. The lines were evaluated in three different years: 2006, 2008, and 2009 (unpublished data of INRA Lusignan). Values of stover
The inbred lines from Exp. 1 were grown in the greenhouse for DNA isolation. Leaves were harvested three weeks after germination, and DNA was extracted using the Maxi CTAB method [
Products were separated by gel electrophoresis on 1.5% agarose gels, stained with ethidium bromide and photographed using an eagle eye apparatus (Herolab, Wiesloch, Germany). Amplicons were purified using QiaQuick spin columns (Qiagen, Valencia, USA) according to the manufacturer instructions, and directly sequenced using internal sequence specific primers and the Big Dye1.1 dye-terminator sequencing kit on an ABI 377 (PE Biosystems, Foster City, USA). Electropherograms of overlapping sequencing fragments were manually edited using the software Sequence Navigator version 1.1, from PE Biosystems.
Full alignments were built for
In Exp. 2, primer pairs for
Lines evaluated in Exp. 1 were genotyped with 101 simple sequence repeat (SSRs) markers providing an even coverage of the maize genome. Population structure was estimated from the SSR data by the
Association between polymorphisms and mean phenotypic values were performed by the General Linear Model (GLM) analyses in
The same parameters were applied to perform the Logistic Regression in
Associations were also tested by the Mixed Linear Model (MLM) in
Significant intron polymorphisms in Exp.1 were analyzed for alterations in motif sequences, using the inbred line W64 as reference sequence (AY323283). Differential splicing was tested by comparing expressed sequences with genomic sequences. The discrimination of favorable and unfavorable alleles at each significant associated polymorphic site at
Guillet-Claude et al. [
The analysis of overlapping sequences of common inbred lines between Exp.1 and Exp.2 (F2, F4, F7, F288, EP1, and W64) revealed consistency of sequences for the
Comparison of sequences of common inbred lines in Exp. 1 and Exp 2
|
|
|
|
||||
|---|---|---|---|---|---|---|
| Inbred lines | Bp Overlap | Differences | Bp Overlap | Differences | Bp Overlap | Differences |
| F2 | 2051 | 0 | 1326 | 0 | - | - |
| F4 | 1947 | 1 SNP | 1327 | 5 SNP, 1 Indel | - | - |
| F7 | 2090 | 0 | 1249 | 0 | - | - |
| F288 | 2052 | 0 | 1328 | 1 SNP, 1 Indel | 655 | 1 SNP, 1 Indel |
| EP1 | 2115 | 1 SNP | 1254 | 9 SNP, 4 Indel | 751 | 0 |
| W64 | 2088 | 0 | - | - | - | - |
Analysis of haplotype diversity for the 40 inbred lines employed in Exp. 1 revealed 13, 9, and 6 haplotypes for
Number of haplotypes based on single nucleotide polymorphisms (SNPs) in the
| WSC | NDF | OMD | DNDF | ||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Gene | No Haplotypes | Min | Max | Range | Min | Max | Range | Min | Max | Range | Min | Max | Range |
|
|
13 | 13.69 | 23.68 | 9.99 | 52.92 | 65.90 | 12.98 | 61.55 | 76.65 | 15.1 | 42.18 | 59.98 | 17.8 |
|
|
9 | 16.04 | 23.28 | 7.24 | 52.93 | 61.93 | 9.00 | 68.02 | 75.55 | 7.53 | 50.20 | 59.98 | 9.78 |
|
|
6 | 13.28 | 20.69 | 7.41 | 54.95 | 63.03 | 8.08 | 71.00 | 74.46 | 3.46 | 54.07 | 59.98 | 5.91 |
Varying alignment lengths and number of polymorphic sites were observed when four different parameter sets were applied for aligning
All subsequent analyses were performed for default sequence alignment settings, in order to be able to compare findings reported here with previous studies [
Significantly associated polymorphic sites of
| Site | Alignments | |||
|---|---|---|---|---|
| 1 (2461 bp) | 2 (2517 bp) | 3 (2508 bp) | 4 (2553 bp) | |
|
|
X | X | X | X |
|
|
X | X | X | X |
|
|
X | |||
|
|
X | X | X | |
|
|
X | X | X | |
|
|
X | |||
|
|
X | X | X | |
|
|
X | X | ||
|
|
X | X | ||
|
|
X | |||
|
|
X | X | X | X |
|
|
X | X | X | X |
|
|
X | X | X | X |
|
|
X | X | X | X |
|
|
X | |||
|
|
X | |||
|
|
X | |||
|
|
X | |||
|
|
X | |||
|
|
X | |||
|
|
X | |||
|
|
X | |||
|
|
X | |||
|
|
X | X | X | X |
|
|
X | X | X | X |
Site numbers denote bp alignment positions of individual SNPs and starting positions of Indels based on reference sequence
Polymorphic sites of
| Site | Snp/Indel | E/I | aa change | GLM | MLM | REG | Other trait* |
|---|---|---|---|---|---|---|---|
| 1233 | C-G | I | - | - | - | - | WSC |
| 1235 | A-T | I | - | - | - | - | WSC |
| 1236 | C-G | I | - | - | - | - | WSC |
| 1240 | C-T | I | - | - | - | - | WSC |
| 1243 | 10 | I | - | - | - | - | WSC |
| 1261 | A-G | I | - | - | - | - | WSC |
| 1296 | C-T | I | - | X | X | X | OMD |
| 1331 | C-T | I | - | X | X | X | OMD |
| 1377 | A-T | I | - | X | X | X | OMD |
| 1381 | 7 | I | - | X | X | X | OMD |
| 1439 | A-C | I | - | X | - | X | - |
| 1449 | 4-8 | I | - | X | X | - | - |
| 1547 | 6 | I | - | X | X | X | OMD |
| 1589 | A-G | I | - | X | X | X | OMD |
| 1638 | 1 | I | - | X | - | X | OMD |
| 1811 | 6 | I | - | X | X | X | OMD |
| 1902 | 1 | I | - | - | - | X | - |
| 1907 | 3 | I | - | - | - | X | - |
| 1916 | A-G | I | - | - | - | X | - |
| 1917 | A-C | I | - | X | X | X | - |
| 1918 | A-G | I | - | X | X | X | - |
| 1919 | C-T | I | - | - | - | X | - |
| 1920 | 28 | I | - | X | X | X | OMD |
| 1948 | C-T | I | - | - | - | X | - |
| 1952 | C-T | I | - | - | - | X | - |
| 1953 | A-T | I | - | - | - | X | - |
| 1954 | 77 | I | - | X | X | X | OMD |
| 2032 | A-C | I | -- | - | - | X | - |
| 2103 | C-T | E2 | Ser - Pro | - | - | X | OMD |
| 2178 | C-G | E2 | His - Asp. | X | X | X | OMD |
| 2185 | C-G | E2 | Arg - Pro | X | X | X | - |
| 2693 | A-C | E2 | Syn. | - | - | - | NDF |
*All polymorphic sites associated with WSC, OMD and NDF were identified only by Logistic Regression. E: exon; I: intron; aa: Amino Acid; Syn.: Synonymous substitution
Site numbers denote bp alignment positions of individual SNPs and starting position of Indels based on reference sequence
No polymorphic sites were observed to be significantly associated with DNDF in the first exon, and most of the polymorphic sites identified for this trait were located in the intron, where 13 SNPs and 9 Indels were detected. In the second exon, three SNPs were significantly associated with DNDF in positions 2103, 2178, and 2185 bp, each leading to amino acid substitutions (Ser/Pro, His/Asp, and Arg/Pro, respectively). Most of the significantly associated polymorphic sites identified for DNDF were in high LD (Fig.
No significant associated SNP or Indel was identified between
Eight significantly associated Indels in the
Association analysis for
Significantly associated polymorphic sites of
| Gene | Trait | Site | Snp/Indel | E/I | aa change | GLM | MLM | REG |
|---|---|---|---|---|---|---|---|---|
|
|
NDF | 726 | C-T | I3 | - | X | - | - |
| NDF | 944 | C-T | I4 | - | X | - | - | |
| OMD | 944 | C-T | I4 | - | X | - | - | |
| WSC | 1299 | C-G | E5 | Ala - Gly | - | - | X | |
|
|
||||||||
|
|
WSC | 404 | 1 | I2 | - | - | - | X |
| WSC | 414 | 4 | I2 | - | X | - | X | |
E: exon; I: intron; aa: Amino Acid.
Site numbers denote bp alignment positions of individual SNPs and starting position of Indels. (Based on reference sequences AY323264 and AY279022, for
Only two Indels showed significant trait associations for
Guillet-Claude et al. [
After taking population structure into account, five sub-populations were identified for this line panel by Camus-Kulandaivelu et al. [
Significant polymorphic sites identified in Exp. 2 in comparison to Exp. 1
| Gene | Exp.2 polymorphisms | Exp.1 | |||
|---|---|---|---|---|---|
|
|
|||||
| Site | Region | SNP/Indel | aa Change | p-value | |
| 342 | Prom | A-T | - | - | |
|
|
659* | Prom | 1 | - | - |
| 749* | E1 | 6 | - | 0.99 | |
| 1948* | I1 | G-T | - | 0.03 | |
| 1981* | I1 | C-T | - | 0.92 | |
|
|
|||||
|
|
704* | I3 | 3 | - | - |
| 734* | I3 | 2 | - | 1 | |
| 944* | I4 | C-G | - | 0.01 | |
| 956* | I4 | 11 | - | - | |
| 972* | I4 | A-G | - | - | |
|
|
|||||
|
|
187* | E1 | A-G | His - Arg | - |
| 688* | I3 | A-G | - | - | |
| 717* | I3 | C-T | - | - | |
| 720* | I3 | C-T | - | - | |
E: exon; I: intron; Prom: promoter; aa: Amino Acid; Syn.: Synonymous substitution; *: singleton
Site numbers denote bp alignment position of individual SNPs and starting position of Indels based on the reference sequences AY323283, AY323264 and AY279022, for
For
Contrary to the findings of Guillet-Claude et al. [
Changing alignments parameters for gap opening and gap extension costs leads to the creation of alignments with different Indel sizes and frequencies [
In this study, cell wall traits of two relatively small populations (40 and 34 lines) were tested for associations with two large sets of polymorphic sites (SNPs and Indels). The large number of independent variables (SNPs/Indels) in relation to number of dependent variables (lines/phenotypic data) increases the chances of false positive associations due to overfitting [
The control of systematic differences due to population structure and kinship has also been suggested to avoid spurious associations and for greatest statistical power in association studies [
Across the investigated genes in this study, most of the significantly associated polymorphisms were detected by Logistic Regression, followed by GLM and MLM. A few previous studies compared results of association analyses from these methods. Andersen et al. [
Across the three genes and four traits within Exp. 1, 13 polymorphisms were consistent across the three methods, while one polymorphism was identified by both GLM and MLM, 3 by both GLM and Logistic Regression. No significant polymorphic sites were consistent between both MLM and Logistic Regression. Logistic Regression was the method detecting the highest number of significant polymorphisms in our study. With this method, permutation tests are performed for individual markers, not controlling for experiment-wise error rates. Consequently, a larger number of false positives are expected for this method of analysis when compared to analyses that control for experiment-wise error, like GLM in our study. Logistic Regression (and GLM in one occasion) revealed significant associations between the three OMT genes with WSC, with none of these three genes being involved in biosynthesis of soluble carbohydrates (Tables
Two SNPs were identified as significantly associated in both datasets (Exp.1 and Exp.2). The SNP at position 1948 bp in the
In addition, three polymorphic sites identified in Exp. 2 were also observed in Exp.1, but not significantly associated with the investigated traits. The relative low number of common polymorphic sites between Exp.1 and Exp.2 can in part be explained by the different regions of genes that each experiment investigated. The overlapping region of each gene common to both studies was not representing the whole gene. Moreover, for some lines overlapping sequences were not completely identical in both studies (Table
Most of the significant SNPs/Indels identified in Exp. 1 were identified for
Andersen et al. [
The association analyses revealed a considerable number of SNPs and Indels associated with cell wall traits for the genes analyzed. Across experiments, methodologies and genes, 11 significant QTNs/Indels were located in coding regions and 52 significant polymorphisms were located in non-coding regions. Some of the significant associations identified in this study may reflect a statistical artifact. The probability of false-positive associations is higher in case of rare alleles and cases where a polymorphic site was only detected with one of the alignments. In addition, several of the identified significant associations can likely be explained by high LD to closely linked causal QTN/QTINDELs. Within an unresolved linkage block, likely only one or few polymorphisms are causative (Figs.
Several significantly associated polymorphisms located in non-coding regions have been reported. Andersen et al. [
Lines containing favorable alleles at putative QTN or QTINDELs were generally, but not always associated with high phenotypic values for cell wall digestibility. As cell wall digestibility is a polygenic inherited quantitative trait, likely unfavorable alleles at other loci affecting this trait are responsible for this finding, either by strong main effects or interaction with
CCoAOMT: caffeoyl-CoA
EAB prepared the manuscript. EAB, JRA, and YC performed data analysis. IZ carried out allele sequencing. GW contributed to experimental design. BD and JE provided phenotypic data. MO provided the SSR data and together with GW contributed to experimental design. UKF performed sequence alignments. YB provided data from Exp.2 and reviewed the manuscript. TL coordinated the project and together with EAB prepared the manuscript. All authors read and approved the final manuscript.
We thank KWS Saat AG (Einbeck) and the German ministry for education and science (BMBF) for financial support of the EUREKA project Cerequal. EAB and YC are supported by the RF Baker Center for Plant Breeding at Iowa State University.