Unigene sequences constitute a rich source of functionally relevant microsatellites. The present study was undertaken to mine the microsatellites in the available unigene sequences of sugarcane for understanding their constitution in the expressed genic component of its complex polyploid/aneuploid genome, assessing their functional significance
The average frequency of perfect microsatellite was 1/10.9 kb, while it was 1/44.3 kb for the long and hypervariable class I repeats. GC-rich trinucleotides coding for alanine and the GA-rich dinucleotides were the most abundant microsatellite classes. Out of 15,594 unigenes mined in the study, 767 contained microsatellite repeats and for 672 of these putative functions were determined
Microsatellite repeats were present in 4.92% of sugarcane unigenes, for most (87.6%) of which functions were determined
Sugarcane (
In sugarcane, Cordeiro
In India, systematic breeding of sugarcane has resulted in the development of a number of varieties with high productivity and stress tolerance by inter-specific hybridization [
The present study was undertaken to mine the available unigene sequences of sugarcane (
The type, frequency and relative distribution of the microsatellites in the unigene sequences of sugarcane are given in Table
Distribution of microsatellites in the unigene sequences of sugarcane
| Characters under study | Unigenes* |
|---|---|
| Number of sequences examined | 15,594 |
|
|
|
| Size (bp) of examined sequences | 9,17,43,95 |
|
|
|
| Number of identified perfect microsatellites | 2,712 (17.4) |
|
|
|
| Number of perfect microsatellite containing sequences | 2,230 (14.3) |
|
|
|
| Number of perfect microsatellite (excluding mononucleotides) containing sequences | 584 (3.7) |
|
|
|
| Number of sequences containing more than one perfect microsatellites | 167 (28.6) |
|
|
|
| Number of sequences containing single and unique perfect microsatellites | 417 (71.4) |
|
|
|
| Number of mononucleotides | 1,871 (12) |
|
|
|
| Number of dinucleotides | 200 (23.8) |
|
|
|
| Number of trinucleotides | 615 (73.1) |
|
|
|
| Number of tetranucleotides | 15 (1.8) |
|
|
|
| Number of pentanucleotides | 7 (0.83) |
|
|
|
| Number of hexanucleotides | 4 (0.47) |
|
|
|
| Number of perfect microsatellites excluding mononucleotides | 841 (5.4) |
|
|
|
| Size (kb) of sequences containing one perfect microsatellite | 10.9 |
|
|
|
| Number of perfect class I microsatellites | 207 (24.6) |
|
|
|
| Size (kb) of sequences containing one perfect class I microsatellite | 44.3 |
|
|
|
| Number of primer pairs designed for perfecta microsatellites | 810 (96.3) |
|
|
|
| Number of compound class I microsatellite containing sequences | 183 |
|
|
|
| Size (kb) of sequences containing one compound class I microsatellite | 50.1 |
|
|
|
| Number of compound class I microsatellites | 183 (1.2) |
|
|
|
| Number of compound interrupting class I microsatellites | 137 (74.9) |
|
|
|
| Number of compound non-interrupting class I microsatellites | 46 (25.1) |
|
|
|
| Number of primer pairs designed for compound class I microsatellites | 151 (82.5) |
*The number in the bracket is the proportion expressed in percentage
aMononucleotides to hexanucleotides repeated up to 100 times without any interruption at a locus
Out of 841 perfect microsatellites identified in sugarcane, 587 (69.8%) were found in the ORFs and the remaining were present either in the 3'UTRs (102, 12.1%) or in the 5'UTRs (152, 18.1%). The trinucleotide repeat motifs were significantly more frequent (about 86%) in the ORFs. In contrast, the GA-rich dinucleotide repeat motifs were more in the 5' (49%) and 3' (32%) UTRs. The density of longer motif containing perfect class I microsatellites was one in every 44.3 kb sequences, which accounted for 24.6% (207) of the total 841 microsatellites identified (Table
The primer pairs could be designed for 810 perfect microsatellites that was 96.3% of the total microsatellites (841) identified in the present investigation. The primer sequences flanking all the perfect UGMS including 207 class I microsatellites with their Tm values and product sizes are given in the Additional file
Evaluation of the amplification efficiency and polymorphic potential of 47 fluorescent dye labeled primers
| Polymorphic potential | ||||||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
|
|
||||||||||||||||
| Sl. No. | Unigene Accession IDs' | Class I |
Repeat-motifs | Location | Forward primer sequences (5'-3') | Reverse primer sequences (5'-3') | Putative functions | Actual annealing temperature (°C) | No. of locus | Total no. of alleles amplified | Size (bp) of allele (s) amplified | Type of allele size distribution | No. of heterozygous loci | Among sugarcane species and genera | Among sugarcane varieties | |
|
|
||||||||||||||||
| P/MB | PIC | |||||||||||||||
| 1 | CA297715 | UGSuM2a | (AT)43 | CDS | CTGTGTATATGTTCGTAGTTTG | CACTTAGTCACACTCTCACACAC | Sucrose phosphate synthase | 55 | 2 | 20 | 206-302 | Step-wise | 1 | P | P | 0.86 |
|
|
|
|||||||||||||||
| UGSuM2b | 4 | 481-512 | Mixed | |||||||||||||
|
|
||||||||||||||||
| 2 | CA278792 | UGSuM5 | (TA)28 | CDS | TCACATCCATCATCCACAGC | TCCAATGCAAGCAAACTCAC | Maize-Cyclin III | 55 | 1 | 23 | 80-170 | Step-wise | 0 | P | P | 0.84 |
|
|
||||||||||||||||
| 3 | AY596609 | UGSuM11 | (TA)21 | 3'UTRs | TGGTAACCCTAGGCAGGTGA | GTGCACCAGATTTGGATGGT | Fructose-bisphosphate aldolase | 56 | 1 | 23 | 92-190 | Step-wise | 0 | P | P | 0.83 |
|
|
||||||||||||||||
| 4 | CA279221 | UGSuM15a | (CCTCGC)6 | CDS | GTTTAAGACAAGATGGTGTAGATG | TACATATTTACATTGTTACTCCGC | Hypothetical protein | 56 | 2 | 17 | 170-247 | Mixed | 1 | P | P | 0.80 |
|
|
|
|||||||||||||||
| UGSuM15b | 3 | 360-378 | Step-wise | |||||||||||||
|
|
||||||||||||||||
| 5 | CA278282 | UGSuM16 | (AT)18 | CDS | GCGTCTTCATCATCTGCAAC | GCGTCTTCATCATCTGCAAC | Pathogenesis-related PRMS protein | 55 | 1 | 22 | 150-244 | Step-wise | 0 | P | P | 0.82 |
|
|
||||||||||||||||
| 6 | CA253277 | UGSuM17 | (AG)18 | CDS | TTTCCATTCTTCCATTCAACTG | GGCAGGCTGAGAGACTGTTC | Abcisic acid-protein kinase | 55 | 1 | 10 | 90-188 | Step-wise | 1 | P | P | 0.78 |
|
|
||||||||||||||||
| 7 | CA227482 | UGSuM18a | (GA)18 | CDS | GGCGAGAGAGAGAGAGAGAGAG | AGGTGGAGATCTTGAGGTAGGC | Glycine decarboxylase | 56 | 2 | 20 | 96-193 | Mixed | 1 | P | P | 0.81 |
|
|
|
|||||||||||||||
| UGSuM18b | 2 | 326-348 | Step-wise | |||||||||||||
|
|
||||||||||||||||
| 8 | CA126180 | UGSuM20a | (TCA)12 | CDS | ATCCCTTATGCTACAGAAATGT | TTAGCCTAGAGGTTTGATTGAT | Acetyl-CoA synthetase | 54 | 2 | 16 | 190-259 | Step-wise | 1 | P | P | 0.83 |
|
|
|
|||||||||||||||
| UGSuM20b | 4 | 392-428 | Step-wise | |||||||||||||
|
|
||||||||||||||||
| 9 | CA177414 | UGSuM21 | (AGGA)9 | CDS | CGCTCCCTCACCGTCATT | CTCCGCATCCTCGTCACC | Transcription regulator protein | 62 | 1 | 20 | 160-252 | Step-wise | 0 | P | P | 0.81 |
|
|
||||||||||||||||
| 10 | CA223153 | UGSuM22a | (GCG)12 | 3'UTRs | CTCCCTCCTCCTCCCGTTG | CTCTTGGGTGTGAACCAG | Polyadenylate-binding protein | 64 | 2 | 19 | 96-187 | Mixed | 1 | P | P | 0.85 |
|
|
|
|||||||||||||||
| UGSuM22b | 3 | 322-358 | Step-wise | |||||||||||||
|
|
||||||||||||||||
| 11 | CA122659A | UGSuM24a | (TTTTC)7 | 5'UTRs | CTGTACAACAGCAATTATGAATCT | CTCGACTACGAGAGGATATGAT | Hypothetical protein | 55 | 3 | 12 | 240-294 | Mixed | 1 | P | P | 0.86 |
|
|
|
|||||||||||||||
| UGSuM24b | 4 | 398-428 | Step-wise | |||||||||||||
|
|
|
|||||||||||||||
| UGSuM24c | 5 | 531-571 | Step-wise | |||||||||||||
|
|
||||||||||||||||
| 12 | BU103692A | UGSuM26 | (CT)17 | 5'UTRs | CTCGATCCCAGAGAGCTCCACAG | AGTACCGAATTCATTAAACTCCT | Beta-amylase | 55 | 1 | 21 | 250-346 | Step-wise | 0 | P | P | 0.85 |
|
|
||||||||||||||||
| 13 | CA073284 | UGSuM27a | (GGC)11 | CDS | CTGCAGTACGGTCCGGAATC | GTACCACCATGGCTCTAGCTTC | 30 S ribosomal protein S16 | 60 | 2 | 2 | 50-54 | Mixed | 1 | P | P | 0.84 |
|
|
|
|||||||||||||||
| UGSuM27b | 22 | 154-288 | Mixed | |||||||||||||
|
|
||||||||||||||||
| 14 | AY596606 | UGSuM33a | (AGC)10 | CDS | CGAGGCACTGAACCCATATC | TGTTTGAACTGGATGGCGTA | Hypothetical protein | 58 | 3 | 17 | 92-189 | Mixed | 1 | P | P | 0.84 |
|
|
|
|||||||||||||||
| UGSuM33b | 2 | 293-305 | Step-wise | |||||||||||||
|
|
|
|||||||||||||||
| UGSuM33c | 2 | 431-491 | Step-wise | |||||||||||||
|
|
||||||||||||||||
| 15 | CA268640 | UGSuM34a | (AAG)10 | CDS | TTACAAATGTAGCCTTGCCTTG | ATCTTTCCTTGCTTGCCTCTC | Soluble acid invertase | 61 | 2 | 19 | 99-178 | Mixed | 1 | P | P | 0.83 |
|
|
|
|||||||||||||||
| UGSuM34b | 4 | 333-351 | Step-wise | |||||||||||||
|
|
||||||||||||||||
| 16 | CA131350 | UGSuM41 | (CCG)10 | CDS | ATCATTCTCCATCATTTCTCA | AGGCTCTTCAACCGTGCT | Unknown protein | 54 | 1 | 20 | 120-210 | Step-wise | 0 | P | P | 0.80 |
|
|
||||||||||||||||
| 17 | CA133924 | UGSuM42a | (CTCTCC)5 | 5'UTRs | TTCATACAGAAGAACCTCCAC | TCCATCAGAGACAAGCAGA | Auxin-independent growth promoter | 54 | 3 | 12 | 130-212 | Mixed | 1 | P | P | 0.83 |
|
|
|
|||||||||||||||
| UGSuM42b | 8 | 325-367 | Step-wise | |||||||||||||
|
|
|
|||||||||||||||
| UGSuM42c | 2 | 475-487 | Step-wise | |||||||||||||
|
|
||||||||||||||||
| 18 | CA136599 | UGSuM43a | (CCG)10 | 3'UTRs | CAAAGTGCTGTAGGGCTG | TTCAATGGGTGATAAGTGTGT | Ribose-phosphate pyrophosphokinase 1 | 55 | 2 | 17 | 90-187 | Mixed | 1 | P | P | 0.85 |
|
|
|
|||||||||||||||
| UGSuM43b | 2 | 332-368 | Step-wise | |||||||||||||
|
|
||||||||||||||||
| 19 | CA139800 | UGSuM44a | (CT)15 | 5'UTRs | TCCATCAAGCCGTTCCTC | GCCAAGCAGATAAAGAAGTG | Rudimentary enhancer | 55 | 2 | 16 | 220-308 | Step-wise | 1 | P | P | 0.84 |
|
|
|
|||||||||||||||
| UGSuM44b | 1 | 419 | - | |||||||||||||
|
|
||||||||||||||||
| 20 | CA171090 | UGSuM45a | (AAAAG)6 | CDS | ATCTCCTCTTATTCGTTCTGG | AGCAGCGTCTTATCTGGG | PAP fibrillin | 56 | 3 | 5 | 91-141 | Step-wise | 1 | P | P | 0.79 |
|
|
|
|||||||||||||||
| UGSuM45b | 2 | 267-287 | Step-wise | |||||||||||||
|
|
|
|||||||||||||||
| UGSuM45c | 2 | 368-393 | Step-wise | |||||||||||||
|
|
||||||||||||||||
| 21 | CA196477 | UGSuM46a | (GAC)10 | CDS | ACTCCTCCCGCCTCCACTAC | CTCACCGAAGCAATCAAG | Hypothetical protein | 60 | 2 | 15 | 101-188 | Mixed | 1 | P | P | 0.81 |
|
|
|
|||||||||||||||
| UGSuM46b | 3 | 341-377 | Step-wise | |||||||||||||
|
|
||||||||||||||||
| 22 | CA228772 | UGSuM47a | (GCC)10 | CDS | ATTTATGGAGGAAGAAACGG | ATTACAAACAAGAAGAGCGG | Transport protein particle component | 55 | 2 | 3 | 82-97 | Step-wise | 1 | P | P | 0.80 |
|
|
|
|||||||||||||||
| UGSuM47b | 14 | 215-270 | Mixed | |||||||||||||
|
|
||||||||||||||||
| 23 | CA161416 | UGSuM50 | (TC)14 | CDS | CTACTGCCGAGGAAAGATCG | GGAAAAGTTTGTGGCAAGGA | Hypothetical protein | 58 | 1 | 16 | 94-188 | Step-wise | 0 | P | P | 0.80 |
|
|
||||||||||||||||
| 24 | CA228375 | UGSuM73a | (CGC)8 | 5'UTRs | CTTTCAACCTCTACACCTCCAC | ACTAGAAGACTGAGAAGAACCAGT | 40 S ribosomal protein S11 | 55 | 2 | 13 | 101-186 | Mixed | 1 | P | P | 0.82 |
|
|
|
|||||||||||||||
| UGSuM73b | 5 | 354-375 | Step-wise | |||||||||||||
|
|
||||||||||||||||
| 25 | CA261182 | UGSuM74 | (ACA)8 | CDS | TCAGCAGCTGTGAAGTTTCATT | CGTCTCTTTTGGGTTTCATCTC | Transcription regulator protein | 55 | 1 | 18 | 190-271 | Step-wise | 0 | P | P | 0.81 |
|
|
||||||||||||||||
| 26 | CA219230 | UGSuM75a | (TA)12 | CDS | TTGTGCTGATGTTTCCTGCT | CAAGAGAAGATGCCATTAGCC | Patatin-like protein | 55 | 2 | 13 | 94-176 | Step-wise | 1 | P | P | 0.83 |
|
|
|
|||||||||||||||
| UGSuM75b | 4 | 333-351 | Step-wise | |||||||||||||
|
|
||||||||||||||||
| 27 | CA093071 | UGSuM96a | (AT)11 | CDS | TCAAACCAGGATCTAAGCTCAC | GGTAGTGCCATTGAGGTTGC | Putative apyrase | 57 | 2 | 11 | 205-290 | Mixed | 1 | P | P | 0.81 |
|
|
|
|||||||||||||||
| UGSuM96b | 5 | 408-458 | Step-wise | |||||||||||||
|
|
||||||||||||||||
| 28 | CA229840 | UGSuM97a | (GA)11 | 5'UTRs | GCGAGAGAGATAGAGGGAGAGA | AGGTGCCGTTCATGAGGTAGT | Glycine decarboxylase | 56 | 2 | 19 | 240-334 | Step-wise | 1 | P | P | 0.85 |
|
|
|
|||||||||||||||
| UGSuM97b | 2 | 453-469 | Step-wise | |||||||||||||
|
|
||||||||||||||||
| 29 | CA112979 | UGSuM149a | (AGC)7 | CDS | GTTCAATCAAATCCCTCTCCTC | AGCTTGGTCAGCTCCTCATCGTT | Ubiquitin-specific protease 4 (UBP4) | 60 | 2 | 4 | 104-141 | Mixed | 0 | P | P | 0.78 |
|
|
|
|||||||||||||||
| UGSuM149b | 3 | 276-297 | Step-wise | |||||||||||||
|
|
||||||||||||||||
| 30 | CA116368 | UGSuM150a | (GGC)7 | CDS | ACACTGACCGATGGATCCTCTT | ATCAACGTGGACCAGATCTTCTT | hAT family dimerisation domain protein | 60 | 2 | 12 | 90-167 | Mixed | 1 | P | P | 0.76 |
|
|
|
|||||||||||||||
| UGSuM150b | 2 | 279-291 | Step-wise | |||||||||||||
|
|
||||||||||||||||
| 31 | CA280782 | UGSuM177a | (TGC)7 | CDS | GGTGCTGTCCCTATCACTAC | GCCCTTGTTTCTTTGTCTACT | Hypothetical protein | 55 | 2 | 3 | 150-167 | Mixed | 0 | P | P | 0.79 |
|
|
|
|||||||||||||||
| UGSuM177b | 3 | 287-314 | Step-wise | |||||||||||||
|
|
||||||||||||||||
| 32 | CA291445 | UGSuM178 | (GAC)7 | 5'UTRs | GGACTACTACGACTACTGCGA | ACCTTGCTTACATCTTCCTCT | O-diphenol-O-methyl transferase | 54 | 1 | 15 | 100-196 | Step-wise | 0 | P | P | 0.80 |
|
|
||||||||||||||||
| 33 | CA244023 | UGSuM186 | (AG)10 | CDS | AACATTTCGGCATTTGAAGC | GGTCTTTCTTGGGGATCTCTC | Ubiquitin C-terminal hydrolase | 56 | 1 | 4 | 190-232 | Step-wise | 0 | P | M | 0.73 |
|
|
||||||||||||||||
| 34 | CA231668 | UGSuM187a | (CT)10 | CDS | CAACAATTGTCGAAGCCTCTC | TTTGCTTACCCCCTGTTGAC | ATP synthase | 58 | 2 | 10 | 200-264 | Step-wise | 0 | P | M | 0.83 |
|
|
|
|||||||||||||||
| UGSuM187b | 6 | 376-418 | Step-wise | |||||||||||||
|
|
||||||||||||||||
| 35 | AY596560 | UGSuM188 | (GA)10 | 3'UTRs | CCCAAGCGAGCTAGAGAGAG | TCTTCTTTCCTTCGCACAGC | Hypothetical protein | 55 | 1 | 19 | 95-186 | Mixed | 0 | P | P | 0.81 |
|
|
||||||||||||||||
| 36 | CA241232 | UGSuM189a | (CT)10 | CDS | CCGCGACTCTCCTCTCTCT | GTTCTTCTCGGCGTTCCTC | Auxin-regulated protein | 55 | 2 | 15 | 106-200 | Step-wise | 1 | P | P | 0.80 |
|
|
|
|||||||||||||||
| UGSuM189b | 3 | 345-377 | Step-wise | |||||||||||||
|
|
||||||||||||||||
| 37 | CA133642 | UGSuM196a | (AAAG)5 | 5'UTRs | GCTACTATGGACAACAGGG | ATGAAGAGACGAGACGAAGA | Cinnamoyl CoA reductase | 54 | 2 | 18 | 90-181 | Mixed | 1 | P | P | 0.81 |
|
|
|
|||||||||||||||
| UGSuM196b | 2 | 373-393 | Step-wise | |||||||||||||
|
|
||||||||||||||||
| 38 | CA134472 | UGSuM197 | (GA)10 | 3'UTRs | GAAGGAGCAGCAGCGCCAGT | GATTTGCCGTCCTAGGGTTT | Epsin N-terminal homology domain | 56 | 1 | 22 | 200-290 | Step-wise | 0 | P | P | 0.84 |
|
|
||||||||||||||||
| 39 | CA297648 | UGSuM343a | (CAG)6 | 5'UTRs | ACTCCTCCTCCTCGCCGT | TCTTGTTGTAGTAGCCCTTGT | SOUL heme-binding family protein | 62 | 2 | 17 | 250-320 | Mixed | 1 | P | M | 0.87 |
|
|
|
|||||||||||||||
| UGSuM343b | 3 | 438-450 | Step-wise | |||||||||||||
|
|
||||||||||||||||
| 40 | CA300679 | UGSuM344 | (CTC)6 | CDS | CTATCCTCTTGTTGGGTCCT | TCCGCACCTCCGTTCACC | Nucleoside diphosphate kinase protein | 55 | 1 | 1 | 260 | Step-wise | 0 | M | M | 0.0 |
|
|
||||||||||||||||
| 41 | CA084691 | UGSuM345 | (TC)8 | CDS | TATACAAGAATGAAAGGTGAGAGA | AAGCATACTCCCTCTATCTCTATG | DC1 domain-containing protein | 55 | 1 | 2 | 210-220 | Step-wise | 0 | M | P | 0.70 |
|
|
||||||||||||||||
| 42 | CA093455 | UGSuM346 | (AG)8 | CDS | TATACGTAGTAGTGATGATGACCG | CTCCTTCGTCCAGTACCAGTAG | DNA-binding protein DF1 | 60 | 1 | 4 | 150-180 | Step-wise | 1 | P | M | 0.71 |
|
|
||||||||||||||||
| 43 | CA110745 | UGSuM347a | (CT)8 | 5'UTRs | TCTGGCTTTATCGTAACTTGTAT | GAGCCTCGTTTGGGTGGCTTTC | Expressed protein | 55 | 2 | 5 | 230-254 | Step-wise | 1 | P | M | 0.74 |
|
|
|
|||||||||||||||
| UGSuM347b | 1 | 365 | - | |||||||||||||
|
|
||||||||||||||||
| 44 | CA112979 | UGSuM348 | (CG)8 | CDS | CTACCTCCTCGTCTCCTCCCTCTT | AACAAGGAATATGGTCCCTGAG | Unknown protein | 61 | 1 | 2 | 240-261 | Mixed | 0 | M | M | 0.0 |
|
|
||||||||||||||||
| 45 | CA116458 | UGSuM349a | (TC)8 | CDS | CAAGATGTACCCGGACATGGCT | TGCTATACTAGCTATCTCCTTCCT | Unknown protein | 55 | 2 | 3 | 220-236 | Step-wise | 0 | P | M | 0.73 |
|
|
|
|||||||||||||||
| UGSuM349b | 2 | 385-395 | Step-wise | |||||||||||||
|
|
||||||||||||||||
| 46 | CA123971 | UGSuM464 | (ACA)5 | 5'UTRs | GGCTACTTCAGACACGCA | TCTACGCATCAACCTCTCA | SNF2-domain-containing protein | 55 | 1 | 1 | 220-232 | Step-wise | 0 | M | M | 0.0 |
|
|
||||||||||||||||
| 47 | CA125310 | UGSuM465 | (GCA)5 | CDS | GCTAACCAACATCAGCAGT | AGGAGATTGACGAAGAAGAAG | Transducin family protein | 53 | 1 | 2 | 260-279 | Mixed | 0 | M | M | 0.0 |
AUGSuM stands for
BPolymorphic (P) and Monomorphic (M)
Due to high polyploidy and heterozygosity of the
Comparison of fragment size (bp) of variant alleles amplified at 75 polymorphic UGMS loci with changes in number of repeats among
In order to assess the functional significance of the UGMS, gene ontology classification of 767 unigenes carrying perfect and compound microsatellites was carried out. Six hundred seventy-two (87.6%) of these (see Additional files
One of these primers (UGSuM26) showing polymorphism targeted the amylase catalytic domain of β-amylase (Figure
The pair-wise similarity among 36 genotypes belonging to
The genetic relationship among the
All the tropical and sub-tropical varieties were included in a distinct and separate cluster (II) with seven distinct sub-clusters (Figure
Unigene resources representing the transcriptome of an organism provide opportunities to understand the sequence organization in the genic regions of complex polylploids and allow design of sequence based robust markers for various genotyping applications. Considering the availability of a large unigene database of sugarcane in public domain, this resource was studied for the presence and functional relevance of different microsatellite repeats. The frequency of microsatellites in sugarcane unigenes (1/10.9 kb) was lower than that obtained in rice,
The observed frequency of the mononucleotides (12%) in sugarcane was much less than that reported earlier [
The unigenes being longer in higher quality sequences offer advantages over the EST sequence resources for the development of microsatellite markers. In the present study, 961 (810 perfect and 151 compound) primer-pairs were designed from the 767 different unigene sequences carrying microsatellite repeats in the expressed component of the sugarcane genome. The primers designed from the unigene sequences flanking the microsatellite motifs were highly efficient with amplification success rate of 94.9% that suggested the utility of the unigene database in designing sequence based robust genic markers. The paucity of usable and robust sequence based markers in sugarcane has been a major limitation in genetic analysis in this important sugar crop. The UGMS markers developed by us have been placed in the public domain and thus would be immediately useful in various genotyping applications in sugarcane.
For most (87.6%) of the unigene sequences from which the primers were designed, putative functions have been predicted. For instance, about 39% and 4% of the primers were from sequences related to sugar metabolic enzymes and disease resistance, respectively. Interestingly, 54% of the microsatellite repeats were present in various functional domains of proteins encoded by the unigenes. Correlation between fragment length polymorphism due to variable number of UGMS repeat-units in the functional domains and alteration of the predicted protein structure and active ligand binding sites suggested functional relevance of the genic microsatellites. The evolutionary and adaptive advantages of such variable microsatellite repeats affecting the structure and function of the encoded proteins to generate favorable alleles for relaxation of environmental stress impact under the action of high natural selection pressure through modulation of mutation/recombination in these loci have been reported in many eukaryotes [
The efficiency of sugarcane UGMS markers to detect polymorphism within
The level of inter-varietal polymorphism (86%, mean PIC of 0.74) detected by the fluorescent dye labeled primers was higher than the level reported previously with the labeled sugarcane EST derived (38%, PIC of 0.23, [
Unigene sequences usually have advantages of unique identity and position in the transcribed regions of the genome. If this is the case then primers designed from the unigene sequences flanking the microsatellite repeats should amplify unique single locus. In contrast, in the present study, multiple loci and thus amplification of multiple sequences of the same gene was observed for 28 (65.1%) of the 43 primers designed from different microsatellite carrying unigenes. This is possibly due to poor representation of unigene sequences in the database that was scanned for microsatellite repeats. Alternatively, all the copies of a microsatellite carrying gene that are PCR amplified might not be transcribed due to dosage compensation leading to silencing of all but one copy in the large polyploid sugarcane genome [
Distribution pattern of size variant alleles amplified at UGMS loci showed higher proportionate (68%) distribution of alleles showing step-wise mutation than that of mixed allele distribution (32%). High-quality sequence alignment of the size variant alleles showing both step-wise and mixed distributions confirmed the presence of variable number of repeat-units in different amplified alleles and additional insertions/deletions in the flanking regions of microsatellite repeat-motifs, which contributed to the UGMS fragment length polymorphism observed in the
It is important to evaluate molecular diversity existing among the members of the
Evaluation of molecular diversity in a set of 28 commercial Indian tropical and sub-tropical sugarcane varieties using unigene based genic microsatellite markers revealed a wider range (0.33 to 0.84 with an average of 0.40) of genetic similarity than the level detected previously with RAPD (0.59 to 0.81 with average of 0.71, [
The present study identified microsatellites in sugarcane unigenes and assessed their functional relevance. A total of 961 primer-pairs were designed targeting 767 different unigenes carrying microsatellite repeats in the expressed component of the sugarcane genome, which would extend the accessibility of such microsatellite markers to researchers for many genetic studies in sugarcane. Precise allele sizing in automated fragment analysis system has encouraging implications to various high-throughput genotyping applications in sugarcane. Assessment of functional genetic diversity revealed that the genetic base of the Indian sugarcane varieties is not narrow.
Fifteen thousand five hundred ninty-four sugarcane (
To assess the amplification success rate of microsatellite markers designed from sugarcane unigenes, 176 primers were used to amplify one genotype of
List of genotypes belonging to five cereal species and
| Sl. No. | Genus and species | Clones/varieties | Parentage/origin | Region of adaptation |
|---|---|---|---|---|
| 1 |
|
AK559 | Unknown | - |
|
|
||||
| 2 |
|
Kalyansona | PJ"S" × GB 55 | - |
|
|
||||
| 3 |
|
IR64 | IR5857-33-2-1 × IR-2061-465-1-5-5 | - |
|
|
||||
| 4 |
|
KA509 | India | - |
|
|
||||
| 5 |
|
Pusa chari6 | India | - |
|
|
||||
| 6 |
|
IJ-76-3-1-9 | Indonesia | - |
|
|
||||
| 7 |
|
Mamjasahe | India | - |
|
|
||||
| 8 |
|
Malari | Unknown | - |
|
|
||||
| 9 |
|
IM-76-256 | Indonesia | - |
|
|
||||
| 10 |
|
1151 | India | - |
|
|
||||
| 11 |
|
Unknown | - | |
|
|
||||
| 12 |
|
S1135 | Unknown | - |
|
|
||||
| 13 |
|
IK76-81 | Indonesia | - |
|
|
||||
| 14 | Co 1158 | Co 421 × Co 419 | Sub-tropical | |
|
|
||||
| 15 | CoJ 64 | Co 976 × Co 617 | Sub-tropical | |
|
|
||||
| 16 | CoS 88230 | Co 1148 × Co 775 | Sub-tropical | |
|
|
||||
| 17 | Bo 91 | Bo 55 × Bo 43 | Sub-tropical | |
|
|
||||
| 18 | CoS 8436 | MS 68/47 × Co 1148 | Sub-tropical | |
|
|
||||
| 19 | Co 1148 | P 4383 × Co 301 | Sub-tropical | |
|
|
||||
| 20 | CoPant 84212 | Co 1148 × Co 775 | Sub-tropical | |
|
|
||||
| 21 | Co 87268 | Bo 91 × Co 62399 | Sub-tropical | |
|
|
||||
| 22 | CoLk 8102 | Co 1158 GC | Sub-tropical | |
|
|
||||
| 23 | Co 7717 | Co 419 × Co 775 | Sub-tropical | |
|
|
||||
| 24 | Co 87263 | Co 312 × Co 6806 | Sub-tropical | |
|
|
||||
| 25 | Co 89003 | Co 7314 × Co 775 | Sub-tropical | |
|
|
||||
| 26 | CoPant 84211 | Co 6806 × Co 6912 | Sub-tropical | |
|
|
||||
| 27 | Co 8347 | Co 419 × CoC 671 | Sub-tropical | |
|
|
||||
| 28 | Co 7219 | Co 449 × Co 658 | Tropical | |
|
|
||||
| 29 | Co 740 | P 3247 × P 4775 | Tropical | |
|
|
||||
| 30 | Co 86010 | Co 740 × Co 7409 | Tropical | |
|
|
||||
| 31 | Co 86032 | Co 62198 × CoC 671 | Tropical | |
|
|
||||
| 32 | Co 85002 | Co 62198 × (-) | Tropical | |
|
|
||||
| 33 | Co 8021 | Co 740 × Co 6806 | Tropical | |
|
|
||||
| 34 | Co 8371 | Co 740 × Co 6806 | Tropical | |
|
|
||||
| 35 | Co 6304 | Co 419 × Co 453 | Tropical | |
|
|
||||
| 36 | Co 419 | PoJ 2878 × Co 290 | Tropical | |
|
|
||||
| 37 | Co 62175 | Co 951 × Co 419 | Tropical | |
|
|
||||
| 38 | Co 7704 | Co 740 × Co 6806 | Tropical | |
|
|
||||
| 39 | Co 86249 | CoJ 64 × CoA 7601 | Tropical | |
|
|
||||
| 40 | CoC 671 | Q 63 × Co 775 | Tropical | |
|
|
||||
| 41 | Co 87025 | Co 7704 × Co 62198 | Tropical | |
Variation in the fragment size (bp) of amplified alleles at each polymorphic UGMS marker locus was compared with changes in the number of microsatellite repeat-units at that target locus, and the occurrence of "stepwise" and "mixed" type of allele size distribution was inferred. When the allele size differences strictly corresponded to the variation in the number of repeat-units, it was considered as stepwise distribution. A mixed allele distribution was assumed when the allele size differences could partly be explained by the stepwise model. Multiple amplicons obtained by a primer-pair with peaks ≥ 1500 fluorescence units showing ≥ 100 bp allele size differences were binned into different loci. To confirm that the primers amplified the target microsatellite repeat motifs in different species and genera, the amplified products were purified using Micropon PCR purification kits (Millipore, Bedford, MA, USA) and sequenced two times in both forward and reverse directions using a capillary-based Automated DNA Sequencer (MegaBACE 1000, Amersham Biosciences, Piscataway NJ, USA). The trace files were base called, checked for quality and assembled into contigs [
The polymorphic information content (PIC) was calculated using the formula, PIC = 1- ∑Pij2 [
Sugarcane, Unigenes, Microsatellites, Functional genetic diversity
SKP conducted mining of UGMS, marker design, large-scale genotyping, polymorphism survey, functional diversity estimation and drafted the manuscript. AP was involved in genotyping and sequencing. KG and TRS participated in microsatellite mining and data analysis. PSS and NKS helped in data analyses, interpretation and drafting of the manuscript. TM designed the study, guided data analysis and interpretation, participated in drafting and correcting the manuscript and gave the final approval of the version to be published. All authors have read and approved the final manuscript.
Click here for file
Click here for file
Click here for file
Click here for file
Click here for file
Click here for file
Click here for file
Click here for file
Click here for file
Click here for file
The work presented in the manuscript was funded by the Department of Biotechnology (DBT), Government of India. We are thankful to the NCBI for making available their databases, Institute of Plant Genetics and Crop Research (IPK) for the availability of microsatellite search tool MISA and Dr. Athiappan Selvi, Sugarcane Breeding Institute, Coimbatore for providing the sugarcane genotypes.