A novel
While the genome sequences for a variety of organisms are now available, the precise number of the genes encoded is still a matter of debate. For the human genome several stringent annotation approaches have resulted in the same number of potential genes, but a careful comparison revealed only limited overlap. This indicates that only the combination of different computational prediction methods and experimental evaluation of such
Here we provide evidence for the transcription of approximately 2,600 additional genes predicted by Fgenesh. Validation of the developmental profiling data by RT-PCR and
The successful design and application of this novel
Knowledge of the complete gene set of a genome is a prerequisite to an integrated view of the network of encoded functions. One major obstacle to this goal is the reliable identification of genes within the vast excess of non-coding sequences within eukaryotic genomes. While bioinformatics methods have substantially evolved in their prediction capabilities, they still compromise on sensitivity versus specificity. Most genome annotations performed so far have concentrated on maximizing the number of 'real' genes by requiring multiple evidence before accepting an annotation, such as: the presence of expressed sequence tags (ESTs); homologies to known proteins or the conservation of genomic sequences between related organisms; or by raising the thresholds of the prediction software in order to keep the number of false positives to a minimum. Another problem of such mere
For
Most whole-transcriptome microarrays will therefore remain incomplete, as the only safe way to include all potential genes (and even their splice forms) is the construction of arrays based on a whole genome tiling path. Whilst the feasibility and superiority of such an approach has been shown for parts of the human chromosome 22 [
We decided to combine the published, conservative BDGP genome annotation Release 2 with an alternative, less stringent gene prediction for the design of a new PCR-fragment based whole-transcriptome microarray for
To overcome the known limitations in gene prediction, we constructed our
The simplest explanation for the high number of HDC unique predictions may be the relaxed stringency criterion applied. Consequently, a careful inspection of the two sets (BDGP/FlyBase versus HDC) showed a high degree of similarity for most common predictions; differences were largely confined to the 5' and 3' ends of the predictions as may be expected. This is not only because
The assumption that most of the HDC unique predictions have been omitted from Release 3.1 due to the lack of significant homologies is further substantiated by homology searches performed against the SwissProt database, which were positive for only 1.5% of the HDC unique predictions. Likewise, only 1.8% scored a hit in an InterPro protein domain search and only 10.2% showed an overlap with EST sequences [
In summary, we find that most of the HDC unique predictions cannot be confirmed by
Using GenomePride [
Together with a number of controls the complete set was spotted in duplicates, resulting in a high density (47,616 spots) transcriptome microarray. Initial hybridizations already indicated on visual inspection that a high percentage of the novel genes are expressed (Figure
After data processing and stringent filtering, we found 13,927 genes to be expressed during the
The latest FlyBase Release 3.1 not only resulted in structural changes to 85% of the transcripts and 45% of the predicted proteins [
Despite the fact that the developmental expression profiling was thoroughly filtered and statistically analyzed, we decided to define the lower limit for the number of novel genes by an additional level of validation. We therefore performed RT-PCR analyses for a semi-random selection (with respect to chromosomal order, see Methods) of newly predicted genes as well as of some previously known ones. For the latter, we confirmed the microarray results at 93.8% (136/145), and confirmed 74.4% (218/293) of the uniquely predicted novel genes. While we cannot exclude the fact that the RT-PCR erroneously missed some weakly expressed genes in the RNA pool made from all stages analyzed in the microarray experiments, this result clearly points to the existence of at least 2,000 additional genes (Table
The additional analysis of gene expression by
One important step during the design of the amplicon set was to exclude regions that showed significant homologies to other regions in the genome. Not only should this step exclude amplicons that would represent conserved protein motifs but it should also prevent most of the transposable elements and repeat structures from being represented in our set. Nevertheless, regions of low stringency identity might have been included in our amplicon set and thus part of the observed expression might result from cross-hybridization of a real gene to, for example, an amplicon representing a non-expressed, degenerated pseudogene. We therefore re-analyzed all amplicons for HDC unique predictions that scored positive in our expression profiling for the existence of additional low stringency blast hits and found no match for 86.5% of them. Careful comparison of the expression profiles for the remaining 13.5% with their second site hits showed that only 15.4% of them were co-regulated, demonstrating that no more than 2% of all novel genes might represent false-positives due to cross-hybridization.
As previously mentioned, the design of HD FlyArray enables further uses of our amplicon set, such as the expression of peptide representatives for each gene for use in antibody production or, as exemplified by the work performed by Boutros
We further analyzed this screen to obtain an estimate of how many of the novel predictions may constitute additional exons for genes already predicted by the FlyBase annotation. To this end, we tested whether FlyBase R3.1 predicted genes neighboring a HDC unique prediction also showed lethality in the RNAi screen. Including 25 kb upstream or downstream, we found such a FlyBase gene for 8/68 (11.8%) of the HDC unique predictions. For 1/68, we saw another HDC unique prediction. Expanding the search space to 50 kb, about 14.7% of the novel genes might be additional exons to a FlyBase gene and 3/68 novel genes might consist of two HDC unique predictions. These numbers are essentially identical to the results obtained for the 369 FlyBase predicted genes influencing cell viability in the RNAi screen. For these, we found 13% (48/369) within 25 kb and 16% (59/369) within 50 kb that could be interpreted as additional exons. We conclude that the vast majority of the HDC unique predictions are not additional exons of genes predicted by FlyBase.
In contrast to the extensive corrections and additions that significantly improved the genome annotation Release 3.1, we did not rework the original set of gene models predicted by Fgenesh as in some cases the arbitrary separation of exons into different genes may have been advantageous in predicting genes that lie within the intron of another gene. Nevertheless, as our expression data not only validates the existence of transcriptional units but also excludes cross-hybridization and the assignment of exons into separate genes as the primary source for most of our expressed novel predictions, we believe it is appropriate to call such transcriptional active regions genes. Ultimately, the correct gene structure for all
Until such experimental confirmation for all gene models exists, the comparison of different predictions may offer a good starting point for judging the reliability of the computed gene structures. We therefore used the Generic Genome Browser (GBrowse) [
Our integrated
In addition, the fact that most of the newly identified genes show no significant homology to known proteins (comparison to SwissProt) or domains/motifs (InterPro search) and also lack considerable conservation between species (
In addition to its value in microarray production, the Heidelberg amplicon set also proved to be a valuable tool for genome-wide RNAi studies. The possibility of using the same set of fragments for both expression profiling and genome-wide RNAi experiments will be of great benefit for further studies on genetic networks.
The complete Heidelberg Prediction is available via download as a FASTA formatted file. Each entry consists of the CDS position information, the strand orientation as well as the sequence of genomic region spanning the prediction. For details on the Fgenesh software and the parameters used refer to [
We extended the BDGP set by genes of the Drosophila Gene Collection (DGC), which were not represented in BDGP, based on their gene names. Sets of predicted genes (for example, DGC/BDGP and Fgenesh) were combined by comparing the overlap in exon sequences of the appropriate orientation. Two gene predictions were defined as reflecting the same gene if the number of common exonic base pairs exceeded 30% of the length of the shorter prediction. If two gene predictions out of one set were covered by a single prediction of the second set we took the shorter predictions as representatives. Finally, we manually added 71 genes that were not included in the genomic sequence.
Since the amplicons should be unique to the gene they represent, the GenomePRIDE software [
After defining the optimal region within a gene, GenomePRIDE computes both PCR primers independently. The optimal position of a primer is hereby defined by the boundaries of the previously selected target region, aiming to amplify a fragment of optimal length. Similar to the PRIDE software used for sequencing, the design of a single primer using GenomePRIDE is based on the evaluation of the thermodynamic stability, strength of the most stable secondary binding site, formation of primer dimers, and the position of the primer (now with respect to the preselected target). The evaluation of potential secondary binding sites of each primer not only includes all exons and introns of the respective gene, but also includes the sequences flanking the gene. We used the default value of 10 kb for the length of these flanking regions, which is generally sufficient to avoid the design of primers that may give rise to a secondary PCR product. Primers were synthesized by Eurogentec (Seraing, Belgium).
PCR amplifications were performed in 96-well microtiter plates. The fragments for all genes of the Heidelberg Collection R1 were initially PCR-amplified from genomic DNA of
The second round PCR-product was spotted in 3x SSC, 150 mM NaPO4, 1.5 M betaine onto QMT Amino slides (Quantifoil, Jena, Germany) using a MicroGrid II arrayer (BioRobotics, Cambridge, UK) and SMP3 pins (TeleChem International Inc., Sunnyvale, USA). Each PCR-product was spotted twice at different positions on the microarray. As controls, PCR-products of
Embryo samples were collected as four-hour egg lays, which were allowed to develop for the desired interval and then snap frozen in liquid nitrogen. A small aliquot was DAPI-stained and inspected for correct staging. The different larval stages as well as the different pupal stages were handpicked and separately snap frozen. Adults were collected as male and female flies and also snap frozen. Total RNA was isolated from all samples using the Trizol reagent (Invitrogen, Karlsruhe, Germany). The concentration and quality was analyzed by separating the samples in an RNA 6000 Nano Assay on the Agilent Bioanalyzer 2100 (Agilent Technologies, Palo Alto, USA). At least three different RNA preparations were pooled for each developmental stage. For larval stage sample, equal amounts of the original pools made from the three separate larval stages were mixed. Likewise, for adult stages, equal amounts of male and female pools were mixed.
At least three independent experiments were performed for each of the eight different stages (embryonic stage 4-8 h, embryonic stage 8-12 h, embryonic stage 12-16 h, pooled larval stage, pupal stage I, pupal stage II, pupal stage III and adult stage) in competitive hybridizations against the common control, which was a sample made from RNA isolated from embryonic stage 0-4 h. At least one dye swap was included in each of the repetitions. We used the indirect labeling method, with 9 μg random hexamer primer (Invitrogen, Karlsruhe, Germany) added to 20 μg total RNA. First-strand cDNA synthesis was performed using 1x dNTP-Mix (25 mM dATP, dCTP, dGTP, 15 mM dTTP, 10 mM amino-allyl-dUTP (Sigma, Heidelberg, Germany), 400 U Superscript II RT (Invitrogen, Karlsruhe, Germany), 1x first-strand buffer (Invitrogen, Karlsruhe, Germany), 0.1 M DTT and incubated at 42°C for 2 h. After purification with QIAquick columns (Qiagen, Hilden, Germany), the cDNA was eluted in 1 M KPO4, pH 8.5, and dried. Then, Cy3 or Cy5 dye-esters (Amersham Pharmacia Biotech, Freiburg, Germany) in 4.5 μl DMSO (Fluka, Buchs, Switzerland) were added to the cDNA in 0.1 M Na2CO3 and incubated in the dark at room temperature for 2 h. After purification on QIAquick columns, the concentration of the labeled cDNA and the dye incorporation rate were determined. Pre-hybridization of the QMT Amino slides was done in 1% BSA, 5x SSC, 0.1% SDS at 55°C for 45 min in order to block the amino-coated glass surface. For hybridization, the labeled cDNAs were taken up in SlideHyb buffer #1 (Ambion, Woodward, USA), denatured at 95°C for 5 min, and applied to the array. Hybridization was performed in appropriate chambers (TeleChem) in a waterbath at 55°C for 16 h. Washing of the slides after hybridization was performed according to the manufacturer's instructions.
The hybridized glass slides were scanned on a ScanArray 5000 (Perkin Elmer, Wellesley, USA) using the ScanArray software (Version 3.1). Directly fluorescence-labeled external controls spotted on the array were used as a reference for the alignment of laser and photomultiplier (PMT) settings. The resulting images (TIFF format, 16-bit grayscale) for each channel-Cy5 (632 nm) and Cy3 (532 nm)-were analyzed further with the GenePix software (Version 4.0; Axon Instruments, Union City, USA). The images were combined and quantified giving rise to the results files (gpr-format). These files contain information about the gene names and/or clone IDs, linked this information to the respective microarray features and the relevant signal intensities.
A first quality filter was applied directly after scanning the array. As each microarray contains two replicates for each gene, we used the standard deviation (SD) of their dye-ratio (Cy5/Cy3) to filter for reproducible hybridizations. Only if at least 30% of all genes on the array had a SD of log(632/532) of less than one third between the replicates, was the microarray included for further analysis. The raw data of these experiments were analyzed further using the M-CHiPS software package [
Correspondence analysis (CA) is an exploratory projection method which displays the associations between genes and hybridizations in a plot diagram. It is well suited for analyzing large data sets - representing transcription intensities together with the corresponding hybridizations in one high dimensional space [
For an independent validation of the expression profiling data we re-used our original primer set and selected several complete 96-well microtiter plates enriched in primer pairs specific for Heidelberg unique predictions. Although this selection is only semi-random, as the resulting amplicons are ordered along the chromosome, no specific bias may be expected. RNA was isolated as described above for Microarray Hybridization. For RT-PCR, RNA from all stages was pooled. Contaminating genomic DNA was removed by digestion of 20 μg pooled RNA in a 50 μl reaction containing 1x NEB restriction buffer 2, 50U RNasin (Promega, Mannheim, Germany) and 25U RNase-free DNase I (Roche Diagnostics, Mannheim, Germany) for 1 h at 37°C. After phenol-chloroform and chloroform extraction and precipitation with 2.5vol of ethanol, RNA was re-dissolved at 1 μg/μl and used for reverse transcription. One microgram of RNA was incubated with 30 pmol random primers (Roche Diagnostics, Mannheim, Germany) at 65°C for 10 min, chilled on ice, and supplemented with 2 μl of 100 mM DTT, 2 μl of dNTPs (10 mM each, Roche Diagnostics, Mannheim, Germany), 10U of RNasin (Promega, Mannheim, Germany), 4 μl of 5x Expand RT-buffer (Roche Diagnostics, Mannheim, Germany), and 50U of Expand reverse transcriptase (Roche Diagnostics, Mannheim, Germany) to a final reaction volume of 20 μl. After 10 min at 30°C, the reaction was incubated for 1 h at 42°C followed by 10 min at 55°C. PCR was performed in a 50 μl reaction inoculated with 1 μl of the RT reaction containing 1x QIAGEN PCR buffer (1.5 mM MgCl2, Qiagen, Hilden, Germany), 0.4 mM each dNTP, 2U QIAGEN Taq polymerase (Qiagen, Hilden, Germany) and 100 pmol of each primer. The plates were incubated for 5 min at 94°C, followed by 35 cycles of denaturation at 94°C for 30 s, annealing for 30 s and elongation at 72°C for 90 s. The annealing temperature was lowered during the first ten cycles from 65°C to 55°C to increase specificity of the amplification. In the last 20 cycles the elongation time was prolonged by 5 s in each cycle to compensate for decreasing enzymatic activity. For all RT-PCRs, a control lacking reverse transcriptase addition (RT-) to detect contamination with genomic DNA was included. PCR amplification success was checked on 1% agarose gels and only scored if the RT-control was negative.
A subset of genes was verified by
All genome sequence related data are presented in the Generic Genome Browser web interface (GBrowse) [
The microarray data described in this article have been submitted to the GEO data library under the accession numbers GPL517, GSM10917-GSM10940. The validated predictions described in this paper have been submitted to the TPA data library at GenBank under the accession numbers BK001800-BK003945.
All primary data are available, including the original Heidelberg Prediction (Additional data file
The original Heidelberg Prediction
Click here for additional data file
The Heidelberg Primer set
Click here for additional data file
Summary of expression profiling results and additional experiments based on the Heidelberg Collection R1 and R2
Click here for additional data file
Correspondence analysis results
Click here for additional data file
We would like to thank A. Kiger, S. Armknecht, K. Kerr and N. Perrimon for sharing information before publication and for a stimulating and open collaboration over the course of this project. We are grateful to M. Leptin for help in analyzing the
The Heidelberg Collection R1. The combination of the BDGP cDNA Collection (BDGC) R1 with the BDGP genome annotation Release 2 contained 13,861 genes. The Heidelberg Prediction based on the Fgenesh
Developmental profiling.
Genomic location and expression patterns of Heidelberg unique predictions. The left part of the figure visualizes the genomic region (10 kb of sequence) for some examples of the novel Heidelberg Predictions. In addition, here is the corresponding amplicon present on the microarray as well as information on conserved regions (
The Heidelberg FlyArray website. Screen shot of the Heidelberg FlyArray website based on the GBrowse platform. After selecting the genomic region of interest, for example by gene name, amplicon name or position, the user is offered a comparative view of the different gene models from the BDGP genome annotations Release 2, FlyBase Release 3.1 and the Heidelberg Prediction, as well as the placement of the amplicons chosen for the Heidelberg FlyArray. In addition, researchers find a comparison to
Conservation between
| Common predictions | Heidelberg Predictions | Expressed Heidelberg Predictions | |
| 10% overlap | 59.8% | 54.2% | 54.9% |
| 30% overlap | 53.5% | 31.3% | 32.5% |
| 50% overlap | 44.9% | 13.4% | 13.7% |
We tested the commonly predicted genes, the Heidelberg unique predictions and the Heidelberg unique predictions that are expressed according to our expression profiling for conservation to
Summary for the Heidelberg Collection R1
| Total | Heidelberg Predictions | BDGP Release 2/ BDGC R1 | Other | Common predictions | |
| Heidelberg Collection R1 | 21,396 | 7,464 | 483 | 71 | 13,378 |
| Heidelberg PrimerSet | 21,306 | 7,463 | 442 | 65 | 13,336 |
| Heidelberg FlyArray | 20,948 | 7,319 | 425 | 62 | 13,142 |
| Expressed during development | 13,927 (66.5%) | 3,497 (47.8%) | 232 (54.6%) | 52 (83.9%) | 10,146 (77.2%) |
| Validation by RT-PCR | 386/478 (80.8%) | 334/424 (78.8%) | ND | ND | 52/54 (96.3%) |
The Heidelberg Collection R1 resulted from the combination of our novel annotation with the published BDGP Release 2 annotation and the sequences of the BDGC R1 clones. The PrimerSet includes only those annotations for which we could successfully design primer pairs, and likewise, the Heidelberg FlyArray sums up the annotations that are included on our novel microarray. The next row presents the results of the developmental profiling; numbers given in parentheses are the percentage of annotations represented on the array that scored positive. The last row shows the validation rate of the microarray results by RT-PCR (amplicon length ≤750 bp).
Summary for the Heidelberg Collection R2
| Total | Heidelberg Predictions | FlyBase Release 3.1 | Other | Common predictions | |
| Heidelberg Collection R2 | 19,879 | 6,224 | 605 | nd | 13,050 |
| Heidelberg PrimerSet | 19,095 (19,389) | 6,224 (294) | 296 | 40 | 12,535 |
| Heidelberg FlyArray | 18,837 (19,123) | 6,143 (286) | 288 | 39 | 12,367 |
| Expressed during development | 12,574 (66.8%) (12,734) (66.6%) | 2,636 (42.9%) (160) (55.9%) | 167 (57.9%) | 30 (76.9%) | 9,741 (78.8%) |
| Validation by RT-PCR | 354/438 (80.8%) | 218/293 (74.4%) | ND | ND | 136/145 (93.8%) |
The Heidelberg Collection R2 resulted from the combination of our novel annotation with the recently published FlyBase Release 3.1 annotation (excluding non-CG annotations, such as TE and CR). Only Heidelberg Predictions, primers and amplicons that matched with high stringency to the FlyBase genomic sequence Release 3.1 were included and re-assigned to the new Heidelberg Collection R2, thus all numbers represent a lower limit. Moreover, numbers in the table are corrected for several amplicons matching a single gene. As before, the PrimerSet includes all annotations for which we successfully designed primer pairs, and likewise, the Heidelberg FlyArray sums up the annotations that are included on the microarray. The next row presents the results of the developmental profiling; numbers given in parentheses are the percentage of annotations represented on the array that scored positive. The last row shows the validation rate of the microarray results by RT-PCR (amplicon length ≤750 bp). In the column for the Heidelberg Predictions we included below (in parentheses) the number of annotations that were unique to BDGP Release 2 and are not part of the FlyBase Release 3.1 CG annotations.
Expression patterns obtained by
| Name | GenBank accession number | Pattern | Evidence | Chromosome | Comments |
| HDC00027 | BK003260 | Ectoderm | ag | 2L | - |
| HDC00627 | BK003299 | Ectoderm | - | 2L | - |
| HDC00658 | BK003302 | Cellular blastoderm subset, salivary glands | - | 2L | - |
| HDC00966 | BK003326 | Cellular blastoderm subset | - | 2L | Intron CG11030 |
| HDC00979 | BK003327 | Yolk nuclei | - | 2L | - |
| HDC02005 | BK003369 | Maternal, subset of cells, embryonic large intestine | dp | 2L | - |
| HDC02009 | BK003370 | Protocerebrum primordium, trunk mesoderm primordium | dp | 2L | - |
| HDC02141 | BK003388 | Embryonic gut, cells in the head (stage 10/11) | - | 2L | - |
| HDC02262 | BK003403 | Weak signal | dp | 2L | - |
| HDC02272 | BK003405 | Weak signal | dp | 2L | - |
| HDC02494 | BK003424 | Mesoderm anlage | dp | 2L | Intron CG15288 |
| HDC02527 | BK003429 | Salivary glands | dp | 2L | - |
| HDC02528 | BK003430 | Protocerebrum primordium, anterior midgut primordium | dp | 2L | - |
| HDC02634 | BK003455 | Cellular blastoderm subset | dp | 2L | - |
| HDC02764 | BK003493 | Cellular blastoderm, ubiquitous, salivary glands | dp, EST | 2L | Intron CG4838 |
| HDC03057 | BK003539 | Maternal, blastoderm, ubiquitous, gut | dp, EST | 2L | Intron CG5803 (overlap) |
| HDC03960 | BK003614 | Trunk mesoderm anlage, head mesoderm primordium | dp | 2R | Opposite strand to CG17921 (overlap) |
| HDC04256 | BK003630 | Subset of mesoderm, gonads | dp | 2R | - |
| HDC05090 | BK003664 | Subset of cells (procephalic ectoderm primordium?), midgut | - | 2R | - |
| HDC05183 | BK003670 | Ubiquitous | dp | 2R | - |
| HDC05573 | BK003699 | Midgut, central nervous system | dp, EST | 2R | - |
| HDC06000 | BK003754 | Cellular blastoderm excluding ventral structures | dp | 2R | Intron CG12369, same staining as HDC05999 |
| HDC06241 | BK003785 | Ventral ectoderm anlage, trunk mesoderm anlage | dp | 2R | - |
| HDC06636 | BK003845 | Maternal | dp, EST | 2R | - |
| HDC07387 | BK003934 | Maternal, subset of cells until stage 12 | - | 2R | - |
| HDC07791 | BK001850 | Weak, ubiquitous at 4-8 h | - | 3L | - |
| HDC08265 | BK001908 | Subset of cells | - | 3L | - |
| HDC08749 | BK001956 | Weak, ubiquitous at 4-8 h | dp | 3L | - |
| HDC09080 | BK002002 | Salivary glands | dp | 3L | - |
| HDC09253 | BK002020 | Posterior spiracles, ectoderm | dp | 3L | - |
| HDC09513 | BK002067 | Weak, ubiquitous at 4-8 h | dp | 3R | - |
| HDC10019 | BK002122 | Salivary gland primordium, salivary glands | - | 3L | Intron CG10741 |
| HDC10028 | BK002123 | Ventral nerve cord | dp | 3L | Intron CG12478 |
| HDC10120 | BK002139 | Trunk mesoderm anlage, cuprophilic cells | - | 3L | Intron CG 17697 |
| HDC10292 | BK002156 | Lateral stripes blastoderm, third wave of neuroblasts | ag | 3L | Predicted in 2.0 as CG17014 |
| HDC10646 | BK002195 | Pole plasm, trunk mesoderm, salivary glands, embryonic midgut | dp | 3L | - |
| HDC10913 | BK002212 | Anterior midgut primordium, posterior midgut primordium | dp | 3L | Intron CG11614 (opposite strand) |
| HDC11249 | BK002252 | Malpighi, gonads | dp | 3L | Intron CG32432 (opposite strand) |
| HDC11512 | BK002283 | Weak, ubiquitous at 4-8 h | dp | 3L | - |
| HDC11876 | BK002318 | Weak, ubiquitous at 4-8 h | dp, EST | 3R | Intron CG12163 (opposite strand) |
| HDC11908 | BK002321 | Ventral nerve cord, embryonic central nervous system | dp | 3R | - |
| HDC12497 | BK002400 | Weak, ubiquitous at 4-8 h | - | 3R | - |
| HDC12511 | BK002404 | Ectoderm | dp | 3R | - |
| HDC12925 | BK002446 | Weak, ubiquitous at 4-8 h | EST | 3R | - |
| HDC13248 | BK002490 | Weak, ubiquitous at 4-8 h | - | 3R | - |
| HDC13350 | BK002511 | Ectoderm | - | 3R | Intron CG7855 (opposite strand) |
| HDC13470 | BK002532 | Cellular blastoderm subset segmentally repeated, ectoderm, embryonic foregut, embryonic hindgut | dp, EST | 3R | - |
| HDC13644 | BK002551 | Embryonic midgut, anal pads | - | 3R | - |
| HDC13905 | BK002563 | Trunk mesoderm anlage, embryonic midgut | dp | 3R | - |
| HDC14221 | BK002623 | Ventral ectoderm anlage, posterior endoderm anlage | dp | 3R | Intron CG31243 (opposite strand) |
| HDC14231 | BK002626 | Maternal, salivary glands | dp, EST | 3R | Short overlap with TE19396 |
| HDC14493 | BK002672 | Dorsal vessel | ag, dp | 3R | Intron CG31175 |
| HDC15090 | BK002773 | Maternal | dp | 3R | - |
| HDC15681 | BK002831 | Weak, ubiquitous at 4-8 h | dp, EST | 3R | - |
| HDC15728 | BK002837 | Maternal | dp | 3R | - |
| HDC16092 | BK002888 | Weak, ubiquitous at 4-8 h | dp | 3R | - |
| HDC16243 | BK002914 | Anterior endoderm anlage, anterior midgut primordium, posterior midgut primordium | dp | 3R | - |
| HDC16874 | BK002959 | Yolk nuclei, anterior endoderm anlage, embryonic midgut, subset of cells | ag | X | - |
| HDC16879 | BK002961 | Invaginating cells (hemocytes?/oenocytes?) | dp | X | - |
| HDC17351 | BK003012 | Embryonic gut | - | X | - |
| HDC18148 | BK003079 | Weak, ubiquitous at 4-8 h | - | X | - |
| HDC18326 | BK003102 | Weak, ubiquitous at 4-8 h | dp | X | Intron CG1691 |
| HDC18410 | BK003108 | Weak, ubiquitous at 4-8 h | dp | X | - |
| HDC19378 | BK003172 | Weak, ubiquitous at 4-8 h | dp | 3R | - |
| HDC19530 | BK003190 | Weak, ubiquitous at 4-8 h | - | X | - |
| HDC19643 | BK003204 | Midgut primordium, embryonic midgut | ag | X | Intron CG32541 (opposite strand) |
| HDC19645 | BK003205 | Cuprophilic cells | - | X | - |
For 40% of the novel genes tested we detected an expression pattern during embryonic development. Any overlap with regions conserved in