This is an Open Access article distributed under the terms of the Creative Commons Attribution License (
Various statistical scores have been proposed for evaluating the significance of genes that may exhibit differential expression between two or more controlled conditions. However, in many clinical studies to detect clinical marker genes for example, the conditions have not necessarily been controlled well, thus condition labels are sometimes hard to obtain due to physical, financial, and time costs. In such a situation, we can consider an unsupervised case where labels are not available or a semi-supervised case where labels are available for a part of the whole sample set, rather than a well-studied supervised case where all samples have their labels.
We assume a latent variable model for the expression of active genes and apply the optimal discovery procedure (ODP) proposed by Storey (2005) to the model. Our latent variable model allows gene significance scores to be applied to unsupervised and semi-supervised cases. The ODP framework improves detectability by sharing the estimated parameters of null and alternative models of multiple tests over multiple genes. A theoretical consideration leads to two different interpretations of the latent variable, i.e., it only implicitly affects the alternative model through the model parameters, or it is explicitly included in the alternative model, so that the interpretations correspond to two different implementations of ODP. By comparing the two implementations through experiments with simulation data, we have found that sharing the latent variable estimation is effective for increasing the detectability of truly active genes. We also show that the unsupervised and semi-supervised rating of genes, which takes into account the samples without condition labels, can improve detection of active genes in real gene discovery problems.
The experimental results indicate that the ODP framework is effective for hypotheses including latent variables and is further improved by sharing the estimations of hidden variables over multiple tests.
Selecting significantly differential genes is one of the most important tasks in studies of gene expression analysis. This, however, is not an easy task because gene expression measurement studies include some difficulties, such as high noise, a small number of measurements (samples) in each condition, severe multiplicity of tests corresponding to a large number of genes, and so on. Our purpose is to improve the gene significance ranking score, by which a larger proportion of truly active genes is referred to as significant, while a greater proportion of truly inactive genes is called insignificant. Here, a gene is called "active" if its mean expression level has a relationship to the biological conditions of interest, and it is "inactive" if there is no such relationship.
Most simply, the significance of genes can be evaluated using p-values with the Bonferroni correction [
Sharing commonality among multiple tests has become another important issue in recent years, in order to increase the detectability from expressed genes that are highly correlated. The significance analysis of microarray (SAM) [
Very recently, Storey et al. [
In this study, we extend the ODP framework to general problems in which alternative distributions are defined by a latent variable model. Specifically, we propose a mixture model, a typical example of a hidden variable model, which assumes a stochastic noise process in each condition and includes a latent (hidden) variable indicating a class label of the condition to which the sample belongs. Since the label is regarded as a hidden variable, it is not necessarily provided. This model is then lent to two new significance scores of genes; an unsupervised significance score that assumes the class labels are not provided, and a semi-supervised significance score that can deal with samples either with or without class labels. We found that there are generally two theoretically natural but different ways to deal with the hidden variables in the ODP manner; namely, the estimated values of hidden variables are shared among multiple tests explicitly or implicitly through the model parameters. We compare them through a simulation experiment and show that explicit sharing of the commonality in the hidden variable improves the detectability of active genes. Using artificial and real data sets, we also show that the unsupervised and semi-supervised score of gene significance can improve the stability of active gene ranking against a small number of labeled samples by incorporating the unlabeled samples.
Our objective is to accurately determine whether a gene
Following the conventional framework of statistical testing, let the null hypothesis that the gene is actually inactive be denoted as
which was stated and proven as Neyman-Pearson's lemma [
Many useful statistical tests, however, assume non-simple hypothetical models so that the null and alternative PDFs include variable parameters to be estimated statistically. For example, in a typical approach for supervised differential gene discovery [
where
where the maximum likelihood estimates (MLEs) of the parameters are given as:
Accordingly, the supervised score LR-S(
For an active gene, gene expression is assumed to relate to the condition of each sample; in other words, it is "on" in samples in certain conditions and "off" in others. Let a hidden variable
For an active gene
where
The null hypothesis is the same as given in the supervised case, i.e., a normal distribution:
The log-likelihood ratio score is then defined as
where
When we have an MLE
Let
In the semi-supervised case, an active gene can be modeled by modifying the prior of the hidden variable so that the prior incorporates the label's inobservability:
This prior states that the hidden label is the same as the observed label when it is observed, but is predominantly predicted by the prior knowledge
According to Neyman-Pearson's lemma, a statistic is defined as being most powerful when its detection probability
This lemma defines an ideal ODP but not a practical one, because it needs information not available in actual situations of multiple testing. The true values of parameters
where
By applying the typical ODP estimation to supervised differential gene discovery problems for artificial and real data sets, they showed that ODP is better at detecting active genes than the existing methods. This means a larger number of genes were obtained at each fixed false discovery rate (FDR), where the FDR was estimated by a permutation test.
The key of the ODP in contrast to the individual likelihood ratio is that ODP shares among all genes common information about the distributions represented as null and alternative models, so that the common information is used when evaluating a single gene. They actually demonstrated that, when the hypothetical models have some global characters shared by the genes, such as an asymmetric or cluster structure or both, ODP can incorporate them to improve the performance of gene discovery, i.e., to increase the ETP for a fixed EFP.
For parametric hypothetical models, the distribution of hypotheses is attributed to the distribution of model parameters. When there are hidden variables, as in our model's case, other definitions of the hypothesis are possible: (a) unknown values of the hidden variables are marginalized out in the hypotheses, or (b) the hidden variables are explicitly included, as are the model parameters in the hypothetical models. In case (b), the distribution of hypotheses is attributed to the distributions of both model parameters and hidden variables. These two definitions of the hypothesis distribution lead to different results. We formalize these two cases in this section and compare them in the next one.
In the likelihood ratio score, we evaluate the ratio of the two likelihood functions,
In case (a), i.e., when only the model parameters are shared by all genes, the likelihood function becomes
In case (b), i.e., when the estimated posterior of the hidden variable,
As we pointed out at the beginning of this section, the difference between the two cases above comes from the difference in the interpretation of what the model is and what the hypothesis is. In fact, if one assumes that the posterior of the hidden variable
The model in case (b) may have a biological meaning. The hidden variable vector
Strictly speaking, a gene is called active if the true mean expression levels are different in different conditions regardless of whether the difference is large or small. However, in real situations it is impossible in principle to know whether a gene is truly active or inactive because its definition must be based on a finite (and often small) number of samples. Note also that it is difficult to confirm ever from any biological experiments whether any gene is truly active or inactive in the strictest sense – it is only a matter of the degree of significance.
In order to compare gene selection criteria, one solid way to consider an artificial situation where expression levels of truly active and truly inactive genes are given by probabilistic sampling from known probability densities on gene expression levels. On the other hand, we should also conduct experiments on real data sets because true distributions of gene expression levels of active and inactive genes are not necessarily the same as artificial distributions. Since we can expect that when the amount of information corresponding to the number of samples increases, gene rankings based on any statistical significance scores become more stable, we prepared a provisional target based on a sufficiently large number of labeled data sets. Based on the provisional target, we compared some gene significance scores with each other by evaluating stabilities against loss of information corresponding to loss of the number of labeled samples in realistic situations.
First of all, we compare various unsupervised scores devised for the detection of significant genes in an artificial setting.
Figure
• LR-U, a gene-wise likelihood ratio, which is independent of the other genes;
• ODPp-U, an ODP based on the case (a) model, which shares the estimation of model parameters; and,
• ODP-U, an ODP based on the case (b) model, which shares the estimation of both model parameters and hidden variables.
Figures
These results show that sharing the estimation of both the parameter and the hidden variable is effective in gene discovery when they have some structures that can be imprinted by making multiple tests cooperative. Also, we found that the more genes there are in the same group, the more effective is the sharing of the hidden variable estimation.
Next, we consider the semi-supervised case. In the 16 samples, we assumed eight samples (index: 1,2,...,8) to have the true class label 1 and the other eight samples (index: 9,10,...,16) to have the class label 2. For each class, five samples were correctly labeled but the other sample labels were unknown. Note that the group (3) genes were assumed to be inactive in this case, because the sample labels represented by these genes (Fig.
• SAM, SAM statistics;
• LR-S, likelihood ratio score;
• ODP-S, an ODP using only the samples with labels; and,
• ODP-SS, an ODP using all samples with the labels being known or unknown.
The three supervised scores, SAM, LR-S, and ODP-S, used the ten samples with labels to calculate their gene scores, while ODP-SS used, in addition, the remaining six unlabelled samples to estimate the unknowns in the hypothesis models.
The ROC curves in Fig.
In this section, we demonstrate the stability of the proposed semi-supervised and unsupervised scores based on two real data sets.
The Prostate data set [
To overview of the data set, we consult the application software EDGE [
The Colon cancer data set [
We randomly masked sample label information except for 4, 6, 8,
Spearman's rank correlation, a well known criterion for evaluating the correlation between two different rankings, is used for evaluating concordance between two rankings:
where
Figure
In the case of
Based on the Spearman's rank correlation to the provisional answer for the Prostate data set, Fig.
For the Colon data set, on the other hand, the Spearman's rank correlation to the provisional answer (D) did not display any prominent difference between any two rankings in the three objectives. Other criteria, AUC10% (E) and AUC1% (F), however, showed prominent the superiority of ODP-SS or ODP-U in cases when the number of labeled samples was smaller
We did not find any prominent difference between the supervised scores of SAM and ODP-S in these cases. That is, the two scores were almost equivalent with respect to the ability to detect a small number of top-ranked genes, and a small difference existed only in rankings among the top genes.
Discriminant analysis is another important aspect, different from differential gene detection, for selecting a small subset of important genes from microarray experiments. In discriminant analysis, a set of genes is determined to optimize the accuracy of sample prediction based on a certain method, such as using a support vector machine (SVM). As a typical example, the recursive feature elimination (RFE) method is considered to be one relevant way of optimizing the performance of SVMs [
Semi-supervised analysis denotes the cases where unlabeled samples are used along with labeled samples. There have been many studies conducted on semi-supervised learning to improve discriminant accuracy by incorporating it with unlabeled samples from the general viewpoint of machine learning, ex. [
Bair and Tibshirani (2004) [
There had been many problem-specific devices proposed for modeling various effects of noise in microarray studies. Parametric models on noise in microarrays [
Although Storey et al. [
The need for better ranking never declines even when FDR estimation is not available. Gene ranking is typically required as an essential process for selecting candidate genes to reduce the cost of experiments in order to examine whether the top-ranked genes can be targets for medical treatment or can be as possible diagnosis markers. Therefore, ranking performance vital regardless of the existence of a way to determine the threshold by considering an FDR.
In this study, we demonstrated an improvement in gene-ranking estimation in semi-supervised and unsupervised cases for artificial and real data sets. We have shown that the stability of gene-ranking for small number of labeled samples can be drastically enhanced by incorporating unlabeled samples in a semi-supervised manner. On the other hand, for a sufficient number of labeled samples, the difference between the supervised and semi-supervised rankings was not prominent. Consequently, the optimal situation where the semi-supervised ranking of genes will exhibit prominent performance is in cases where the number of labeled samples is small whereas the number of unlabeled samples is large. There will be such a typical semi-supervised situation because sometimes it is not very difficult to measure a lot of samples by means of cheap microarrays, while to apply a label to each sample still incurs a large amount of time, money, and other miscellaneous costs. For example, to apply dead/alive in clinical applications may take several years to follow up.
Although the unsupervised score also exhibited fairly good concordance to the provisional supervised ranking corresponding to the difference between normal and tumor, it will not necessarily behave similarly in other situations corresponding, for example, to the difference between different types of tumors because the unsupervised score may pick up any genes that actively correspond to any differences. This characteristic of the unsupervised score was shown in the experiment using artificial data, where many genes in group (3), whose distributions have multimodality, were picked up even though they had no relationship to the main objective normal/tumor label. Consequently, the unsupervised score can effectively be used as an alternative way of filtering insignificant genes out in the pre-processing of gene expression analyses.
Applying various latent variable models such as a mixture of the t-distribution model, Cox's proportional hazards model, and so on, may lead to useful gene ranking within our framework, and such an application will be our future work.
In general, sharing commonarities of hypothetical models over multiple tests has been thought as an important key of multiple testing problem. And, the ODP framework [
We, in this paper, found that application of the ODP to latent variable models generally has two different possibilities, namely, the estimated hidden variables for other genes are shared or not. As an application of the ODP to latent variable models, we proposed new unsupervised and semi-supervised significance scores of genes based on a simple latent variable model, normal mixture model.
The simulation results indicated that the ODP framework is effective for hypotheses including latent variables and is further improved by sharing the estimations of hidden variables over multiple tests. The real data experiments showed that incorporation of the unlabeled samples by using the proposed unsupervised or semi-supervised score can improve the detectability of active genes.
SO: Conducted research, designed study, designed artificial and real data simulation, wrote manuscript. SI: Wrote manuscript. Both authors read and approved the final manuscript.
This work was supported by KAKENHI (Grant-in-Aid for Scientific Research) on Priority Areas "Applied Genomics" from the Ministry of Education, Culture, Sports, Science and Technology of Japan. The authors declare that they have no competing financial interests.