Individual cells in genetically homogeneous populations have been found to express different numbers of molecules of specific proteins. We investigated the origins of these variations in mammalian cells by counting individual molecules of mRNA produced from a reporter gene that was stably integrated into the cell's genome. We found that there are massive variations in the number of mRNA molecules present in each cell. These variations occur because mRNAs are synthesized in short but intense bursts of transcription beginning when the gene transitions from an inactive to an active state and ending when they transition back to the inactive state. We show that these transitions are intrinsically random and not due to global, extrinsic factors such as the levels of transcriptional activators. Moreover, the gene activation causes burst-like expression of all genes within a wider genomic locus. We further found that bursts are also exhibited in the synthesis of natural genes. The bursts of mRNA expression can be buffered at the protein level by slow protein degradation rates. A stochastic model of gene activation and inactivation was developed to explain the statistical properties of the bursts. The model showed that increasing the level of transcription factors increases the average size of the bursts rather than their frequency. These results demonstrate that gene expression in mammalian cells is subject to large, intrinsically random fluctuations and raise questions about how cells are able to function in the face of such noise
Counting individual mRNA molecules produced by a given gene provides evidence of random fluctuations in mammalian gene expression from cell to cell.
Many recent experiments show that genetically identical populations of bacteria and yeast can exhibit cell-to-cell variations in the amount of protein a gene produces [
While studies in bacteria have shown that variations have partially [
However, direct detection of the proposed gene activation and inactivation events was not possible because new proteins from individual activation events were masked by proteins remaining from previous events as a result of the long half-lives of the fluorescent proteins used as the reporter. The use of fluorescent proteins is further limited by their low sensitivity: because individual molecules of fluorescent protein produce only small amounts of fluorescence, they are difficult to detect in the low numbers produced by many genes. This limitation is particularly troublesome in eukaryotic cells in general, and higher eukaryotes in particular, as the cellular volumes are much larger than those of bacteria, thus diluting the fluorescent protein concentration.
Given these limitations, the most direct way to detect gene activation and inactivation is to directly monitor the mRNA produced from the gene at the resolution of single molecules. Because the half-life of mRNA is typically much shorter than that of fluorescent proteins, their levels reflect more accurately the state of the gene. Moreover, by detecting single molecules, one would sidestep the issue of sensitivity. Furthermore, the presence of integral molecule counts would be especially valuable in precisely evaluating models of stochastic gene expression.
In this study, we explored cell-to-cell variation in gene expression in mammalian cells by accurately counting single molecules of mRNA through the use of fluorescence in situ hybridization (FISH). By obtaining precise measurements of these most immediate (and fast decaying) products of gene expression, we show direct evidence that genes transition infrequently between active and inactive states, resulting in large cell-to-cell variations in gene expression in clonal cell lines. In contrast to the mostly extrinsic variations observed in yeast, we show that these transitions are intrinsically random and not due to extrinsic factors and that they affect the expression of entire genomic loci. Furthermore, we found that the mRNA produced by the gene encoding the large subunit of RNA polymerase II is also produced in bursts, and that these bursts are uncorrelated with those from our reporter gene, indicating that the level of RNA polymerase II is not an important extrinsic determinant of cell-to-cell variations. We also analyzed the effect that the mRNA variations had on the proteins they encode and found that having a slow protein degradation rate can serve to buffer the mRNA variations. A mathematical model of gene activation and inactivation indicates that the mean number of mRNA produced per activation event (the “average burst size”) can be controlled by varying the amount of transcriptional activators in the cell.
To directly observe random events of gene activation and inactivation, we measured cell-to-cell variations in the number of molecules of a specific mRNA. We accomplished this by integrating a reporter gene possessing a tandem array of probe binding sites into mammalian cells and utilized fluorescently labeled probes to visualize mRNA transcribed from the gene by FISH. To obtain single molecule sensitivity, we introduced 32 tandem copies of a 43–base-pair probe-binding sequence at the 3′ end of a coding sequence for a fluorescent protein (throughout this paper, we refer to this sequence array as M1). The construct was inserted into Chinese hamster ovary (CHO) cells by electroporation, and a stable cell line was isolated in which a single copy of the gene was integrated into the cells' genome. These cells were then fixed and subjected to hybridization with a single-stranded oligodeoxyribonucleotide probe that was both complementary to the tandemly repeated sequences and labeled with five well-dispersed fluorophore moieties (
(A) Schematic diagram depicting the mRNA detection method. Multiple fluorescent probes bind to each mRNA molecule, yielding a bright, localized signal.
(B) Merged image of a three-dimensional stack of images from a CHO cell expressing the 7x-tetO gene depicted in (A), where each mRNA is hybridized to FISH probes that bind to the multimeric probe-binding sequence in its 3′-UTR (probe P1-TMR binding to the M1 multimer). Each spot corresponds to a single mRNA molecule.
(C) Identification of the spots in the three-dimensional image stack in (C). Each particle found by the image-analysis algorithm is colored differently, showing that the algorithm is accurate and that individual molecules are uniquely identified. The scale bars are 5 μm long.
To measure cell-to-cell variations in mRNA numbers in clonal cell lines, we generated stable CHO cell lines expressing our construct. The gene was placed under the control of a promoter whose expression could be controlled in mammalian cells (
(A) Schematic diagram of the doxycycline-controllable promoters and the reporter genes that they control. Doxycycline binds to the tTA protein, thereby preventing it from binding to the
(B, C) Representative fields of cells from cell lines E-YFP-M1-1x and E-YFP-M1-7x, containing the 1x-tetO and 7x-tetO promoters, respectively, where each mRNA is hybridized to FISH probe P1-TMR and the image was obtained by merging a three-dimensional stack of images.
(D) Two sister cells from cell line E-YFP-M1-7x displaying mRNA hybridized to FISH probe P1-TMR (red) and costained with DAPI (blue). The image represents one focal plane. The scale bars are 5 μm long.
Quantitative evidence of the burst-like nature of transcription comes from comparing the number of mRNA in cells containing active transcription sites to those without active transcription sites. We found that of 97 randomly selected cells from cell line E-YFP-M1-7x (details of construct discussed below), the 23 containing transcriptional foci had an average of 244 mRNA per cell, as compared to 33 mRNA per cell in the 74 without any active transcription site (
Further evidence for transcriptional bursts comes from an analysis of the statistics of the distribution of mRNA molecules per cell over the entire cell population. If mRNA were produced at a constant rate, one would expect a Poisson distribution of mRNA per cell, in which case the mean number of mRNA molecules per cell and the variance (the square of the standard deviation) would be equal. However, we found that the mean was approximately 40 mRNA molecules per cell, while the variance was roughly 1,600 molecules squared, indicating that the mRNA is not synthesized at a constant rate, consistent with the occurrence of transcriptional bursts.
To investigate the mechanisms controlling transcriptional bursts, we altered the overall level of transcription both by changing the amount of transcriptional activator present in the cells and by altering the number of binding sites for that activator in the promoter. To accomplish this, the gene was inserted downstream from a minimal cytomegalovirus promoter, and either one or seven copies of the tetracycline-sensitive
Two constructs (1x-tetO and 7x-tetO) were stably integrated into CHO cells that had previously been modified to express tTA, resulting in the cell lines E-YFP-M1-1x and E-YFP-M1-7x, each containing a single copy of the respective reporter gene. Representative snapshots of clonal cell fields are shown in
(A) Histograms showing the distribution of mRNA molecules per cell for three doxycycline concentrations for both cell lines E-YFP-M1-1x (left) and E-YFP-M1-7x (right).
(B) Graphs showing the population mean (top) and noise (defined as the standard deviation divided by the mean [bottom]) as a function of doxycycline concentration. Statistics were taken from the mRNA counts used in (A), and error bars were obtained by bootstrapping.
(C) Activation rate (λ/δ) and average number of mRNA molecules produced per burst (μ/γ) for 1x-tetO (blue) and 7x-tetO (red) obtained from fitting the mRNA count data to the model of gene activation and inactivation by the maximum-likelihood method. Error bars reflect 95% confidence intervals.
Both results are inconsistent with conventional stochastic models of gene expression [
If the variation in expression levels from one cell to another were truly due to random gene activation events, then the presence of multiple independently activating copies of the gene would result in less cell-to-cell variability in mRNA numbers (i.e., the noise should decrease). Intuitively, this can be seen by considering simultaneous coin tosses: if only one coin is tossed, it is either heads or tails, but if several coins are tossed at once, the chances of the set of them being close to 50% heads and 50% tails increases with the number of coins used. To test this possibility, we integrated multiple copies of our reporter gene into one region of the genome via cationic lipid-based transfection (lipofection), which simultaneously integrates tens to hundreds of gene copies, often in tandem and at the same locus, and isolated cell line L-GFP-M1-7x. Generally, the number of mRNA produced in this cell line was much larger than in E-YFP-M1-7x (with only one copy of the reporter gene), but the cell line still displayed massive cell-to-cell variations (
(A) Representative field from cell line L-GFP-M1-7x, generated by lipofection, where the mRNA was hybridized to FISH probe P1-TMR; the image was obtained by merging a three-dimensional stack of images.
(B) Histogram showing the distribution of mRNA molecules per cell over for cell line L-GFP-M1-7x when grown in media containing no doxycycline.
(C) Graphs showing the population mean (top) and noise (defined as the standard deviation divided by the mean [bottom]) as a function of doxycycline concentration. Error bars were obtained by bootstrapping.
There are two alternate explanations for this observation. One possibility is that the massive fluctuations seen in the number of mRNA molecules per cell are due to fluctuations in global factors that simultaneously affect the expression of all of the reporter genes (e.g., fluctuations in the levels of tTA or RNA polymerase II); this is usually referred to as extrinsic noise [
To explore these two alternative explanations, we constructed another reporter gene, CFP-M2, that encoded a cyan fluorescent protein and contained a different tandem array of probe-binding sequences in its 3′-UTR, denoted M2. This allowed its mRNA to be distinguished from the mRNA synthesized from reporter genes containing the M1 sequence array by performing FISH with an additional probe that binds to the M2 array but is conjugated to a differently colored fluorophore. In one series of experiments, this gene was integrated into a cell line (L-GFP-M1-7x) that already expressed a reporter gene containing the M1 array, resulting in the CFP-M2 reporter gene being integrated into a locus distant from the site of integration of GFP-M1 (
(A) Schematic depicting dual-reporter integration experiments, where the two reporters are either integrated in separate locations in the genome (left) or at the same locus (right).
(B) Two-color overlay showing a merged three-dimensional stack of images of the cell line in which two different reporter genes (one expressing mRNA containing the M1 sequence array and the other expressing mRNA containing the M2 sequence array) are integrated into separate loci (cell line L-GFP-M1-7x+L-CFP-M2-7x); green corresponds to the signal from GFP-M1 mRNA and red corresponds to the signal from CFP-M2 mRNA (using probes P1-TMR and P2-Alexa-594, respectively). The relative mRNA levels of both genes were quantified by counting the mRNA of each color in a single optical slice in each cell (inset).
(C) Two-color overlay showing a merged three-dimensional stack of images of the cell line in which the two distinct reporter genes are integrated into the same locus (cell line L-YFP-M1-CFP-M2), where the same FISH probes were used as in (B). The relative mRNA levels of both genes were quantified by counting the mRNA of each color in a single optical slice in each cell (inset). The scale bars are 5 μm long.
When integrated at separate loci (as evidenced by the presence of two distinct transcription sites), the two reporter mRNAs each individually displayed the large fluctuations observed previously (
To further examine the role of global, extrinsic factors, we checked for fluctuations in a putative extrinsic factor, RNA polymerase II, to see if the level of expression of its mRNA correlated with the level of expression of the mRNA from a reporter gene. We were able to image individual molecules of the natural mRNA encoding the large subunit of RNA polymerase II by exploiting the presence of a naturally occurring 21-nucleotide-long sequence that is repeated 52 times in the mRNA. We used a FISH probe for the repeated sequence that was labeled with a distinctively colored fluorophore, and we counted fluorescent spots similar to those observed previously (
(A) Representative field of cell line E-YFP-M1-7x upon performing FISH with differently colored probes for both YFP-M1 mRNA and the mRNA encoding the large subunit of RNA polymerase II. The image is a two-color overlay, where green corresponds to the signal from YFP-M1 mRNA (one optical slice) and red corresponds to the signal from RNA polymerase mRNA (merged three-dimensional image stack). The probes used were P1-Cy5.5 and P3-TMR.
(B) Histogram showing the distribution of RNA polymerase mRNA molecules per cell (top) and scatterplot (bottom) showing reporter mRNA levels (quantified by counting the mRNA in a single optical slice) and RNA polymerase mRNA levels (quantified by counting all mRNA) in each cell. Gene activation (λ/δ), inactivation (γ/δ), and transcription rates (μ/δ) (normalized by the mRNA decay rate δ) are given with 95% confidence intervals as indicated. The scale bar is 5 μm long.
To investigate the effects that burst-like transcription of mRNA had upon intracellular protein levels, we simultaneously quantified the number of mRNA and the fluorescent proteins they encoded in individual cells. To assess the effects that the rate of protein degradation had upon protein variability, we performed this analysis on a cell line expressing a fluorescent protein that was actively degraded and another cell line expressing a fluorescent protein with no active degradation.
For the case of active degradation, we used a cell line stably expressing the GFP-M1 reporter gene, which encoded a green fluorescent protein (GFP) that had been tagged at the C-terminus with a short amino acid sequence rich in proline, glutamic acid, serine, and threonine (d2EGFP) that targets the protein for active degradation (with a half-life of approximately 2 h) and examined the correlations between the mRNA and protein levels (
(A) Scatterplot of total GFP and mRNA numbers in individual cells from cell line L-GFP-M1-7x grown under conditions of no doxycycline (blue), 0.08 ng/ml doxycycline (red), and 0.16 ng/ml doxycycline (green). Marginal histograms indicate the distribution of reporter mRNA per cell (top) or total GFP (right) for all the growth conditions.
(B) Scatterplot of total CFP and mRNA levels (obtained from a single optical slice) in individual cells from cell line L-GFP-M1-7x+L-CFP-M2-7x grown under conditions of no doxycycline. Marginal histograms indicate the distribution of reporter mRNA per cell (top) or total GFP (right) for all the growth conditions.
(C) Histogram of total YFP per cell from cell line E-YFP-M1-7x. YFP was quantified in live cells to minimize loss of fluorescence due to fixation and permeabilization.
(D) Scatterplots and associated marginal histograms showing the results of stochastic simulations of the model of mRNA and protein dynamics presented in
For the case of no active protein degradation, we used another cell line, this time expressing the CFP-M2 reporter gene, in which the fluorescent protein did not contain any degradation tags. We found that the correlation between the mRNA and protein levels was significantly lower than before (
To explain the differences in protein distributions and correlation between the actively degraded and nondegraded cases, we added protein dynamics to the model of mRNA dynamics and examined the behavior of this model through the use of Gillespie's stochastic simulation algorithm [
We have shown that the mRNA levels of both reporter genes and native genes display large cell-to-cell variations in mammalian cells due to intrinsically random, infrequent events of gene activation. These burst-like fluctuations are not restricted to engineered reporter genes, but occur in natural genes as well, as demonstrated for the mRNA encoding the large subunit of RNA polymerase II. We have further shown that these events are controlled by gene regulatory mechanisms, such as the level of activator proteins and the number of transcription factor binding sites, and can affect regions of the genome rather than just specific genes.
Moreover, we have found that the variations are intrinsically random, rather than due to global extrinsic factors. This contrasts with the results of previous studies in lower eukaryotes [
We have also shown that the statistics of these variations are well described by a model in which the only sources of randomness are random events of gene activation and inactivation, implying that one can safely ignore the randomness inherent in the chemical reactions describing transcription and translation. This is qualitatively different from the bacterial case, where such reactions are thought to be the dominant source of variability in gene expression [
The most likely sources for the transcriptional bursts are random events of chromatin remodeling [
If the decondensation of chromatin is a prerequisite for gene activation, then the nucleation of this decondensation will be a significant rate-limiting step. The structure of chromatin at the level of nucleosome stacking suggests that the “breathing” events that permit the entry of transcriptional regulators will be infrequent [
The fact the mRNA is produced in bursts points to new means by which the cell may control transcription. There are three apparent means by which a cell would be able to upregulate a gene's transcription: it could (i) increase the rate of gene activation, (ii) increase the rate of transcription when the gene is in the active state, or (iii) decrease the rate of gene inactivation (the opposite behaviors, of course, apply should a cell decide to downregulate a gene's transcription). These mechanisms, while all resulting in the same
Our mathematical treatment of stochastic gene expression is rather different than methods based on moment generating functions [
In a wider sense, the presence of such large, unpredictable fluctuations in gene expression may initially appear to be a significant impediment to the functioning of a cell. In particular, given its essential role in cellular function, it is surprising that the gene encoding the large subunit of RNA polymerase II also displays fluctuations on the order of those seen in the reporter genes. Our analysis of protein levels yields a resolution to this apparent paradox: if the degradation rate of the proteins is sufficiently small, then the variations in protein level will be buffered because the proteins from new bursts serve only to “top up” the proteins already present from previous bursts. This suggests that essential genes whose mRNA expression is burst-like should have relatively stable proteins. Moreover, there are other manners in which protein variations may be further reduced. For instance, should two different proteins, each bursting independently, form a heteromeric complex, then the variations in the number of complexes will be somewhat buffered from the variations in each component.
Conversely, there may also be situations in which burst-like expression of unstable proteins may be desirable as well. Many examples of such situations exist in bacteria and yeast, often as a result of multistable behavior [
Construction of the DNA fragment with 32 probe binding sites in the pGEM-f11(z-) cloning vector (Invitrogen, Carlsbad, California, United States) was performed by the method described in Robinett et al. [
The underlined portions of the sequence correspond to the SalI, XhoI, and BamHI sites used for integration into the host vector. A similar procedure was used to produce the M2 32-mer described by Vargas et al. [
The reporter genes were constructed by adding an open reading frame for yellow fluorescent protein (YFP) and cyan fluorescent protein (CFP) upstream of the M1 and M2 multimers in the pGEM-M1-32x and pGEM-M2-32x plasmids, respectively. The sequences encoding YFP and CFP were amplified via PCR from pDH3 and pDH5 (University of Washington, Yeast Resource Center, Seattle, Washington, United States) and inserted in front of the M1 and M2 multimers between the SacI and SalI restriction sites, also introducing a BglII restriction site between the SacI site and the start codon of the open reading frame.
These reporter genes were then integrated into expression vectors enabling their expression in mammalian cells. The base vectors chosen were the pTRE2Hyg, pTRE2Pur, and pTRE-d2EGFP vectors (Clontech, Palo Alto, California, United States). Each contains a tetracycline-responsive promoter containing seven copies of the tet operator followed by a minimal cytomegalovirus promoter and a polyadenylation signal. Additionally, the pTRE2Hyg and pTRE2pur vectors enabled selection by appropriate quantities of hygromycin B (Invitrogen) or puromycin (Sigma, St. Louis, Missouri, United States).
To create a vector with one copy of the tet operator, we amplified the promoter region of pTRE2Hyg using a primer containing one copy of the tet operator. This was then cloned back into the pTRE2Hyg promoter site, replacing the native promoter and creating the plasmid pTRE2Hyg1x. The YFP-M1 construct was then extracted from the pGEM host vector with BglII and NotI and then inserted into the pTRE2Hyg and pTRE2Hyg1x vectors between the BamHI and NotI sites. The CFP-M2 construct was similarly inserted into the pTRE2Hyg and pTRE2Pur vectors. This created the plasmids pTRE2Hyg-YFP-M1, pTRE2Hyg1x-YFP-M1, pTRE2Hyg-CFP-M2, and pTRE2Pur-CFP-M2. The M1 multimer was also inserted into 3′-UTR in the pTRE-d2EGFP vector between the BamHI and EcoRI sites, creating the plasmid pTRE-d2EGFP-M1. All constructs were verified by sequencing.
All cell lines were derived from the CHO-AA8-Tet-off cell line (Clontech), which possesses a stably integrated gene expressing the tetracycline-controlled Tet-off transactivator. Cell lines E-YFP-M1-1x and E-YFP-M1-7x, containing the 1x-tetO and 7x-tetO constructs, were generated by electroporation (Bio-Rad, Hercules, California, United States) using plasmids pTRE2Hyg1x-YFP-M1 and pTRE2Hyg-YFP-M1, respectively. The electroporator settings were 360 V and 975 μF using 10 μg of plasmid DNA linearized with XmnI added to 107 cells in 500 μl of PBS in a 4-mm gap cuvette. The multiple-copy integration clones were generated using LipofectAMINE 2000 (Invitrogen) by following the manufacturers instructions. Cell line L-GFP-M1-7x was created by transfecting the CHO-AA8-Tet-off cell line with the pTRE-d2EGFP-M1 plasmid linearized with ScaI. To create the cell lines in which two reporter genes were integrated in different loci, cell line L-GFP-M1-7x was transfected using LipofectAMINE 2000 with the plasmid pTRE2Hyg-CFP-M2, linearized with XmnI; the resultant cell line is L-GFP-M1-7x+L-CFP-M2-7x. To create the cell line in which two reporter genes were integrated in the same locus, the CHO-AA8-Tet-off cell line was simultaneously cotransfected with equal amounts of pTRE2Hyg-YFP-M1 and pTRE2Pur-CFP-M2, resulting in cell line L-YFP-M1-CFP-M2. Cell lines were isolated after transfection by either electroporation or lipofection via selection with the appropriate antibiotic (hygromycin B or puromycin) and then purified by serial dilution. That only one copy of the transgene was integrated into cell lines E-YFP-M1-1x and E-YFP-M1-7x was verified by Southern blotting upon digestion of genomic DNA with the restriction enzyme BglII. Several cell lines were isolated following transfection; all exhibited similar phenotypes to the cell lines analyzed in this paper. Cell lines obtained from lipofection with pTRE-d2EGFP-M1 were isolated not by antibiotic selection but instead by directly identifying fluorescent cell clusters and purifying by serial dilution. Stability of the gene was verified by DNA FISH (unpublished data).
Cells were cultured in the alpha modification of Eagle's minimum essential medium (Sigma) supplemented with 10% TET-System-Approved fetal bovine serum (Clontech). The growth medium was supplemented with a low concentration of the selective antibiotic to ensure stability of the transfected gene. Appropriate amounts of doxycycline were added as indicated, and cells were grown at the desired concentration of doxycycline for 4 d to minimize any transient effects. The doxycycline concentration experiments were all performed in parallel with the same batch of media to minimize differences due to serum composition, etc.
The probes used for the in situ hybridization were DNA oligonucleotides synthesized on an Applied Biosystems (Foster City, California, United States) 394 DNA synthesizer using mild phosphoramidites (Glen Research, Sterling, Virginia, United States). The oligonucleotide sequences were P1: 5′-CGGCRGGTAAGGGRTTCCATARAAACTCCTRAGGCCACGA-3′; P2: 5′-RCGAGGTCGARCAGCTGGCTGGRGCTCTTCGRCCACAAACA-3′; and P3: 5′-AGAGGRGGGCGAGRAGCRGGGAGAGGRGGGCGAGRAGCRGGG-3′, where P1 and P2 are complementary to the repeat sequences in M1 and M2, respectively, and P3 is complementary to a consensus sequence for the 52 repeats in the RNAPII subunit A cDNA sequence. The “R”s represent locations where an amino-dT was introduced in place of a regular dT. The oligonucleotides were synthesized on a controlled pore glass column (Glen Research) that introduced an additional amine group at the 3′ end of the oligonucleotide. The probes' amine groups were then coupled to the fluorophores Cy5.5, Alexa 594, and tetramethylrhodamine (Molecular Probes, Eugene, Oregon, United States) to create the following probes: P1-TMR, P1-Cy5.5, P2-Alexa-594, and P3-TMR. The probes were purified on an HPLC column to isolate oligonucleotides displaying the highest degree of coupling of the fluorophore to the amine groups.
Cells were cultured in multichambered coverglass (Lab-Tek, Nalge Nunc, Rochester, New York, United States) coated with gelatin. The cells were fixed with 3.7% formaldehyde for 10 min at room temperature, washed with 1× PBS, and permeabilized for at least 1 h in 70% ethanol. FISH was then performed using combinations of probes P1, P2, and P3 at a concentration of 1 ng/μl each following the procedure outlined in Femino et al. [
After in situ hybridization, cells were imaged using an Axiovert 200M inverted fluorescence microscope (Zeiss, Oberkochen, Germany), equipped with a 100× oil-immersion objective and a CoolSNAP HQ camera (Photometrics, Pleasanton, California, United States), and cooled to −30 °C; then, standard filter sets obtained from Omega Optical (Brattleboro, Vermont, United States). Openlab acquisition software (Improvision, Sheffield, United Kingdom), was used to acquire the images. For three-dimensional imaging, randomly chosen fields were imaged by taking adjacent Z-axis optical sections that were 0.3 μm apart. The particles were counted in three dimensions using custom software written in MATLAB (The Mathworks, Natick, Massachusetts, United States). The general procedure was to (i) manually select the individual cells in a field, then (ii) run a median filter on each optical slice taken, then (iii) run a custom linear three-dimensional filter, designed to enhance particulate signals and loosely based on the discrete Laplacian, on the stack of images, then (iv) manually select a threshold for the enhanced images, and then (v) count the total number of isolated signals (i.e., connected components) in three dimensions. For each manually selected threshold, other thresholds 25% above and below were also analyzed to verify that the particle count did not depend significantly on the particular threshold chosen. Our best estimate is that the number of spots counted by our algorithm is accurate to within 10% of the actual number. In cells with transcription sites, the transcription site itself was subjected to the same counting procedures, usually resulting in it being counted as a single molecule. This is justified, since the nascent RNAs present at the transcription site are mostly likely unprocessed pre-mRNA that have not yet been subjected to the various post-transcriptional modifications required for an mRNA to be considered functional [
In experiments where fluorophores other than TMR were used, we instead quantified the relative amount of mRNA from cell to cell by counting the number of mRNA in one optical section (chosen near the bottom of the cellular volume). This was done because the relatively poor photostability of the Cy5.5 and Alexa 594 dyes meant that the particulate signal became quite weak during the acquisition of the image stacks, making the imaging and counting of individual molecules progressively more difficult and thus significantly less accurate. In the case of the L-GFP-M1-7x clone, the large number of mRNA molecules often resulted in signals too intense to quantify using the segmentation method above due to overlap in the diffraction-limited spots. In this case, we quantified the mRNA by integrating the total fluorescence over the entire cellular volume. To relate this to the absolute number of mRNA in the cell, we quantified the mRNA in several test cells where the counting procedure was reliable using our particle counting algorithm and correlated that to the total fluorescence within the volume. The relationship was found to be linear (
The fluorescent protein levels were quantified by a single fluorescence image toward the lower focal plane of the cells. The total fluorescence was found by integrating the difference between the pixel intensities and the average background over the entire cellular area. In the case of the live cell YFP images, OptiMEM (Sigma) was used as growth medium because of its reduced autofluorescence as compared with regular MEM.
All software is available upon request.
The error bars for the means and noises reporter were obtained by the bootstrap method. The parameters from the model were estimated using the maximum-likelihood method based on an explicit formula derived for the complete mRNA distribution as outlined in
The
The
The mRNA decay rate was found by using real-time RT-PCR on RNA isolated from cell line L-GFP-M1-7x grown in medium containing 10 ng/ml doxycycline for a range of times using the Qiagen One-Step RT-PCR kit (Valencia, California, United States). The real time RT-PCR was performed for both the GFP transgene and the highly expressed elongation factor 1 gene, which served as an internal control not likely to change in response to doxycycline concentration. We used primers and molecular beacons specific to each gene to perform real-time PCR. The difference in threshold cycle between the GFP and the EF1 signals was linearly related to the time since transcription was halted, allowing an accurate determination of the half-life of the mRNA from the transgene. The results are shown in
Stochastic simulations of the stochastic mRNA and protein model described in
Plot shows the difference in threshold cycle between PCRs performed on the GFP reporter gene and the EF1 housekeeping from a real-time RT-PCR experiment performed on total mRNA extracted from cell line L-GFP-M1-7x. At time 0, the cellular media was replaced with media containing 10 ng/ml doxycycline, effectively shutting down transcription of the reporter gene, thus allowing for a determination of the mRNA degradation time. In determining the half-life, we only considered the rightmost three points, since early time points may display non–first-order degradation due to transient effects of mRNA processing and export.
(61 KB PDF)
Click here for additional data file.
Plot shows the noise (defined as the standard deviation divided by the mean) of mRNA concentrations for the 1x-tetO construct (red) and the 7x-tetO construct (blue) over a range of doxycycline concentrations. The mRNA concentration was determined by dividing the number of mRNA by the total volume of the cell, as determined by microscopy. Compare to
(61 KB PDF)
Click here for additional data file.
Plot shows the correlation between the number of mRNA molecules and total fluorescence integrated over the cellular volume (background subtracted). Cells were taken from random fields of cell line L-GFP-M1-7x and cells were chosen for both reasonable levels of mRNA to allow for segmentation and for lack of brightly fluorescent features that could potentially influence the calibration. The linear fit indicated a value of roughly 53,500 fluorescence units per molecule of mRNA.
(60 KB PDF)
Click here for additional data file.
Mechanistic model of bursts in mRNA synthesis, basic model of gene activation and inactivation, fitting of parameters to experimental distributions, determination of protein mean and variances, and relationship between model and previous studies of intrinsic versus extrinsic noise.
(106 KB PDF)
Click here for additional data file.
(21 KB XLS)
Click here for additional data file.
Each white spot represents a molecule of mRNA. The dense white spot in one of the cells is an active transcription site.
(6.0 MB MOV)
Click here for additional data file.
Green corresponds to the signal from YFP-M1 mRNA, and red corresponds to the signal from CFP-M2 mRNA. The dense yellow spot in both cells is an active site of transcription, indicating that both mRNAs are being transcribed at the same genomic locus.
(4.0 MB MOV)
Click here for additional data file.
We thank F. Kramer for a critical reading of the manuscript, S. Marras for assistance with synthesizing the fluorescent probes, and S. Isaacson and P. van den Bogaard for discussions.
Chinese hamster ovary
fluorescence in situ hybridization
green fluorescent protein
tet-transactivator