Conceived and designed the experiments: DD MC. Performed the experiments: DD MC. Analyzed the data: DD MC. Contributed reagents/materials/analysis tools: DD MC. Wrote the paper: DD MC.
One of the main obstacles to understanding complex biological systems is the extent and rapid evolution of information, way beyond the capacity individuals to manage and comprehend. Current modeling approaches and tools lack adequate capacity to model concurrently structure and behavior of biological systems. Here we propose Object-Process Methodology (OPM), a holistic conceptual modeling paradigm, as a means to model both diagrammatically and textually biological systems formally and intuitively at any desired number of levels of detail. OPM combines objects, e.g., proteins, and processes, e.g., transcription, in a way that is simple and easily comprehensible to researchers and scholars. As a case in point, we modeled the yeast mRNA lifecycle. The mRNA lifecycle involves mRNA synthesis in the nucleus, mRNA transport to the cytoplasm, and its subsequent translation and degradation therein. Recent studies have identified specific cytoplasmic foci, termed processing bodies that contain large complexes of mRNAs and decay factors. Our OPM model of this cellular subsystem, presented here, led to the discovery of a new constituent of these complexes, the translation termination factor eRF3. Association of eRF3 with processing bodies is observed after a long-term starvation period. We suggest that OPM can eventually serve as a comprehensive evolvable model of the entire living cell system. The model would serve as a research and communication platform, highlighting unknown and uncertain aspects that can be addressed empirically and updated consequently while maintaining consistency.
Recent years have witnessed unprecedented increases in the number, variety and complexity of information resources available to researchers in the life sciences. We are at a turning point in biological research, where emphasis is shifting from the study of a single molecular process to studying complete cellular pathways and the entire cell as a system. This pivotal time for the life sciences is captured by the words of Kitano: “a
In November 2006, Nature Cell Biology and Nature Reviews Molecular Cell Biology published jointly on the Web (
Also in the editorial, after noting how molecular cell biology is emancipating itself from an informal, reductive, hypothesis-driven approach by embracing high-throughput data acquisition, rigorous quantification and mathematical modeling, it is forecast by Kritikou et al.
Adopting a mechanistic view, Davidson et al.
The number of interactions, processes, and transport activities in the living cell is enormous. Therefore, often preceding a quantitative problem of how much or to what extent is the qualitative one of figuring out how and what. Thus, a combined qualitative then quantitative conceptual modeling approach, such as the one adopted in this work, plays a crucial role in facilitating human comprehension of complex cellular mechanisms. Conceptual models advocate the construction of primarily qualitative models, in which biological concepts are put in context with each other in an attempt to gain insights into the function, structure, and dynamics of the biological systems under study. Once a particular, relatively small subsystem in a specific cell location is understood well enough, mathematical tools, such as differential equations, can successfully describe time varying changes. In what follows we survey briefly current approaches and software environments for modeling biological systems, highlighting their advantages and disadvantages.
Modeling efforts in biology are sometimes classified according to their focus on quantitative vs. qualitative aspects. However, oftentimes the two approaches cannot be separated; for example understanding a qualitative process such as the mechanism regulating cell division requires quantitative understanding of this system's dynamics. It must be appreciated that a complex network of protein interactions that influence the activities of cyclin-dependent kinases control major events of the cell cycle, including DNA synthesis, mitosis and cell division.
One quantitative modeling environment is E-Cell
Another example of a mathematics-oriented software modeling environment is the Virtual Cell
Investigating multi-cellular organisms by constructing their conceptual models has been promoted by Harel
Some quantitative approaches combine the concept of intelligent computer programs, commonly known as software agents, with mathematical models. Applying an OO and agent-based approach, Webb and White
Conceptual modeling originated with efforts to streamline software development some three decades ago. Therefore, the object-oriented approach, which is the currently accepted paradigm in the software engineering community, has been very popular in recent years for modeling systems in general and biological systems in particular. The object-oriented (OO) approach advocates that objects are the prime entities or building blocks of software systems, and this notion has been recently extended via SysML (
BioUML,
UML is built on the premise that “Modeling is the designing of software applications before coding”
In general, the current Object-Oriented approach to modeling and developing software systems is not suitable for representing effectively biological concepts, because, as argued, it cannot model concurrently in a single type of diagram both the objects, e.g., a protein, and the processes, e.g., transcription, that transform (create, consume, or change the state of) these objects. UML 2.0
Systems Biology Markup Language, SBML
The recent Systems Modeling Language, SysML
Kohn
In general, when considering human-readable diagrammatic representations, it is notable that the current informal ways most biologists draw diagrams means that correct biological interpretation depends entirely on the reader's knowledge
Using different arrowhead shapes CellDesigner diagrams focus on conveying the semantics of several process types prevalent in signaling, such as translocation, catalysis, splitting, phosphorylation, or state transition. Formalized process diagrams have been used to describe signal-transduction cascades and pathway maps, and are readable and precise as long as the network is not too large. However, scalability is an issue. There is no way to refine mechanisms and designating new pictograms for each new reaction type is problematic. Moreover, as Blinov et al.
In an attempt to solve this problem, Blivnov et al. have introduced graphical rules to allow the connectivity of proteins in a complex to be represented explicitly. These rules provide a means to visualize comprehensibly protein-protein interactions. Nevertheless, for more general biological modeling, process maps are likely to be of unmanageable size due to combinatorial complexity.
An example of a combined quantitative-qualitative model is the work of Tyson et al.
BioTapestry
The current state of affairs, reflected in this survey, is that there are quite a number of modeling approaches and software environments for modeling biological systems, but many of them are object-oriented, hindering direct and explicit process modeling, which is at the heart of systems biology. Moreover, since most approaches are non-scalable and specialized for specific types of cellular reactions or subsystems, they cannot be extended naturally to modeling the entire cell, not to mention organisms, societies, habitats and ecologies.
Perhaps most importantly, according to the definition of systems biology, models are supposed to advance research, yet none of the existing modeling approaches or systems have been shown to promote innovative questions that trigger experiments to confirm or refute assumptions emerging from the model. Such a disappointing situation indicates that a totally different modeling approach is in order. This situation was a major stimulus for the work presented here, a non-traditional conceptual modeling approach that has already stimulated a new empirical finding concerning the mRNA lifecycle.
Representing the vast amount of ever increasing knowledge formally, yet accessibly, can be compared to putting the pieces of a gigantic puzzle together, mandating adoption of a common evolving grand model. The model needs to be founded on a compact generic set of the most basic ontological building blocks in order for it to be general enough to serve as a basis for modeling the gamut of biological systems, from molecules to ecosystems.
We submit that stateful objects and processes that transform them, along with several types of links, as advocated by OPM—Object-Process Methodology
In the present study we aimed to carry out a modest first step toward this admittedly ambitious goal. As the yeast mRNA life cycle is a key cellular system, we chose it to be our case in point for conceptual modeling that employs Object-Process Methodology. In the next section we explain in more detail why we chose Object-Process Methodology, OPM.
Object-Process Methodology, OPM
The elements of OPM ontology are entities and links. A complete list of OPM elements with their symbols and definitions is provided in
A link can be structural or procedural. A structural link expresses a static, time-independent relation between pairs of entities. The four fundamental structural relations are: aggregation-participation, generalization-specialization, exhibition-characterization, and classification-instantiation. An example of using the aggregation-participation structural relation is derived from the phrase “The eukaryotic cytoskeleton is composed of microfilaments, intermediate filaments and microtubules.” (
The Gene-Ontology (GO)
Two semantically equivalent modalities, one graphic and the other textual, are used to describe each OPM model. A set of inter-related Object-Process Diagrams (OPDs), showing portions of the system at various levels of detail, constitute the graphical, visual OPM formalism. Each OPM element is denoted by a symbol in an OPD, and the OPD syntax specifies correct and consistent ways by which entities can be connected via structural and procedural links, such that each legal entity-link-entity combination bears specific, unambiguous semantics. OPCAT
The Object-Process Language (OPL), which is the textual counterpart of the graphical OPD, is a dual-purpose language, oriented towards humans as well as machines. Catering to human needs, OPL is designed as a subset of English, which serves domain experts (e.g., biologists) and system architects, engaged jointly in modeling a complex system. Every OPD construct is expressed by a semantically equivalent OPL sentence or phrase. According to the modality principle of the cognitive theory of multimodal learning
mRNA Lifecycle consumes Amino Acid Set, Ribonucleotide Set, and mRNA.
mRNA Lifecycle yields mRNA, Ribonucleotide Set, and Protein.
(A) An Object-Process Diagram (OPD) showing the top-level, bird's eye view of the process mRNA Lifecycle, which generates the object Protein from the object Amino Acid Set. The process, marked as the blue ellipse, shows the generation of mRNA from Ribonucleotide Set and the degradation of mRNA back to Ribonucleotide Set. (B) The corresponding Object-Process Language (OPL) text that was generated automatically by OPCAT, the software that supports OPM modeling.
Examining these sentences, we see that each arrow from an object to the process (mRNA Lifecycle) has the semantics of consumption (degradation, or destruction), while each arrow from the process to an object has the semantics of result (creation, or generation). During the process mRNA Lifecycle, Ribonucleotide Set and mRNA are both consumed and generated, while Amino Acid Set is consumed to create Protein.
A major problem with most graphic modeling approaches is their scalability. As the system's complexity increases, the graphic model becomes cluttered with symbols and their connecting links. The limited channel capacity
(A) An OPD in which the mRNA Lifecycle process from
The expression of protein encoding genes is a complex process that determines which genes are expressed as proteins at any given time, as well as the relative levels of these proteins
Transcription by pol II, the first stage in the expression of protein-encoding genes, produces RNA–the primary transcript. This primary transcript is processed to yield an mRNA (usually shorter than the primary transcript) that contains a 5′ cap (m(7)GpppN) and 3′ poly(A) tail. These two tags are critical for the appropriate function, localization and stability of the mRNA
Recently, a new venue for mRNA localization was uncovered that revolutionized our view of how gene expression is regulated post-transcriptionally. Specifically, yeast mRNA can be localized in discrete cytoplasmic foci together with a number of mRNA decay factors and limited repertoire of translation factors, mostly translation repressors. These foci, termed processing bodies (P bodies), represent complexes where mRNA degradation can occur, since mRNA decay intermediates
A different kind of cytoplasmic foci, referred to as stress granules, have been discovered in higher eukaryotes under various stress conditions (recently reviewed in
Most published works to date have used fluorescent microscopy to study P bodies or stress granules. This technology allows the detection of large complexes whose fluorescence is above that of the background, but small P bodies might escape detection. Interestingly, though, Aragon et al.
It has been contended both for yeast
OPM allows us to model the system under study—the mRNA lifecycle—at various hierarchically arranged levels of detail. We started modeling only established knowledge concerning the mRNA lifecycle.
In the cytoplasm, mRNA can be located in various complexes, such as the ribosome or P body. It is quite possible that the yeast cytoplasm contains other large bodies that accommodate mRNA. However, as such bodies have not been reported, the only two cytoplasmic mRNA locations in our model are the ribosome (including also poly-ribosome) and P bodies. It has not been established whether mRNA can move in the cytoplasmic matrix, unattached to any complex. Trying to include this option in our initial model rendered the model more complex with no tangible benefit. Following Occum's Razor, we therefore assumed the simplest option, i.e., that mRNA does not reside in the cytoplasmic matrix as an unbound molecule. Thus,
The model leaves only two options: either ribosome or P body. Although ribosomes have been implicated as the mRNA acceptor in the cytoplasm
(A) Three mRNA transport processes, marked in cyan, have been added to the OPD of
In
(A) Zooming into the Translation process exposes its three subprocesses that follow P Body-Ribosome Transport (initiation): Elongation, Termination, and Protein Cleavage&Releasing. The corresponding factors involved as instruments in these processes are also shown linked to the processes with an instrument link (a line ending with a circle at the process end). The object eEF Set is the instrument for the Elongation process, while eRF Set with its members, the factors eRF1 and eRF3, is the instrument for the Protein Cleavage&Releasing process. Our conjecture is that eRF3 is the factor which is also involved as instrument for the Ribosome-P body Transport process. Since there is no proof for this as yet, the instrument link from eRF3 to Ribosome-P body Transport is colored red, denoting uncertainty. (B) The OPL text of the OPD in (A). Note that the word requires in the OPL sentence Ribosome-P body Transport requires eRF3. denotes the same uncertainty regarding the role of eRF3 as instrument to the Ribosome-P body Transport process, analogous to the instrument link from eRF3 to Ribosome-P body Transport in (A).
eRF3 is an instrument for translation termination (see
This prediction prompted us to examine whether eRF3 is a component of P bodies.
Some P bodies are indicated by arrows (different kinds of arrows point at different P bodies). P bodies were examined after 7 days starvation in stationary phase as described in (Lotan et al., 2005). Merging of the green (eRF3-GFP) and the red (Dcp2p-RFP or Rpb4p-RFP, as indicated) channels was done using PhotoShop software. Note that if the green and red foci colocalize the resulting foci are yellow.
Continuing our OPM-based conceptual modeling process, Post-ribosomal Processing from
(A) Zooming into Post-ribosomal Processing exposes its subprocesses Checkpoint, Storing, and Degradation, as well as the objects D-factor and Decision. As before, red indicates uncertainty or hypothesis: We propose that D-factor is the instrument for the process we call Checkpoint, which in turn, determines whether to store, degrade, or unload the mRNA for reuse. Green links denote an uncertain conjecture that was confirmed in this work by our experiments. Here, the structural link from P Body to eRF3 and the tag contains along it are green, denoting that we demonstrated experimentally our model-based conjecture that P-body contains eRF3. (B) The OPL text of the OPD in (A). Note the red color of the words in the sentences Checkpoint requires D-factor. and in Checkpoint yields Decision. The red denotes that we are not certain whether Checkpoint and D-factor exist, and if so whether Checkpoint requires D-factor. On the other hand, the green color of the word contains in the sentence P-body contains eRF3 indicates our success at experimentally proving our model-based conjecture that P-body contains eRF3.
Having presented the modeling process for specific portions of the mRNA lifecycle in some detail, we now outline general guidelines for OPM-based conceptual modeling of a biological system. At the molecular level, all biological systems obey relatively few common rules. Molecules can interact with other molecules, change their molecular environment if they have enzymatic activity, or affect localization within the cell. All these activities are included in our mRNA lifecycle model, which can therefore be potentially used as a case in point for modeling other subsystems of increasing complexity, and ultimately, the entire cell.
Using the OPCAT modeling environment (downloadable from
In this work, we have applied Object-Process Methodology to create a conceptual model of the mRNA lifecycle. This relatively simple model demonstrates the usefulness of conceptual modeling not just for understanding and communicating the structure and behavior of cell-level systems, but also for provoking conjectures and triggering ideas for experiments that can confirm or refute such conjectures. The findings update the model, which keeps evolving by repeating this cycle. Our model includes basic, well-known aspects of the mRNA lifecycle as well as recently discovered features. We modeled deliberately both established and less established knowledge in order to demonstrate that in both cases the model generates useful predictions. The relatively new feature of the mRNA lifecycle that we focused on here is the capacity of mRNA-protein (RNP) molecules to bundle together and form large cytoplasmic complexes, termed P bodies. We hypothesized that the purported regulation of P body biology is governed by D-Factor (which may be composed of several distinct components). This hypothetical D-Factor controls the fate of mRNA by “deciding” whether each mRNA is degraded, transported to the ribosome for reuse, or sequestered in the P body.
Recently, we proposed that P bodies are heterogeneous complexes
OPM has rich and flexible modeling capabilities with high expressive power. For example, as
Modeling other types of variants, e.g., protein phosphorylation, protein ubiquitination, RNA editing, can be done in a similar manner.
We emphasize that the OPM conceptual model of the mRNA lifecycle presented in this work is by no means complete. It illustrates superficially certain parts of this extremely complex cycle (which can be referred to as a pathway, if nucleotide recycling is ignored) as we conceive it today. Still, our model was detailed enough to provoke research questions, and a solution for one of those questions was found experimentally. The mRNA lifecycle is studied extensively, so any sections of this model can be in-zoomed further to provide ever more detailed descriptions.
In general, systems biology can benefit from using OPM as a generic framework for knowledge capture and representation, particularly since it enables balanced and unified representation of the system's structure and behavior using both objects and processes in the same diagram. Using the refinement-abstraction mechanisms that are built into OPM, the system under study can be clearly understood and communicated at various detail (granularity) levels. Moreover, as demonstrated here, OPM-based modeling provokes consideration of links missing from the process chain and stimulates ideas for experiments to prove or disprove new theories. Without the intellectual activity underlying conceptual modeling, such gaps or inconsistencies in the model can easily go unnoticed, evading the researchers' attention. Indeed, while engaged in modeling, we encountered portions of the system which we were uncertain how to model. The knowledge gaps become more apparent as we tried to further zoom (drill down) into specific subprocesses of the mRNA lifecycle. Graphically, this was manifested by increasing red color in the diagrams. Based on our positive experience, we propose that the friendly, yet formal, OPM modeling framework is a tool for modeling biological systems, whose adoption would benefit the emerging domain of systems biology.
Due to the ability to forge a holistic conceptual model of a biological system, we maintain that this research should be valuable to biologists and computer scientists who work on developing a Systems Biology understanding of the living organisms. The OPM-based conceptual modeling framework provides a clear, unambiguous way to describe our current knowledge of the state of the cell. It allows incremental resolution of different parts of the cellular machinery, helps pinpoint areas where our current understanding of the model is lacking, and finally, what type of experiments we may conduct to improve our understanding.
Following intensive future research and development, we envision an evolving model that would be developed, maintained, and updated constantly by the research community at large. When incorporating new data into the existing model, it would be possible to determine if it is consistent with previous knowledge. Unproven conjectures could also be incorporated into this evolving model, tagged as uncertain. Their inclusion in the model would help point towards evidence required to support these hypotheses. Such a modeling framework would help researchers tackling the huge challenge of understanding holistically the intricacies of the living cell.
A Quick Guide to the Syntax and Semantics of the Object-Process Methodology (OPM) Language.
(1.06 MB DOC)
Click here for additional data file.
We thank Roy Parker for RFP DCP2-RFP plasmid.