Present addresses: A. Xayaphoummine, Université Paris 7, UMR 7056, 2, place Jussieu, 75251 Paris cedex 05, France
V. Viasnoff, ESPCI, 10 rue Vauquelin, 75005 Paris, France
S. Harlepp, IPCMS–GONLO, 23 rue du Loess BP 43, 67034 Strasbourg cedex 2, France
RNA co-transcriptional folding has long been suspected to play an active role in helping proper native folding of ribozymes and structured regulatory motifs in mRNA untranslated regions (UTRs). Yet, the underlying mechanisms and coding requirements for efficient co-transcriptional folding remain unclear. Traditional approaches have intrinsic limitations to dissect RNA folding paths, as they rely on sequence mutations or circular permutations that typically perturb both RNA folding paths and equilibrium structures. Here, we show that exploiting sequence symmetries instead of mutations can circumvent this problem by essentially decoupling folding paths from equilibrium structures of designed RNA sequences. Using bistable RNA switches with symmetrical helices conserved under sequence reversal, we demonstrate experimentally that native and transiently formed helices can guide efficient co-transcriptional folding into either long-lived structure of these RNA switches. Their folding path is controlled by the order of helix nucleations and subsequent exchanges during transcription, and may also be redirected by transient antisense interactions. Hence, transient intra- and inter-molecular base pair interactions can effectively regulate the folding of nascent RNA molecules into different native structures, provided limited coding requirements, as discussed from an information theory perspective. This constitutive coupling between RNA synthesis and RNA folding regulation may have enabled the early emergence of autonomous RNA-based regulation networks.
RNA molecules exhibit a wide range of functions from essential components of the transcription/translation machinery (
It has long been proposed (
Several inspiring reports have demonstrated the importance of co-transcriptional folding (
To circumvent these limitations, we propose to use artificial RNA switches, presumably void of biological functions, and investigate how to efficiently encode their folding paths by exploiting simple sequence symmetries, instead of extensive (and possibly non-conclusive) mutation studies. Beyond specific examples of natural or designed RNA sequences, we aim at delineating general mechanisms and coding requirements for efficient co-transcriptional folding paths.
In a nutshell, we have designed a pair of synthetic RNA switches sharing strong sequence symmetries, so that both molecules partition, at equilibrium, into equivalent branched and rod-like nested structures with nearly the same free energy. Yet, in spite of this structural equivalence between the two RNA switches at equilibrium, we demonstrate that their folding path can be encoded to guide the first RNA switch exclusively into the branched structure, while the other switch adopts instead the rod-like nested structure by the end of transcription. This shows that folding paths do not simply result from the sequential formation of native helices in their order of appearance during transcription (i.e. sequential folding, see Discussion). Instead, efficient folding paths rely on the relative stability between native and non-native helices together with their precise positional order along the 5′–3′ oriented sequence (i.e. encoded co-transcriptional folding). Furthermore, we show that efficient folding path can be redirected through transient antisense interaction during transcription, suggesting an intrinsic and possibly ancestral coupling between RNA synthesis and folding regulation.
RNA switches with encoded folding paths depicted in
Encoded co-transcriptional folding path of a bistable RNA switch. (
Opposite co-transcriptional folding paths of a pair of RNA switches with ‘direct’ and ‘reverse’ sequences (i.e. 5′-ABCD-3′ versus 5′-DCBA-3′). Structures 1D and 1R (respectively, 2D and 2R) of the direct and reverse switches are energetically equivalent because of helix symmetries; dashed lines indicate mirror symmetry of Pa, Pb, Pc and Pd which are therefore conserved under sequence reversal relating direct and reverse switches. Despite these strong similarities between D and R structures at equilibrium, direct and reverse switches display ‘opposite’ co-transcriptional folding paths (direct switch into structure 1D and reverse switch into structure 2R) guided through a helix encoded persistence (left) or exchange (right) during
Correspondence between branched versus rod-like structure and migrating bands. A single mutation U38/C38 on the reverse sequence, Ru/c (see blue u/c mutation in Figure 2) unambiguously demonstrates the correspondence between the stabilized branched structure and the lower band on the gel (see text).
Sequences were inserted into pUC19 plasmid (between KpnI and BamH1 restriction sites) using enzyme removal kits (Qiagen) and cloned into calcium competent
Influence of temperature and transient antisense interactions on co-transcriptional folding. Equilibrium and native structures of reverse switch (R) with
This section is organized into two complementary subsections. The first one is primarily experimental and demonstrates, using sequence symmetries, the basis for encoding efficient folding paths with a pair of ‘symmetrically equivalent’ RNA switches adopting either their branched or rod-like structure by the end of transcription. The second subsection is theoretical and discuss, from an information content perspective (
We decided to investigate the basic mechanisms and coding requirements for efficient RNA folding paths with a stringent test case. Following the RNA switch design depicted on
In practice, however, 5′–3′ versus 3′–5′ folding paths cannot be probed on the same RNA sequence, as there is no RNA polymerase known to perform transcription in ‘opposite’ (3′–5′) direction. Hence, instead of studying a single RNA sequence, we have actually used a pair of RNA switches with exactly opposite sequences, i.e. 5′-ABCD-3′ and 5′-DCBA-3′ (see Materials and Methods). It is important to note that, in general, such pairs of RNA molecules do not adopt related structures at equilibrium, due to the large asymmetry between free energies of stacking base pairs with reversed orientation (e.g. 5′-GC/GC-3′ ≃ −3.4 kcal/mol and 3′-GC/GC-5′ ≡ 5′-CG/CG-3′ ≃−2.4 kcal/mol). For this reason, the pair of direct (D) and reverse (R) RNA switches, we have designed (
These results strongly support the co-transcriptional folding principles depicted on
Overall, this demonstrates that the competition between native and non-native helices can lead to efficient co-transcriptional folding paths of RNA switches independently from their actual equilibrium structures.
Moreover, we found that the folding path of the reverse switch could be significantly redirected toward structure 1R (≃50%) through transient antisense interactions, (
Antisense regulation of co-transcriptional folding paths. Interpretation of the encoded (left) and redirected (right) co-transcriptional folding paths of the reverse switch (Figure 4). This is based on simulations performed using the kinefold server (
Hence, transient intra- and inter-molecular base pair interactions can efficiently regulate the folding of nascent RNA molecules between alternative long-lived native structures, irrespective of their actual thermodynamic stability. Indeed, once formed, the co-transcriptional structures 1D and 2R remain trapped out-of-equilibrium for more than a day at room temperature (data not shown) demonstrating that these RNA switches can reliably store information on physiological time scales with their co-transcriptionally folded structures. In another context, the ability to control folding between distinct long-lived structures of nucleic acids using electrical (
Although our conclusions are based on particular examples of synthetic RNA switches related by sequence reversal and helix symmetries, we want to stress that these strong symmetry constraints are solely instrumental in demonstrating the possible independence between encoded folding paths and low-free energy RNA structures. These symmetries are not directly used nor necessary to achieve efficient co-transcriptional folding. On the contrary, imposing such strong sequence symmetries greatly limits the additional ‘information content’ that can possibly be encoded on the sequence. In the next subsection, we discuss how this use of sequence symmetries can actually be formalized to provide quantitative estimates on the minimum coding requirement for selective folding paths of generic RNA swiches.
In this subsection, we discuss how sequence symmetries can be utilized to estimate necessary base pairing conditions to encode efficient co-transcriptional folding paths. This requires, however, to reformulate base pairing conditions from an information content perspective, following the approach developped for biomolecular sequences (
In the following, we first establish a simple conservation law for information content. We then argue that upperbounds for the coding requirement of selective folding paths (or other molecular features) can be estimated by restricting the available coding space with strong sequence symmetries. Ultimately, upperbounds on coding requirements are related to the likelihood that a particular feature might arise from natural or
Let us first recall what the information content of a biomolecule is, before showing how it can actually be estimated for designed RNA switches using sequence symmetries.
The information content
With these crude initial assumptions, the information content of a short RNA sequence adopting a unique stable secondary structure can be estimated as
Similar estimates can be made including wobble base pairs (GU and UG) in addition to Watson–Crick base pairs (GC, CG, AU and UA). In that case, the available sequence entropy becomes
The previous coding requirement estimates demonstrate that structural information
This can be applied to estimate the minimum information that might be required to obtain two efficient opposite folding paths from the generic bistable RNA switch sequence of
This limited coding requirement concerning overlapping base pairs reinforces,
Hence, if all non-functional sequence symmetries are lifted, we expect that selective folding paths can indeed be readily achieved for a wide class of RNA sequences, as they require little encoded information beyond small asymmetries between alternative helices to guide or prevent their successive exchanges during transcription. Interestingly, this pivotal role of a few unpaired or transiently paired bases for efficient folding paths is also observed for other encoded molecular functions of RNAs. For instance, a few unpaired conserved bases usually prove essential for ribozyme functions or
Although many convincing reports have shown the importance of co-transcriptional folding (
In particular, studying the equilibrium folds of increasingly longer 3′-truncated transcripts has been argued to miss important out-of-equilibrium intermediates on the folding path of full-length molecules (
Simple sequential folding of a bistable RNA switch under sequence reversal and circular permutation. (
In contrast, our results demonstrate that folding paths can efficiently guide RNA transcripts into distinct alternative structures even when competing branched-like conformations exist and could, in principle, form during transcription. This competition between local overlapping helices and even global alternative structures is, in fact, ubiquitous to the folding dynamics and thermodynamics of RNA molecules. For instance, co-transcriptional folding has long been known to induce structural rearrangements as the nascent RNA chain is being transcribed (
The present study, following earlier stochastic folding simulations reported in (
Non-coding RNAs typically tolerate a significant number of neutral mutations and co-variations in their sequence, which presumably facilitates their continuous adaptation to environmental changes. From an information content perspective, this tolerance to (concerted) mutations also suggests that (partly) unconstrainted nucleotides may be used to encode other alternative structures and functions on the same RNA sequence, a feature which might have favored the emergence of new functional RNAs and RNA switches in the course of evolution (
In this study, we showed that co-transcriptional folding can efficiently guide RNA folding either toward branched structures (as for the ‘direct’ switch,
Finally, these results suggest that efficient folding pathways might have easily emerged and continuously adapted in the course of evolution the same way functional native structures have done so through mutation drift in sequence space; non-deleterious mutations are mostly neutral and conserve sequence folds and activity, while new functions may occasionally arise by rare hopping between intersecting networks of neutral mutations (neutral networks) (
We thank H. Putzer and L. Hirschbein for critical reading of the manuscript and R. Breaker, D. Chatenay, B. Masquida, K. Pleij, P. Schuster, J. Robert and S. Woodson for discussions. We acknowledge support from CNRS, Institut Curie, Ministère de la Recherche (ACI grant no. DRAB04/117) and HFSP grant (no. RGP36/2005). Funding to pay the Open Access publication charges for this article was provided by Institut Curie.