- Journal List
- HHS Author Manuscripts
- PMC4256716

Multiple intermediates on the energy landscape of a 15-HEAT-repeat protein
Maksym Tsytlonok
aMRC Cancer Cell Unit, Hutchison/MRC Research Centre, Hills Road, Cambridge CB2 0XZ, UK
bUniversity of Cambridge Department of Chemistry, Lensfield Road, Cambridge CB2 1EW, UK
Patricio O. Craig
cDepartment of Chemistry and Biochemistry, and Center for Theoretical Biological Physics (CTBP), University of California San Diego (UCSD), 9500 Gilman Drive, La Jolla, CA 92093-0374, USA
Elin Sivertsson
bUniversity of Cambridge Department of Chemistry, Lensfield Road, Cambridge CB2 1EW, UK
David Serquera
aMRC Cancer Cell Unit, Hutchison/MRC Research Centre, Hills Road, Cambridge CB2 0XZ, UK
Sarah Perrett
dNational Laboratory of Biomacromolecules, Institute of Biophysics, Chinese Academy of Sciences, 15 Datun Road, Chaoyang District, Beijing 100101, China
Robert B. Best
bUniversity of Cambridge Department of Chemistry, Lensfield Road, Cambridge CB2 1EW, UK
Peter G. Wolynes
eCenter for Theoretical Biological Physics & Department of Chemistry, Rice University, Houston, Texas, USA
Laura S. Itzhaki
bUniversity of Cambridge Department of Chemistry, Lensfield Road, Cambridge CB2 1EW, UK
Associated Data
Abstract
Repeat proteins are a special class of modular, non-globular proteins composed of small structural motifs arrayed to form elongated architectures and stabilised solely by short-range contacts. We find a remarkable complexity in the unfolding of the large HEAT repeat protein PR65/A. In contrast to what has been seen for small repeat proteins in which unfolding propagates from one end, the HEAT array of PR65/A ruptures at multiple distant sites, leading to intermediate states with non-contiguous folded subdomains. Kinetic analysis allows us to define a network of intermediates and to delineate the pathways that connect them. There is a dominant sequence of unfolding, reflecting a non-uniform distribution of stability across the repeat array; however the unfolding of certain intermediates is competitive, leading to parallel pathways. Theoretical models accounting for the heterogeneous contact density in the folded structure are able to rationalize the variation in stability across the array. This variation in stability also suggests how folding may direct function in a large repeat protein: The stability distribution enables certain regions to present rigid motifs for molecular recognition while affording others flexibility to broaden the search area as in a fly-casting mechanism. Thus PR65/A uses the two ends of the repeat array to bind diverse partners and thereby coordinate the dephosphorylation of many different substrates and of multiple sites within hyperphosphorylated substrates.
Introduction
Proteins may fold and unfold repeatedly during their life in the cell. For example, partly folded states are populated both at the start of a protein’s life (co-translational folding, intracellular trafficking) and at the end (proteasomal degradation); in between, small and large fluctuations of the protein are required to carry out functions such as signal transduction, allostery, catalysis and mechanical work. The relative stabilities of the different conformations must be tightly controlled to prevent the population of problematic species that lead to malfunction or misfolding. This multitude of states characterizes the so-called “energy landscape” of a protein. Experimental folding studies have, however, classically focused on small, globular proteins that can only provide information about one or two species within the landscape. In order to access more of the landscape one must look at larger proteins, but such data are often not easy to interpret and therefore new insights remain elusive. To resolve this dilemma we have turned to a special class of proteins known as tandem repeat proteins. They are composed of small motifs (20–40 amino acids), arrays of which pack in a roughly linear fashion to produce elongated, super-helical architectures (1). Their structures comprise only short-range interactions between residues within a repeat or in adjacent repeats and as such they are distinct from the topologically complex structures of globular proteins. Herein we exploit this simplicity and modularity to decipher the folding of a very large repeat protein, thereby enabling us to visualise a broad energy landscape to an unprecedented degree.
The smaller repeat motifs (e.g. ankyrin and tetratricopeptide) have been the main focus of study to date (2–11). The 39-residue, α-helical HEAT motif (huntingtin, elongation factor 3, PP2A subunit and the lipid kinase TOR) is larger than these and it is usually found in large repeat arrays. It is thought that the absence of sequence-distant contacts affords repeat proteins inherent flexibility, and crystallographic data and molecular dynamics simulations hint at how such a property may be employed in the function of large repeat proteins like HEAT repeats (12, 13). However, experimental information on the stability and dynamics of these proteins has to date been lacking. The 590-residue protein PR65/A, made up of 15 tandem HEAT repeats, is the scaffolding subunit of the heterotrimeric serine/threonine protein phosphatase 2A (PP2A) (Fig. 1). PP2A regulates many cellular signaling pathways (reviewed in (14–16)), and this diversity derives from the fact that the PP2A enzyme constitutes over 200 complexes assembled from different combinations of the catalytic C subunit, the PR65/A scaffolding subunit and one of a large number of regulatory B subunits. PP2A is a tumor suppressor and its function is subverted in some cancers by viral oncoproteins that bind the core A–C heterodimer, thereby inhibiting activity and altering substrate specificity (17–20). The regular packing of the HEAT repeats of PR65/A is interrupted by a 22° rotation between HEAT 3 and HEAT 4 and a 45° rotation between HEAT 12 and HEAT 13 (21–23). The catalytic subunit C binds to the C-terminal HEAT repeats 11–15 of PR65/A, whereas diverse regulatory B subunits bind to the N-terminal HEAT repeats 1–10 (24). Binding alters the packing of the N- and C-termini (HEAT 1–3 and HEAT 12–15) relative to the middle of the protein and thereby brings the B–bound substrate into close proximity with the active site of the C subunit (21, 25).
(A) The HEAT-repeat scaffolding subunit PR65/A is in green, the catalytic subunit is in yellow and the regulatory subunit B55 is in pink (PDB 3DW8, image generated using PyMOL (http://www.pymol.org). (b) Structure of PR65/A with the mutations introduced in this study highlighted in blue.
Here we have elucidated the folding mechanism of PR65/A using a combination of experimental and computational approaches. We show that the HEAT array ruptures at multiple sites, rather than unfolding propagating from one end of the array as observed for small repeat proteins, leading to an equilibrium intermediate consisting of folded subdomains (HEAT 1–2 and HEAT 11–13) that are non-contiguous in the sequence. Kinetic analysis enables us to delineate a network of partly folded intermediates and to define both sequential and parallel pathways that connect them. The folding/unfolding of the repeat protein architecture has been likened to driving along a highway, and the unfolding of PR65/A nicely illustrates this analogy: the reaction can only propagate in a one-dimensional manner and is therefore very sensitive to retardation by intermediate species akin to a traffic jam; the longer is the highway, the greater the potential for traffic to slow progress (26, 27). We can destabilize these intermediates by truncation or mutation, smoothing the energy landscape sufficiently in one construct that it unfolds in a single, rapid step. Folding and function may be coupled in the HEAT array of PR65/A, and Wolynes and colleagues have used a fly-casting analogy to describe a similar type of behavior in other proteins (28, 29). Thus, HEAT 1–2 and HEAT 11–13 constitute the rigid core of PR65/A that is important for the recognition of partner proteins, whereas the less stable HEAT 3–10 and HEAT 14–15 provide the flexibility with which to widen the search area, coordinating the binding of PP2A regulatory and catalytic subunits at opposite ends of the HEAT array and thereby the timely dephosphorylation of many different substrates. The fluctuations of PR65/A could also facilitate the dephosphorylation of multiple sites within hyperphosphorylated PP2A substrates like securin and Tau (13, 24).
Results
PR65/A unfolds at equilibrium via a hyper-fluorescent intermediate
PR65/A has five tryptophan residues at positions 140 (HEAT 4), 257 (HEAT 7), 417 (HEAT 11), 450 (HEAT 12) and 477 (HEAT 13). The protein unfolds reversibly in a range of different buffers, as monitored by fluorescence. Using a buffer of 50 mM MES at pH 6.5 and a temperature of 25 °C we obtained a denaturation profile in which two transitions are well resolved (Fig. S1) (Table S1). There is no change in the profile after different incubation times, and the unfolding and refolding denaturation curves superimpose yielding the same midpoints and m-values within error, indicating reversibility (Figs. S1A, S1B). The far-UV CD-monitored denaturation curve also shows two transitions (Fig. S1D). An intermediate is populated between ~2 M urea and 4 M urea that has higher fluorescence intensity than have either the native or fully denatured states. The CD spectrum of the intermediate at 2.2 M urea has ~40% of the native ellipticity and, like the native state, it displays the double minima characteristic of α-;helical structure. These data are consistent with approximately five or six HEAT repeats being structured in the intermediate. Analysis by gel filtration and native PAGE shows that the intermediate oligomerises (see SI Text; Figs. S2–S4).
Concomitant rupture of the HEAT repeat array at multiple sites under denaturing conditions
Mutations were made of the conserved valine residue at position 24 of the HEAT motifs and of non-conserved valine residues at other positions (Table S2). The remaining mutations, S142V, Q169P, S445P and H496V are to consensus residues, these being natural substitutions that ought to be accommodated by the structure of the repeat. Mutations in HEAT repeats 3–7 and HEAT repeats 14–15 (Figs. 2A,B; Table S2) shift the midpoint of the first unfolding transition to lower urea concentrations, indicating destabilization of the native state relative to the intermediate. Mutations in HEAT repeats 8–10 have little or no effect on the first unfolding transition when monitored by fluorescence; however, destabilisation of the native state relative to the intermediate is observed when unfolding of these mutants is monitored by CD (Fig. 2C; Table S3). The lack of fluorescence change associated with unfolding of HEAT repeats 8–10 is likely due to the absence of tryptophan residues in these repeats. Mutations in HEAT repeats 1–2 and 11–13 have little or no effect on the stability of the native state relative to the intermediate when measured by either fluorescence or CD (Fig. 2D; Tables S2 and S3). The simplest interpretation of the results is that the first transition is associated with the unfolding of HEAT 3–10 and HEAT 14–15, and the second transition is associated with the unfolding of HEAT 1–2 and HEAT 11–13, i.e. HEAT 1–2 and HEAT 11–13 are structured in the hyper-fluorescent intermediate and the other repeats are unstructured.
Denaturation curves of the mutants are in red and wild type is in blue. Effects of mutations in HEAT repeats 3–7 and 14–15 are shown measured by fluorescence (A) and by CD (B); for simplicity, only the CD data for the first unfolding transition are shown. Effects of mutations in HEAT repeats 8–10 are shown in (C). Effects of mutations in HEAT repeats 1–2 and 11–13 are shown in (D).
Truncations map the stability distribution across the repeat array and define subdomains
The one-dimensional, modular nature of repeat protein architectures allows the use of truncations to interrogate the structure and provide a complement to site-directed mutagenesis (4, 30, 31). We truncated the protein at the end of each repeat or in the middle of repeats in those cases where the truncation at the end of a repeat was not expressed in a soluble form. Several truncated variants (referred to as HEAT 1–13, HEAT 1–10, HEAT 1–9, HEAT 1–8, HEAT 1–5, HEAT 4–15, HEAT 5–15) were not soluble, and two variants (HEAT 1–6 and HEAT 2–15) were unstable and oligomerised under native conditions. A total of eight truncated variants (HEAT 1–14, HEAT 1–12, HEAT 1–11, HEAT 1–9.5, HEAT 1–8.5, HEAT 1–7, HEAT 3–15, HEAT 3–9.5) were monomeric and were analysed further (Table S4). Although the ellipticity signal at 222 nm could potentially provide information about how much native helical structure is maintained upon truncation, in the case of PR65/A the five tryptophan residues appear to contribute significantly to this ellipticity (as shown by our observation that mutations of the tryptophan residues in full-length PR65/A cause large changes in the 222 nm ellipticity (Fig. S5), and these changes are not the result of partial unfolding in the tryptophan mutants as gel-filtration shows that they are fully folded). It is therefore impossible to correlate the magnitude of the 222 nm ellipticity of the fragments, which contain varying numbers of tryptophan residues, with their helical content. Instead, urea-induced unfolding of the variants was used to compare stability and structural integrity across the whole fragment set. A pattern emerges from the CD-monitored denaturation profiles (Fig. 3A). Deletion of the most C-terminal repeat (HEAT 1–14) results in a large shift in the first unfolding transition to lower urea concentrations and a large decrease in the apparent m-value; upon further truncation (HEAT 1–12 and HEAT 1–11) this transition is shifted back to higher urea concentrations; a similar pattern is evident in the fluorescence-monitored denaturation profiles (Fig. 3B; Table S5). Deletion of HEAT 13 and HEAT 12 also results in a large shift to lower urea concentrations in the second unfolding transition monitored by fluorescence and less oligomerisation of the intermediate as judged by native PAGE (Fig. 3C); moreover, no oligomerisation is observed for constructs comprising fewer than ten repeats. These results are consistent with the site-directed mutagenesis data, and they show that HEAT repeats 11–13 are required for the stability of the oligomeric intermediate. A second transition in the CD-monitored denaturation profile is evident for HEAT 1–9.5, HEAT 1–8.5 and HEAT 1–7 and absent for HEAT 3–9.5, and therefore it most likely corresponds to the unfolding of HEAT 1–2. The CD-monitored denaturation profile does not change greatly between HEAT 1–9.5, HEAT 1–8.5 and HEAT 1–7. The CD- and fluorescence-monitored denaturation data of these three variants together suggest that HEAT repeats 8–9.5 are weakly structured in the absence of any C-terminal repeats.
Denaturation curves of truncated variants measured by CD (A) and fluorescence (B). The truncations are in red and the full-length protein is in blue. The variant HEAT 3–9.5, which is both N- and C-terminally truncated is shown in yellow with HEAT 1–9.5 in red for comparison. The protein concentration used was between 0.5 µM and 3 µM, depending on the sample, and the data were normalized to 1 µM protein concentration. (C) Native gels of wild type and truncated variants. Samples were incubated in 0–8 M urea and then diluted with a loading dye, resulting in final urea concentrations of 0–4 M urea.
The truncations are consistent with the model for the unfolding of full-length PR65/A inferred from the mutants, in which the hyper-fluorescent intermediate has structured HEAT 1–2 and HEAT 11–13. The truncations help us to further delineate the subdomains with the HEAT-repeat array and provide information about the relative stabilities of these subdomains. Deletion of the N-terminal two repeats (variant HEAT 3–15) shifts the first transition to a much lower urea concentration, indicating that HEAT 1–2 acts as a cap for the HEAT 3–10 moiety. The variant HEAT 1–13 was insoluble, suggesting that HEAT 14–15 acts as a cap for the HEAT 11–13 moiety; however, the mutational data both at equilibrium and under kinetic conditions (see below) indicate that HEAT 11–13 remains folded when HEAT 14–15 is unfolded, suggesting that HEAT 11–13 is an independently stable moiety and that the HEAT 14–15 cap acts to protect the HEAT 11–13 moiety from oligomerisation rather than stabilising it against unfolding. Deletion of HEAT 15 destabilises the C-terminal moiety resulting in it unfolding at a much lower urea concentration in the variant HEAT 1–14 than in the full-length protein. Further deletions, of HEAT 13 and HEAT 12, destabilise the HEAT 11–13 moiety causing this moiety to unfold at very low urea concentrations in variants HEAT 1–12 and HEAT 1–11.
Additional intermediates revealed in the unfolding kinetics
The unfolding kinetic traces are characterized by three phases at urea concentrations above 5 M: a fast phase associated with a decrease in fluorescence, a slower phase associated with an increase in fluorescence and a very slow phase associated with a decrease in fluorescence (Fig. S6A). The fast and slower phases are of similar amplitudes and the slowest phase is much smaller in amplitude. At urea concentrations below 5 M the slowest phase is not present. However, the traces still fit to the sum of three exponential phases because a different slow phase appears, one that is associated with an increase in fluorescence (Fig. S6B); this phase is small in amplitude. The unfolding traces obtained by stopped-flow far-UV CD fit to a single exponential phase at higher urea concentrations and to the sum of two exponential phases at lower urea concentrations; the rate constants for these phases are similar to those of the fast and slower phases measured by fluorescence (at higher urea concentrations these two rate constants are very similar and they can only be distinguished in the fluorescence experiments because they have opposite signs). The fast phase is larger in amplitude than the slower phase, which is consistent with the assignment of these phases to the unfolding of specific subdomains of different sizes that we make below.
The fluorescence-monitored refolding kinetic traces can be fitted to the sum of three exponential phases (Fig. S6C). There is a fast phase associated with an increase in fluorescence and two slower phases associated with decreasing fluorescence. Interestingly, the rate constants increase with increasing urea concentrations between 1.9 M and 2.3 M urea, in contrast to the regime below 1.9 M urea and above 2.3 M urea where they show the usual decrease upon increase in urea concentration (Fig. 4A). Non-linearity can be caused by transient aggregation; however, in the case of PR65/A the refolding rate constants are independent of protein concentration between 0.2 µM and 5 µM (Fig. S7); the unfolding kinetics is also independent of protein concentration (Fig. S7). Another explanation is that a misfolded species accumulates that needs to unfold in order to progress to the native state. Due to the complexity of the refolding kinetics and the potential for additional difficulties in interpreting the data arising from the large number of proline residues (26), our subsequent analysis focuses instead on the unfolding kinetics. The plot of the logarithm of the rate constants for unfolding versus urea concentration is shown in Figure 4A.
The urea dependence of the rate constants is shown for the wild type in blue and the mutants and truncated variant HEAT 3–15 in red. The fast phase is in squares, the slow phase is in triangles, the slowest phase between 3.5 M and 5.5 M urea is in circles and between 5.5 M and 8.5 M urea in inverted triangles
Unfolding of HEAT 3–10 occurs rapidly
Upward curvature is evident in the urea dependence of the rate constants of the fast unfolding phase (Fig. 4A), which suggests that there are parallel pathways accessible, as has been observed observed previously for the large ankyrin repeat protein D34 (36). The additional, downward curvature evident in this limb and in the limb of the slower phase can be interpreted as a denaturant-dependent shift between two (or more) sequential transition states separated by a high-energy intermediate (sequential barriers model) (32) or a gradual movement of the top of a broad energy barrier along the reaction coordinate in accordance with Hammond behavior (33) (34). Mutations in HEAT repeat 3 (V103A) and in HEAT repeats 8–10 (V283A, V333A, V340A, V390A) speed up the fast phase, suggesting that it corresponds to the unfolding of these repeats (Fig. 4C). The mutation V103A does not show upward curvature, suggesting that only one pathway is now accessible. The mutation V340A also shows no upward curvature, suggesting rerouting of unfolding via the other pathway. Since V283A, V333A and V390A are not very destabilizing, neither pathway is abrogated and therefore there is upward curvature for these mutants as for the wild type. Mutations in HEAT 4–7 (S142V, Q169P, V200A, V219A, V258A) have only small effects on the fast unfolding phase and no significant effect on the other unfolding phases (Fig. S8).
A second intermediate is formed subsequently by the unfolding of HEAT 14–15
Very destabilizing mutations in HEAT repeats 14–15 (V535A, V575A) have the most dramatic effect on the slower unfolding phase, which is associated with an increase in fluorescence, suggesting that this phase arises from unfolding of these elements of structure. Interestingly, however, these mutations also increase the rate of the fast phase described in the previous section (Fig. 4C). This result can be understood if the mutations accelerate the rate of unfolding of repeats 14–15 such that this process becomes competitive with unfolding of 3–10, and the unfolding kinetics of these two regions can no longer be considered as sequential steps. A simple scheme that can explain the data allows for 3–10 to unfold first, followed by 14–15, or vice versa. The rate for the fast phase is then the sum of the unfolding rates of 14–15 and 3–10, and can therefore be increased by strongly destabilizing mutations of 14–15. The dominant slow phase is still unfolding of 14–15. For V535A, the slow phase is completely absent (Fig. 4C); this is most likely because the mutation is so destabilizing that the initial state observed has repeats 14–15 already unfolded – because either they are already unfolded at low denaturant as suggested by the equilibrium fluorescence, or they unfold in the dead time of the instrument. A similar scenario is most likely responsible for the absence of the slow phase for V575A at higher urea concentrations. For V532A the slow phase is speeded up to a lesser extent and therefore can still be observed. Later we describe in detail how such a scheme quantitatively fits the kinetic traces for all the above mutants.
A truncated variant comprising HEAT 13–15 was made in order to simplify the unfolding kinetics; W477 provides a spectroscopic probe for this variant. The unfolding traces of HEAT 13–15 are monophasic and associated with an increase in fluorescence (Fig. S9A); this result is consistent with our assignment of the slow phase for full-length PR65/A to the unfolding of HEAT 14–15. These kinetic data are also consistent with the equilibrium data showing that HEAT 14–15 is unfolded in the hyper-fluorescent intermediate. Interestingly, V535A has an effect on the slow unfolding kinetics also: whereas for the wild type there is a slow unfolding phase below 5 M urea associated with an increase in fluorescence (corresponding to unfolding of HEAT 1–2, see below), this phase is absent in the mutant V535A and instead there is a slow phase associated with a decrease in fluorescence.
HEAT 1–2 unfolds to form a third intermediate
The biggest effects of mutations V25A and V44A (HEAT 1 and HEAT 2) were on the slowest unfolding phase, observed at urea concentrations below 5 M, which is associated with an increase in fluorescence (Fig. 4C). The rate constants for this phase are speeded up by these mutations. Moreover, the phase is absent in the N-terminally truncated variant, HEAT 3–15 (Fig. 4C). The results suggest that this phase corresponds to the unfolding of HEAT 1–2, which occurs separately from HEAT 3–10 below 5 M urea whereas HEAT 1–2 unfolds concomitantly with HEAT 3–10 above 5 M urea. Interestingly, V44A also speeds up the phase assigned to the unfolding of HEAT 14–15.
Unfolding of HEAT 11–13 is the slowest step
V435A (HEAT 11) and V454A (HEAT 12) speed up the slowest unfolding phase observed for the wild type above 5.5 M urea (Fig. 4C). The absence of this phase at urea concentrations below 5.5 M urea is consistent with it corresponding to unfolding of the HEAT 11–13 moiety: the equilibrium experiments show that this moiety is folded at lower urea concentrations. V435A also has an effect on the other phases; the slow phase, corresponding to unfolding of HEAT 1–2 and associated with increasing fluorescence, is absent in its kinetics.
Unfolding kinetics of truncated variants
The unfolding traces of C-terminally truncated variants HEAT 1–9.5, HEAT 1–8.5 and HEAT 1–7 (all of which have W140 and W257) show a decrease in fluorescence (Fig. S9B). The absence of the phase observed for full-length PR65/A associated with an increase in fluorescence suggests that the hyperfluorescence of the intermediate arises from the unfolding of repeats in the C-terminal moiety and it is consistent with the absence of a hyper-fluorescent intermediate for these truncations at equilibrium. The unfolding of HEAT 1–9.5 can be fitted to the sum of three exponential phases (Fig. S9B and Fig. 4B). There is a very fast phase that is faster than any of the phases observed for the full-length protein, a fast phase that has the same rate constants and urea dependence as the fast phase of the full-length protein, and also a slow phase. The slow phase is very low in amplitude (~5%). The very fast phase is the major phase at high urea concentrations whereas the fast phase is the major phase at lower urea concentrations. The unfolding of HEAT 1–8.5 can be fitted to the sum of two exponentials and shows the same very fast and fast phases as HEAT 1–9.5; the fast phase has very low amplitude and the slow phase observed for HEAT 1–9.5 is absent (Fig. 4B). The unfolding kinetics of HEAT 1–7 shows a single, very fast phase, which is similar to what is observed for HEAT 1–8.5 and HEAT 1–9.5 (Fig. 4B). The presence of a fast phase for variants HEAT 1–8.5 and HEAT 1–9.5, similar to that observed for full-length PR65/A, and its absence in the kinetics of HEAT 1–7, is consistent with the assignment of this phase to the unfolding of HEAT 8–10. We therefore propose the following to explain the appearance of the very fast unfolding phase in these three truncated variants and the presence at low amplitudes in HEAT 1–8.5 and HEAT 1–9.5 of the fast unfolding phase characteristic of full-length PR65/A: in the absence of HEAT 8–10, the HEAT 1–7 variant is able to unfold in a single very rapid step. For HEAT 1–8.5 and HEAT 1–9.5 there is a mixture of ‘native’ states, arising because the HEAT 8–10 moiety is truncated; the truncated moiety is unstructured in the major population and so unfolding proceeds as it does for HEAT 1–7, whereas the moiety is structured in the minor population and so unfolding proceeds as it does for the full-length protein.
Unfolding of HEAT 4–7 in a very fast unfolding step revealed by double-jump experiments
Two observations suggest that HEAT 4–7 unfolds after the unfolding of HEAT 8–10 in a very fast step, which, since it is faster than the previous step, is not detected in the kinetic measurements. First, mutations in HEAT repeats 4–7 do not have a clear effect on the any of the five kinetic unfolding phases observed for full-length PR65/A, whereas mutations elsewhere in the protein have clear effects that allow each phase to be assigned to the unfolding of a specific subdomain. Second, when repeats adjacent to HEAT 4–7 are deleted (i.e. the truncated variants HEAT 1–7, HEAT 1–8.5, HEAT 1–9), a very fast phase is observed that is absent in the unfolding of the full-length protein (Fig. 4). To test whether there is an additional phase that is missed in the single-jump method we performed a double-jump experiment in which unfolded protein was allowed to refold for a delay time of between 30 ms and 50 s before unfolding was initiated. The unfolding kinetics of this refolded protein was multiphasic and all but one of the phases had similar rate constants to those observed by single-jump unfolding of native protein (Fig. S10). The new phase is very fast and its amplitude first increases and then decreases with increasing refolding delay time, consistent with the accumulation and decay of an intermediate. We therefore propose that this very fast phase is the ‘missing’ phase that corresponds to the unfolding of HEAT 4–7.
Kinetic model for unfolding of PR65/A
The above mutation and truncation studies establish qualitatively that the unfolding of PR65/A can be understood in terms of a predominant order of unfolding events. We have used these data to construct a six-state phenomenological kinetic model that explains both the wild-type data and the data for mutants over a wide range of denaturant concentrations. The scheme is summarized in Fig. 5A and includes the native and unfolded proteins as well as four intermediates, labeled I–IV. The initial unfolding of the wild-type protein occurs mainly via unfolding of repeats 3–10 (rate k1), followed by unfolding of 14–15 (rate k2). However, an alternative possibility is that these two regions could unfold in the reverse order, and, as described below, these parallel pathways are able to explain several features of the kinetic data. The last two steps are unfolding of repeats 1–2 (rate k3), followed by 11–13 (rates k4, k−4). Unfolding of 11–13 has been treated as being reversible in the model, since it does not occur at equilibrium below ~5 M urea for the wild-type protein. To allow for downward curvature in the denaturant-dependence of the rate k1, it was modeled using a sequential barriers scheme in which unfolding from native to intermediate I occurs via a high energy intermediate X, i.e. k1 = knxkxi/(kxu+kxn) (32). This assumption is reasonable, as k1 describes the folding of the largest contiguous “block” of repeats (3–10), and a similar model has been used to describe downward curvature in the unfolding of other long repeat arrays (7, 35). Note that this extra complication is relatively unimportant, and is only necessary to describe the unfolding data below ~4 M urea concentration. The scheme has five kinetic eigenvalues corresponding to the rates k1, k2 and k3 and the sums k1+k2 and k4+k−4 (see SI for details).
(A) Kinetic scheme in which the native protein unfolds via a sequence of (parallel) intermediates to the unfolded state. Names of species are given in boxes and rate coefficients for their interconversion are shown next to the arrows. Color code refers to the region which is unfolding in each step. (B) Lower: Relaxation rates from kinetic model (lines) overlaid on the rates obtained by fitting of exponentials to raw data; color code corresponds to that in (A), as far as possible. Upper: Amplitudes corresponding to each relaxation. Positive amplitude corresponds to a decrease of fluorescence with time. (C) Left axes: Raw fluorescence traces (blue symbols) overlaid with the global fit from the kinetic model (red lines); Right axes: Time evolution of the populations of the species shown in (A). A range of concentrations are shown for wild-type in the left column, and in the right column are shown representative traces for five mutants chosen from the different structural regions.
The fluorescence signal was modeled as a sum over populations of each species, weighted by their relative fluorescence quantum yield. The model was globally fit to the raw fluorescence traces at all urea concentrations for the wild-type protein and several mutants: V25A representing repeats 1–2, V333A for repeats 3–10, V436A for repeats 11–13, and V532A and V575A for repeats 14–15. The fit parameters for the wild type were the rates at 4 M urea, their denaturant dependences (m-values) and the quantum yields for each species. The mutants shared all the parameters with wild type, with only the rate coefficient(s) pertaining to unfolding of the relevant region being allowed to vary. However, for V436A it was found to be necessary to optimize separately the relative quantum yields; in retrospect, this can easily be rationalized by a direct Van der Waals contact between V436 and W477 in the crystal structure, resulting in changes in environment and solvent exposure of that tryptophan on mutation. The final optimized parameters are summarized in Table S6.
The model formally has five relaxation time scales, which are plotted versus urea concentration in Fig. 5B (lines). In Fig. 5C are shown the raw fluorescence traces fitted to the model. For clarity, we will describe separately the situation at urea concentrations below and above 5 M. At low denaturant concentrations, the protein unfolds via a series of hypofluorescent intermediates (I,II,III) to the hyperfluorescent intermediate IV which is stable up to ~5M, resulting in a “dip” in the fluorescence with time. This can clearly be seen to be correlated with the population of the intermediates also plotted in Fig. 5C. There are three exponential phases evident in the raw data (the relaxation rates for these are plotted as symbols in Fig. 5B). In the model, the fastest (green) phase is associated with the formation of intermediates I and II with rate k1+k2, which has a large positive amplitude, i.e. causes a decrease in fluorescence (upper part of Fig. 5B). Decay of II occurs with a very similar rate k1≈ k1+k2 with small negative amplitude, but because of the similarity of time scales in the wild type, and the small population of intermediate II, it cannot be distinguished in the fluorescence traces. The second (magenta) exponential phase arises from decay of the populated intermediate I with rate k2, and has a significant negative amplitude. This is primarily responsible for the rise in fluorescence at the end of the “dip”. Finally, a third (cyan) exponential phase with very small negative amplitude comes from conversion of intermediate III to IV at long times.
Above 5 M urea, the fastest phase is still associated with the formation of intermediates I and II with rate k1+k2. However, upward curvature is seen in this phase, resulting from the fact that at (very) high urea concentration, k2>k1 and there is a switch of pathway toward that going via intermediate II. As a result of the larger k2, the second fastest phase now contains contributions from the decay of both intermediates I and II, since intermediate II is now significantly populated and the rates k1 and k2 are now more comparable. At very high denaturant, the second fastest phase overlays well with the rate k1 as this phase is now mostly associated with decay of intermediate II, which has become larger in population than I. The relaxation with rate k3 is still present, but less evident in the fluorescence traces, because above 5 M urea, intermediate IV decays further to the unfolded state with rate ~k4. The amplitude of this final relaxation is of course initially zero and only emerges with significant positive amplitude at urea concentrations above ~6 M (Fig. 5B, upper panel).
The fluorescence traces for the mutants can be well described by the same model, by changing only the rates associated with the unfolding of the relevant subdomain. The fits for representative urea concentrations are shown in Fig. 5C. For example, the V25A data can be reproduced by an increase of rate k3. The result is a faster depletion of intermediate III such that the fluorescence has almost reached a “plateau” value at long times. The V333A data are explained by an increase in k1 (via knx), such that intermediate I is populated more rapidly leading to a faster initial fluorescence decay and a larger overall “dip” in the total fluorescence. Mutation V436A results in a large increase in k4 and decrease in k−4. Consequently, the population of intermediates is much reduced at high denaturant concentration (being converted quickly to unfolded), and the overall signal appears almost single-exponential. Mutations V532A and V575A affect rate k2. Because of the parallel pathways, an increase in k2 can have multiple effects. The V532A results in the smaller increase of k2 and the primary effect of this is that intermediate I decays more rapidly, such that the rise in fluorescence occurs earlier and the overall signal amplitude is smaller. The effect of V575A on k2 is much larger, such that k2 becomes larger than k1. This also increases the rate of the initial fluorescence decay (with rate k1+k2), and results in intermediate II being the predominantly populated one. Because II has a larger fluorescence quantum yield than I, and decays with a fast rate (k1), the result is that the magnitude of the “dip” in fluorescence is very small, and signal to noise ratio is small.
Structure-based model simulations of PR65/A
The structure-based model simulations were next used to analyze the folding energy landscape of PR65/A. This type of model relies on the dominantly funneled nature of protein energy landscapes, achieved by evolution through the realization of consistent inter-residue interactions that largely favors the native topology compared to other alternative configurations (36). Pure structure based models take into account in the energy function solely the interactions found in the native structure. This energy function is then used in molecular dynamics simulations to characterize the tradeoff between the loss of configurational entropy and the gain of favorable contact interactions along folding trajectories, to characterize the folding mechanism and to identify the dominant pathways in a context where they can solely be influenced by topology. Many different flavors of structure based models that incorporate heterogeneity in the strength of inter-residue interactions or allow for non-native interactions have been successfully used in the past to characterize the existence of folding intermediates and the interpretation of folding pathways (37).
We used two types of coarse-grained models. One model involves a perfectly funneled energy function (Gō like model) with a homogeneous contact potential. The other model treats the short- and medium-range interactions with a heterogeneous native structure based potential, but allows for frustration and non-native contacts between residues distant in sequence (i-j≥12). These tertiary interactions involve a generalized statistical contact potential that depends on the spatial distance and sequence identity of the interacting residue pairs. It includes water-mediated interactions and a burial energy term. We call this model smAMW for single memory associative memory Hamiltonian with water-mediated interactions.
We first ran a series of short simulations at different temperatures to determine the Tf and then performed a more exhaustive sampling at a temperature near the Tf using umbrella sampling. The Tf observed was 0.9 and 1.0 in reduced temperature units for the pure structure-based and smAMW models, respectively. Figures 6A and 6B show the evolution of the equilibrium average fraction of native contacts (qic) for all residues of PR65/A as a function of the global reaction coordinate, Q, which represents the total fraction of native contacts for the whole protein. This type of plot gives an indication of which regions become folded at different depths in the funnel i.e. values of the reaction coordinate. For both models, at low values of Q the elements that have the highest number of native contacts formed correspond to HEAT 1 (residues 1–42), 11–12 (397–473) and half of HEAT 13 (residues 474–495) (Fig. 6C and 6D). These results suggest that these elements fold first in agreement with the experimental observations. The other elements fold at a later stage. The analysis also shows that at low values of Q there is also some partial structure in restricted regions of HEAT 5, HEAT 6 and HEAT 9, although lower than the amount found in the previous repeats. Although the heterogeneity of the contact energy and the influence of non-native contacts might modulate the order of folding of the different elements, the similarity in the patterns observed for the two models indicates that the native topology plays a dominant role in determining the folding mechanism of this protein.
The evolution of the fraction of native contacts for each residue (qic) as a function of the global reaction coordinate Q is shown in panels (A) and (B) for the Gō and smAMW models, respectively. The structure of the native state of PR65/A (pdb code: 1B3U chain A) coloured by the value of qic of each residue at Q=0.47 is shown in panels (C) and (D) for the Gō and smAMW models, respectively. The numbers indicate the HEAT repeat index number.
Discussion
Hierarchical unfolding of PR65/A
The equilibrium unfolding of PR65/A shows two transitions when monitored by either fluorescence or CD. However, the m-value obtained for the first transition when monitored by fluorescence (4 kcalmol−1M−1) is much smaller than that predicted for the longest contiguous segment that we know unfolds in this regime (HEAT 3–10, ~320 residues). Moreover, this m-value is bigger than that measured by CD (3 kcalmol−1M−1), and the m-values for mutants in these repeats are smaller than that of the wild type. These results indicate that the unfolding of this segment is not cooperative. The mutant and fragment data are therefore consistent with unfolding of PR65/A in a multi-step process involving five subdomains: HEAT 1–2, HEAT 3–7, HEAT 8–10, HEAT 11–13 and HEAT 14–15. We emphasise that we cannot define the boundaries of these subdomains in a very precise manner: for example, a mutation in HEAT 3 has a similar effect on the kinetics to mutations in HEAT 8–10, but we have grouped HEAT 3 together with the adjacent repeats into the HEAT 3–7 subdomain. We are able to delineate the relative stabilities of the subdomains: HEAT 3–7 and HEAT 8–10 are the least stable subdomains; they have similar stabilities to one another and the interface between them is weak. Consequently, unfolding of these two subdomains is decoupled and the single transition that is observed at equilibrium has a lower m-value than would be expected for a cooperative process; interestingly, the unfolding kinetics of the truncated variants and the double-jump experiments suggest that HEAT 3–7 and HEAT 8–10 unfold in separate steps. The HEAT 14–15 subdomain has low stability, like HEAT 3–7 and HEAT 8–10, and the unfolding of these subdomains may occur in parallel. Subdomains HEAT 1–2 and HEAT 11–13 have the highest stabilities.
We observe that folding of HEAT 1 and HEAT 11–13 are somewhat favored topologically, as shown by the pure structure-based model simulation. This feature is accentuated in the smAMW simulation, indicating that the heterogeneous contact potential and burial energy terms in this model reinforce the different stability of these repeats compared to the others. The simulation results therefore support the experimental evidence of the late unfolding (early folding) of HEAT 1–2 and HEAT 11–13. However, a prediction of the folding order of the other repeats at later stages of the folding funnel was more difficult, most likely because of small differences in their stabilization energy.
Origins and consequences of cooperativity in repeat proteins
Based on the order of unfolding events inferred from mutant studies, we have constructed a six-state kinetic model. Using a common set of rate coefficients and relative quantum yields for the different states, we are able to fit the kinetic data quantitatively over a wide-range of denaturant concentrations, for both wild-type and mutant proteins. The only parameter that is varied between the wild-type and each mutant protein is the unfolding rate of the region in which the mutation is located. Although there is a predominant order of unfolding, the model includes parallel unfolding pathways, which explain both the upward curvature in the fastest relaxation rate at high denaturant, as well as the fact that some mutations are able to alter more than one relaxation rate. HEAT 3–10 unfolds most rapidly, followed by HEAT 14–15, and then HEAT 1–2 and HEAT 11–13 unfold in slower steps. Because the distribution of stability between different subdomains of the array is very uneven, this order of unfolding events is strongly preferred, such that only very destabilizing mutations can shift the order, for example making unfolding of HEAT 14–15 competitive with HEAT 3–10. Similarly, our group and others have shown that alternative kinetic (un)folding routes are accessible for a number of ankyrin-repeat proteins such as myotrophin (7), D34 (35) and Notch (38), as was originally predicted based on the symmetric topology of repeat proteins (39). When intermediates are destabilized, unfolding of the HEAT repeats can occur in a single step, as occurs for example in the truncated variant HEAT 1–7.
The origins and limits of cooperativity have been explored for ankyrin and TPR repeats as well as more recently for LRRs, using both experimental and computational approaches (4, 5, 8, 39–42). These studies indicate that repeats that are distant from one another and therefore not directly in contact can nevertheless unfold in a coupled manner if they and the intervening units have similar intrinsic, low stabilities and the interfaces between them are strong. Cooperativity is thereby maintained over the medium range (4). If the distribution of stabilities is uneven, an intermediate state will be populated; for example in D34, a protein comprising twelve ankyrin repeats, there is an intermediate in which the C-terminal six repeats are structured and the N-terminal six repeats are unstructured (8). Moreover, cooperativity in repeat proteins can be easily modulated in a rational way, by changing the stabilities of the individual units; thus, we were able to make variants of D34 in which the unfolding of ten of the twelve repeats was coupled (8). In PR65/A, the stability is unevenly distributed across this large array with the result that multiple partly folded intermediates are populated.
Folding and function in repeat proteins
Partial unfolding of low stability regions of small repeat proteins is important for binding to partner proteins and for post-translational modifications (43–47). If some regions remain folded when others are unfolded, then processes such as translocation required for secretion and for proteasomal degradation may also be facilitated (31, 48), as shown for IκBα(46). An uneven stability distribution in long repeat arrays may be important for a distinct type of function: long arrays bind multiple partners by making use of distantly located sites. The ability to uncouple the folding of these sites means that they can employ a so-called ‘fly-casting’ mechanism for molecular recognition (28, 29). PR65/A uses its C-terminal HEAT repeats 11–15 to bind the catalytic C subunit of PP2A and the N-terminal HEAT repeats 1–10 to bind diverse regulatory B subunits (24). Intrinsic flexibility within the HEAT scaffold is apparent even in the x-ray structures of the complexes, with B- and C-subunit binding both resulting in significant reorientation of the inter- and intra-repeat packing; moreover, the low stability HEAT 3–10 region may be already partially unfolded under physiological conditions as shown by the denaturation profile of PR65/A at 37 °C (Fig. S1E). Thus, the central moiety could function to broaden the search area for binding partners and thereby coordinate molecular recognition at the two ends of the repeat array. Our results show how the absence of sequence-distant contacts in repeat proteins affords them particular flexibility, consistent with the proposal (12) that repeat proteins are a distinct structural class midway between globular structured proteins and intrinsically disordered proteins (49).
Materials and Methods
Molecular biology
The expression vector for PR65/A with an N-terminal hexahistidine tag was a kind gift from David Barford (ICR, London, UK). Mutant and truncated variants were made using the QuickChange site-directed mutagenesis kit (Stratagene) and confirmed by sequencing.
Protein expression and purification
Wild-type PR65/A, mutants and truncated variants were expressed and purified as described (22). Protein concentration was determined spectrophotometrically (50). The protein was frozen after purification and stored at −80 °C.
Fluorescence spectroscopy
Fluorescence was measured using a Perkin Elmer luminescence spectrometer LS55. Urea samples were prepared using a Hamilton Microlab 500 Series. Protein stock was added to a final concentration of 0.5 µM in 50 mM MES buffer pH 6.5, 1 mM DTT. The samples were equilibrated at 25 °C for at least two hours before measurement. The 1 cm pathlength cuvette was thermostatted using a water bath (Thermo Neslab). An excitation wavelength of 280 nm was used and the excitation and emission bandwidths were 5 nm. Wavelength scans between 320 nm and 370 nm were performed for each sample at a scan speed of 1 nm s−1.
Circular dichroism
Sample preparation was as for fluorescence except that the final protein concentration was 1.5 µM. Samples were incubated at 25 °C for 2 h in 25 mM MES buffer pH 6.5, 1 mM DTE and different urea concentrations. Far-UV CD spectra between 215 nm and 225 nm were then collected on a Chirascan™ circular dichroism spectrometer (Applied Photophysics) using a 0.2 cm pathlength cuvette thermostatted with a waterbath.
Native PAGE
Protein samples at 5 µM concentration were incubated in 50 mM MES buffer pH 6.5 containing 0 M, 1 M, 2 M, 3 M, 4 M, 5 M, 6 M, 7 M or 8 M urea for 2 h and then diluted with a loading dye resulting in final urea concentrations of 0 M, 0.5 M, 1 M, 1.5 M, 2 M, 2.5 M, 3 M, 3.5 M and 4 M urea. Samples were subsequently loaded on a Novex® Tris-glycine gel (Invitrogen) and run for 3 h.
Limited proteolysis
PR65/A, either in buffer or after 2 h incubation in 2.5 M urea, was digested with Proteinase K (dilution of 5 units per ml) at room temperature. The reaction was then terminated by boiling the samples for 10 minutes. Digestion fragments were resolved by SDS-PAGE and blotted on a membrane for N-terminal sequencing (performed by the Department of Biochemistry, University of Cambridge).
Simulation methods
For the analysis of the folding energy landscape of PR65/A we used two Hamiltonians with different contact energy terms. One of these represents an off-lattice native structure based model of the type introduced by Eastwood and coworkers (51, 52) to investigate the role of topology on folding. The other model is based on the AMW Hamiltonian (53). The transferable AMW Hamiltonian includes non-native contacts for long-range interactions in sequence as well as water-mediated interactions. Here we used it along with native-based interactions for residues close in sequence. The energy functions of both models and the simulation protocols are described in the Supplementary Material.
Acknowledgements
We thank the Center of Theoretical Biological Physics for computational resources. P.G.W. and P.O.C. acknowledge NIH grants GM071862 and R01 GM44557. L.S.I., M.T. and E.S. are supported by the Medical Research Council of the UK (MRC) (grant G1002329) and the Medical Research Foundation. R.B.B. is supported by a Royal Society University Research Fellowship.





