Logo of nihpaAbout Author manuscriptsSubmit a manuscriptHHS Public Access; Author Manuscript; Accepted for publication in peer reviewed journal;
Structure. Author manuscript; available in PMC 2016 Nov 3.
Published in final edited form as:
PMCID: PMC4811334
NIHMSID: NIHMS723584
PMID: 26439765

A naturally occurring repeat protein with high internal sequence identity defines a new class of TPR-like proteins

Associated Data

Supplementary Materials

SUMMARY

Linear repeat proteins often have high structural similarity and low (~25%) pairwise sequence identities (PSI) among modules. We identified a unique P. anserina (Pa) sequence with tetratricopeptide repeat (TPR) homology, which contains longer (42 residue) repeats (42PRs) with an average PSI >91%. We determined the crystal structure of five tandem Pa 42PRs to 1.6Å, and examined the stability and solution properties of constructs containing three to six Pa 42PRs. Compared to 34-residue TPRs (34PRs), Pa 42PRs have a one-turn extension of each helix, and bury more surface area. Unfolding transitions shift to higher denaturant concentration and become sharper as repeats are added. Fitted Ising models show Pa 42PRs to be more cooperative than consensus 34PRs, with increased magnitudes of intrinsic and interfacial free energies. These results demonstrate the tolerance of the TPR motif to length variation, and provide a basis to understand the effects of helix length on intrinsic/interfacial stability.

INTRODUCTION

Linear repeat proteins consist of arrays of a structural motif, typically 20–40 residues in length. Adjacent motifs pack together to create elongated superhelical structures defined by geometric relationships between units (Kloss et al., 2008; Main et al., 2005; Kajava, 2002; Kobe and Kajava, 2000). One such motif is the tetratricopeptide repeat (TPR), a 34-residue motif found in a wide range of proteins from all three kingdoms of life.

TPR domains mediate protein-protein interactions. Although TPR sequences and functions vary widely, repeats have nearly identical geometries (D’Andrea and Regan, 2003). The structure of the TPR consists of two anti-parallel α-helices, termed “A” and “B”, which stack at an angle of ~160° (Blatch and L ässle, 1999). The structure and folding of 34 residue TPRs has been studied extensively using a series of consensus repeats, in which each repeat has the same sequence, based on multiple sequence alignments (Main et al., 2003; Kajander et al., 2005; Cortajarena and Regan, 2011). The application of consensus design methods to repeat proteins (Binz et al., 2003; Main et al., 2003; Mosavi et al., 2002; Parmeggiani et al., 2008; Urvoas et al., 2010), which is usually based on hidden Markov models (HMMs), highlights conserved residues of each motif. However, HMM programs are infrequently used to generate new motifs (Frith et al., 2008), and length variations in aligned sequences are therefore reduced to insertion and deletion probabilities within HMMs. This has the potential to mask significant length differences among distinct motif subfamilies.

One particularly interesting aspect of the TPR motif, compared to other linear repeat motifs, is the diversity of its repeat sequence. The Pfam 27.0 (Finn et al., 2014) TPR superfamily contains over 100 family members. Of these, 21 family members are classified TPR_1 through TPR_21. Although some family members have very similar HMM logos and consensus sequences (e.g. TPR_1 and TPR_2), other families differ in length and composition (length range: 26–280 residues). Though some of the longer families result from a classification of tandem repeats as a single motif (presumably due to high similarity between nonadjacent repeats), there is considerable length variation among families representing single repeats. This differs from other helical repeats such as ankyrin repeats, where sequence lengths are more tightly distributed (~33 residues/repeat).

A striking example of length variation in TPRs can be found in sequences classified as TPR_10. These sequences are 42 residues in length, as opposed to the founding 34 residue motif (Sikorski et al., 1990). Owing to the length variation observed in TPR sequences, we adopt a nomenclature that better reflects motif length: nPRs (the name TPR derives from the tetratrico prefix, meaning thirty-four; the “T” (for “tetra”, four) cannot capture variation in the tens digit). Here, n corresponds to the number of residues in a single repeat. For example, we refer to 42 residue nPR motifs as 42PRs, and 34 residue motifs as 34PRs.

There is little high-resolution structural information for 42PRs. The closest structural homologs are the TPR domains of human kinesin light chain (hKLC) isoforms 1 and 2, which were solved to 2.8 and 2.75 Å, respectively (Zhu et al., 2012). Some of the TPRs in hKLC1 and hKLC2 belong to TPR_10, perhaps due to the low identity between repeats. This limits an understanding of the structural features defining repeats belonging to 42PRs.

To explore the structural and thermodynamic implications of this new class of extended repeat sequences, we identified and characterized a rather unusual 42PR array from the Podospora anserina (Pa) genome (Espagne et al., 2008), containing 15 tandem 42PRs of nearly complete identity. We used the repeats in this sequence to design central (AB) and capping (NAB and ACB) 42PRs to create Pa 42PR constructs of the type NAB(AB)xACB. Here, x signifies the number of central AB units, ranging from one to four. We find NAB(AB)xACB constructs to be soluble, stable, and predominantly (>95%) monomeric below 10 μM. We determined the X-ray structure of NAB(AB)3ACB to 1.6 Å. The structure reveals a five-repeat Pa 42PR right-handed superhelix, with longer A and B-helices compared to canonical 34PRs. We find Pa 42PRs to be significantly less stable, yet more cooperative, than consensus 34PRs (c34PRs) of equivalent repeat, based on analysis of fitted one-dimensional (1D) Ising models to unfolding transitions.

RESULTS

Identification of a new class of 42PRs

Comparison of the lengths of TPR families 1–21 in Pfam 27.0 revealed two well represented repeat sequence lengths: 34 and 42 residues. We found many 42PR containing sequences to have a rather high average pairwise identity among repeats (internal sequence identity, ISI). This differs from the ISI in 34PRs, and the majority of other repeat protein motifs, which is typically only ~25%.

To define the shared and unique sequence features of 42PRs and 34PRs, we generated HMM sequence logos using seed sequences from Pfam 27.0. Seed sequences TPR_1 and TPR_2 were combined (924 total) to create a 34PR HMM, and 291 TPR_10 seed sequences were used to create a 42PR HMM. Alignment of the HMM logos shows conserved sequence features along the entire 34PR AB unit (Figure 1). Residues 2–31 in 34PRs and 6–35 in 42PRs show high sequence identity and have identical spacing between conserved positions. In this alignment, the 42PR sequence logo contains four-residue extensions on the N- and C-termini (residues 1–4 and 34–38 in the 42PR HMM, Figure 1B), which could conceivably extend the A and B-helices of the 34PR motif.

An external file that holds a picture, illustration, etc.
Object name is nihms723584f1.jpg
HMM logos and helix definitions of 34PR and 42PR sequences

(A) 34PR HMM sequence logo. (B) 42PR HMM sequence logo. A and B-helix boundaries are from the structures 1NA0 (c34PR) and 4Y6W (42PR). Special characters represent systematic hydrogen bonding interactions in 4Y6W (42PR) within (*) and between (#) repeats (see Figure 5). Logos were generated using Skylign (Wheeler et al., 2014) and aligned manually in the central region (residues 5–29 in the 34PR logo), which contains the greatest similarity between the two motifs.

Design of Podospora anserina repeats

We used the 42PR HMM to search for sequence members of this family and identified a 42PR-containing ORF, Pa_6_8860, present in the genome of the fungus P. anserina. The predicted domain architecture of Pa_6_8860 consists of an N-terminal partial NB-ARC domain (van Ooijen et al., 2008), followed by a region containing modest similarity to the heptad repeats in kinesin light chains (Cyr et al., 1991). The remainder of Pa_6_8860 encodes 15 putative 42PRs (Figure 2). These 42PRs are an extreme example of high (>91%) ISI.

An external file that holds a picture, illustration, etc.
Object name is nihms723584f2.jpg
Sequence features of Podospora anserina Pa_6_8860 ORF and Pa 42PR repeat design

(A) Predicted domain organization of Pa_6_8860. The predicted 42PR motifs are labeled as grey boxes. Sequence positions in the 15 central 42PRs that differ from the derived consensus after alignment are highlighted in yellow. (B) Designed Podospora anserina 42PR sequence (Pa AB) from the alignment in (A). Capping sequences NAB and ACB were created by substitution of non-polar residues for polar residues (red) to promote solubility. RS substitutions (blue) resulted from cloning of repeat arrays. The consensus 34PR sequence (Main et al., 2003), c34PR, is aligned as in Figure 1. A and B-helix boundaries, and special characters to indicate hydrogen bonding interactions are labeled as described in Figure 1. (C) Single-helix representation of Pa 42PR constructs used in this study, where x signifies the number of internal Pa AB units, ranging from one to four. A and B-helices are shown in red and magenta, respectively. Together with connecting turns, these two helices make up a full Pa 42PR. NA and CB capping helices are shown in grey.

To characterize the structure and stability of this unique Pa 42PR, and to compare it with the more common 34PR family, we designed constructs to express arrays of a core AB repeat (Pa AB; Figure 2B and 2C). Although proteins constructed solely from Pa AB units expressed well, they were insoluble and could not be characterized. Therefore, polar substitutions to putative solvent exposed hydrophobic residues were designed from helix wheel representations and made to the N-terminal A and C-terminal B-helices, to create capping repeats NAB and ACB (Figure 2B). Substitutions on surface-exposed sites are expected to minimize structural and energetic perturbations. This approach has been successful in promoting solubility in other repeat protein studies (Main et al. 2003; Wetzel et al. 2008; Aksel et al., 2011).

We used these capping repeats, along with unmodified AB units, to create constructs of the type NAB(AB)xACB, where the subscript x represents the number of AB core repeats, ranging from one to four (Figure 2C). We were able to express and purify these capped constructs with reasonable yields for biophysical characterization.

Solution structure of Pa 42PRs

To determine the hydrodynamic properties of our Pa 42PR constructs, we conducted analytical ultracentrifugation sedimentation velocity (AUC-SV) experiments. We modeled SV data using direct boundary (ΔC/ΔT) methods, as well as c(s) methods (Schuck 2000). The SV data indicate that all constructs populate predominantly (>97%) monomeric species at low (<10μM) concentrations (Figure 3). Higher concentrations result in weak associations; the extent and nature of association differs for each construct (Table 1).

An external file that holds a picture, illustration, etc.
Object name is nihms723584f3.jpg
Sedimentation velocity analytical ultracentrifugation of Pa 42PR NAB(AB)xACB constructs

Continuous c(s) distributions and global ΔC/ΔT fits of NAB(AB)ACB (panels A,B), NAB(AB)2ACB (panels C,D), NAB(AB)3ACB (panels E,F), and NAB(AB)4ACB (panels G,H), respectively. Lower panels are residuals between ΔC/ΔT values and fitted models. For clarity, only a subset of ΔC/ΔT values and curves are shown. See Table 1 for a summary of fitted models.

Table 1

Hydrodynamic properties of Pa 42PR constructs

ProteinMw (kDa)ModelaRMSDbs(A)cs (A2)cs (A2’)cs (A3’)cKD (mM)cr (An/A) (%)c,d
NAB(AB)ACB16.5M + IT4.50E-031.86 (1.85–1.86)N/AN/A3.57 (3.53–3.6)N/A1.33 (1.29–1.38)
NAB(AB)2ACB21.1M + IT6.56E-032.39 (2.37–2.4)N/AN/A4.14 (4.11–4.17)N/A2.73 (2.69–2.78)
NAB(AB)3ACB25.8SA + ID9.19E-032.39 (2.38–2.4)3.27 (3.24–3.3)4.08 (4.05–4.12)N/A1.2 (1.13–1.27)1.67 (1.64–1.7)
NAB(AB)4ACB30.4SA + ID7.51E-032.72 (2.7–2.75)4.58 (4.47–4.75)4.88 (4.82–4.94)N/A0.27 (0.24–0.32)2.63 (2.54–2.70)
aM, single species monomer; SA, equilibrium self-association; IT, incompetent trimer; ID, incompetent dimer.
bRoot-mean-squared-deviation of the global fit of the model to ΔC/ΔT data.
cFitted parameter 95% confidence intervals calculated from F-statistics (Johnson and Straume, 1994) are shown in parenthesis.
dRatio of incompetent species to total loading concentration of cell, expressed as a percentage.

Figure S1 shows a summary of models used in fitting.

NAB(AB)ACB and NAB(AB)2ACB SV c(s) distributions are consistent with predominantly monomeric species with higher concentrations displaying small peaks at higher s-values (Figures 3A and 3C). These higher s-value peaks are shifted from what would be expected for dimeric species. Sedimentation velocity ΔC/ΔT curves spanning a wide concentration range were globally fit using a monomer-incompetent trimer model (Figures 3B and 3D). This model assumes a small proportion of the loading concentration is present as a trimeric species that does not equilibrate with the monomer (Figure S1A). Incompetent species have been identified in a number of other AUC studies (Lemaire et al., 2005; Wowor et al., 2011; Xu, 2004). Although the fitted s- values of the incompetent species are consistent with molecular weights corresponding to trimers of each protein, given the low fitted concentrations of these species (less than 3% of the total loading concentration), it is possible they represent small amounts of impurities.

NAB(AB)3ACB and NAB(AB)4ACB SV c(s) distributions are consistent with predominantly monomeric species with higher concentrations displaying small peaks at s-values consistent with dimeric species (Figures 3E and 3G). Sedimentation velocity ΔC/ΔT curves were globally fit to a monomer-dimer, incompetent dimer model (Figures 3F and 3H). This model assumes a rapid and reversible equilibrium between monomer and dimer, and an additional dimeric species that does not equilibrate with the monomer (Figure S1B). A summary of hydrodynamic models and parameters used for each construct is shown in Table 1.

Based on sequence elements that match 34PR motifs, we expect the 42PRs derived from Pa to adopt an α-helical structure. Consistent with this expectation, far-UV CD spectra of all NAB(AB)xACB constructs show a high level of α-helical character with well-defined minima at 222 and 208nm (Figure 4A). CD spectra for Pa 42PR constructs suggest a higher level of α-helical character than c34PR constructs (Main et al. 2003). Observed variations in molar residue ellipticity values are likely due to uncertainties in protein concentrations. Each repeat contains only one tyrosine, and no tryptophans (Figure 2); thus, extinction coefficients are low.

An external file that holds a picture, illustration, etc.
Object name is nihms723584f4.jpg
Far-UV CD spectroscopy and global Ising modeling of equilibrium unfolding of 42PR and 34PR constructs

(A) Far-UV CD spectra of Pa 42PR constructs NAB(AB)ACB (black), NAB(AB)2ACB (purple), NAB(AB)3ACB (blue), and NAB(AB)4ACB (cyan). (B) Normalized equilibrium unfolding transitions of Pa 42PRs and c34PRs. Closed circles show Pa 42PR constructs, and are colored as in (A). Open circles show c34PR constructs: B(AB)S (grey), B(AB)2S (black), B(AB)3S (purple), and B(AB)4S (blue). With this color scheme, constructs with the same integer number of repeats have the same color. Solid and dashed lines result from globally fitting one-dimensional Ising models to Pa 42PRs and c34PRs, respectively. c34PR samples contain 50 mM sodium phosphate, 150 mM NaCl, and are at pH 6.8. Pa 42PR samples contain 25 mM Tris-HCl, 350 mM NaCl, and are at pH 8.0. All samples are at 25°C. See Figure S2 and Table S1 for two-state analysis of the same data.

Thermodynamic stability of Pa 42PRs

To measure the thermodynamic stability of designed Pa 42PRs, we carried out urea-induced equilibrium unfolding of constructs ranging in length from three to six total repeats (Figure 4B, closed circles), in 350 mM NaCl, pH 8. Under these conditions, all transitions are completely reversible and have fully resolved native baselines. As repeats are added, the transition midpoints increase and become sharper, indicating both increased stability and a high level of cooperativity. These observations are reflected in apparent free energies and m-values for unfolding obtained from a two-state fit (Table S1).

To compare stabilities and cooperativities of Pa 42PRs to shorter 34PR motifs, we also measured the urea-induced equilibrium unfolding of c34PRs of equivalent numbers of total repeats (Figure 4B, open circles), at 150 mM NaCl, pH 6.8. To generate c34PRs with integral numbers of whole repeats, we added a single c34PR B-helix to the N-termini of c34PR constructs studied by Regan and coworkers (Main et al., 2003; Kajander et al., 2005; Cortajarena and Regan, 2011). We term these constructs B(AB)xS, where x signifies the number of central c34PRs ranging from one to four, and S is the “solvation helix” designed by Regan and coworkers. The addition of the N-terminal B helix slightly increases the stability of each construct, and is in agreement with the intrinsic and interfacial helical coupling energies measured by Regan and coworkers (Kajander et al. 2005).

Although Pa 42PRs are larger, they have significantly lower urea midpoints, sharper transitions, and increased m-values than their c34PR counterparts (Figures 4, S2, and Table S1). In contrast to the m-values of c34PRs, which plateau at four repeats, Pa 42PR m-values continue to increase through six repeats. This reflects a higher level of cooperativity in Pa 42PRs than in c34PRs. The magnitudes of the fitted Pa 42PR two-state m-values correlate well with their larger motif size and the expected solvent accessible surface area (SASA) changes for unfolding.

One-dimensional Ising analysis

To obtain a mechanistic understanding of the apparent increase in cooperativity in Pa 42PRs compared to c34PRs, we analyzed unfolding transitions using a nearest-neighbor 1D Ising model (Aksel and Barrick 2009; Kajander et al. 2005; Mello and Barrick 2004; Aksel, Majumdar, and Barrick 2011). We globally fit all Pa 42PR unfolding transitions to an Ising model (Figure 4B, closed circles and solid lines) and compared them to a separate global fit of c34PR unfolding transitions (Figure 4B, open circles and dashed lines). The global parameters obtained from these fits are shown in Table 2. For both nPR series, cooperativity arises from unfavorable intrinsic repeat folding (ΔGi), and from favorable interfacial coupling between adjacent folded repeats (ΔGi,i+1). The cooperativity enhancement observed for Pa 42PR transitions results from an increase in magnitude of both ΔGi and ΔGi,i+1 compared to c34PR transitions, i.e. lower intrinsic stability and higher interfacial stability.

Table 2

Fitted Ising thermodynamic parameters to Pa 42PR and c34PRs

Repeatχ2/νaΔGibΔGi,i+1bmic
Pa 42PRs6.5E−52.01± 0.026 (1.81,2.2)d−4.63 ± 0.038 (−4.97,−4.38)d−0.572 ± 0.0041 (−0.604,−0.541)d
c34PRs1.64E−41.39 ± 0.042 (1.05,1.73)d−4.3 ± 0.067 (−4.93,−3.83)d−0.383 ± 0.005 (−0.426,−0.346)d

Ising parameters were obtained from a global fit of a nearest-neighbor model to Pa 42PR and c34PR equilibrium unfolding curves. Three or more independent unfolding transitions for each construct were included.

aReduced chi-squared
bkcal*mol−1
ckcal*mol−1*M−1
dFitted parameter 95% confidence intervals calculated using F-statistics (Johnson and Straume, 1994).

Structure determination of Pa 42PR motifs

To determine the atomic structure of tandem 42PR motifs from the Pa_6_8860 gene product, we crystallized a five-repeat Pa 42PR, NAB(AB)3ACB. This construct crystallized in space group P21212, and contained one molecule in the asymmetric unit, with a calculated solvent fraction of 0.45. Crystals diffracted X-rays past 1.59 Å, with an I/Iσ of ~7 in the highest resolution shell.

Although the data merged with good statistics (Table 3), we were unable to solve the structure using molecular replacement with various search models, including single and five repeat arrays of various nPR structures. Anomalous diffraction data collected from selenium-methionine containing protein failed to provide adequate anomalous signal to determine experimental phases, perhaps because the protein contained only two Met residues located at the N-terminus, outside of the nPRs.

Table 3

Data collection and refinement statistics

CrystalNativeQ17M SeMet
PDB accession code4Y6W4Y6C

Wavelength (Å)1.03750.978 (SAD peak)

Refinement resolution range40.12–1.587 (1.643–1.587)39.76–1.772 (1.836–1.773)

Space group (hkl)P21212P21212

Unit cell dimensions

a, b, c, (Å)81.071, 92.341, 30.81781.103, 91.242, 30.839

α, β, γ (°)90, 90, 9090, 90, 90

Rsym/Rmeas/Rpim0.112/0.114/0.021 (0.737/0.798/0.153)0.172/0.175/0.033 (0.348/0.360/0.093)

CC(1/2)/CC*0.999/1 (0.969/0.992)0.996/0.999 (0.981/0.995)

<I/σI>18.6 (7)17.7 (11.7)

Redundancy14.2 (14)14.3 (13.9)

Completeness (%)99.68 (98.14)99.92 (99.16)

No. Reflections

Total909,077 (85,701)659,856 (64,472)
Unique32,123 (3,110)23,013 (2,236)

Rwork/Rfree0.1778/0.2042 (0.1809/0.2281)0.1730/0.2135 (0.1885/0.1901)

No. Atoms17961904

Protein16811725
Water/solvent115169

RMS deviations

Bond lengths (Å)0.0050.006
Bond angles (Å)0.950.92

Ramachandran analysis

Most favored (%)9999
Allowed (%)11

Values enclosed in parenthesis represent the highest resolution shell.

Rsym = Σhkl | I(hkl) − <I(hkl)> | / Σhkl I(hkl)

Rmeas=∑hkl(n/(n-1))∣I(hkl)-<I(hkl)>I/∑hklI(hkl)

Rmeas=∑hkl(1/(n-1))∣I(hkl)-<I(hkl)>I/∑hklI(hkl)

Rwork = Σhkl | Fobs − Fcalc | / Σhkl Fobs; Rfree = test set 6.23% (Native) and 5.14% (Q17M)

CC∗=(2CC(1/2))/(1+CC(1/2))

To obtain a stronger anomalous signal we introduced a single Met substitution into each of the three central AB repeats. Of four substitution sites tested (L12M, Q17M, N19M, and I35M; numbering is with respect to the position within the 42-residue AB unit), only one substitution site (Q17M) yielded crystals that diffracted to high resolution. Collection of single-wavelength anomalous dispersion (SAD) data at the Se peak wavelength allowed for determination of experimental phases. These phases produced electron density maps of excellent quality that enabled building and refinement of the Q17M variant structure. The phases and model of Q17M were used to build and refine a model of NAB(AB)3ACB using the native data. Refinement and data collection statistics for both structures are shown in Table 3.

The structure reveals a five repeat right-handed superhelix, with an overall architecture similar to 34PR domains (Figure 5). However, each of the A and B helices is approximately one helical turn longer than in 34PRs. These extensions occur on the N-terminus of the A-helix, and the C-terminus of the B-helix, compared to canonical 34PR helices. The same architecture is maintained across the entire Pa 42PR array. The average backbone RMSD for all repeats is 1.63 Å. For repeats with identical sequence, the average backbone RMSD is 1.20 Å. Much of this deviation comes from the third central repeat, which differs from the first two by ~1.5 Å. This difference is exemplified by a shorter (6.92 Å) calculated Ai:Bi helical packing distance compared to the first and second central repeats (8.1 and 8.26 Å, respectively), and appears to be a result of slight helical distortions. These helical distortions may be induced by crystal lattice interactions, as there is a large interface between symmetry mates, centered on the third repeat (Figure S3). In contrast, the first two central repeats have an RMSD of 0.54 Å. For helices of each type, the backbone RMSD is 1.1 and 0.98 Å, for A and B-helices, respectively. All possible pairwise repeat alignments are shown in Figure S4.

An external file that holds a picture, illustration, etc.
Object name is nihms723584f5.jpg
Crystal Structure of Pa 42PR NAB(AB)3ACB

(A) Cartoon representation of the NAB(AB)3ACB crystal structure 4Y6W, colored from N-terminus (blue) to C-terminus (red). Crystallographic waters are omitted for clarity. (B) Surface representation of NAB(AB)3ACB molecule. View is same as in (A), with coloring scheme as in Figure 2C. Each 42-residue repeat includes an A and B helix. (C–E) Representative electron density (2FO-FC, contoured at one sigma) of interactions along the repeat array. (C) and (D) Hydrophobic residues in an intra-repeat (Ai:Bi) and inter-repeat (Bi:Ai+1) helical interface, respectively. (E) One of four conserved Tyr OηH--−OεC Glu hydrogen bonds present within each of the inter-repeat helical interfaces. Additional structural features are shown in Figures S3, S5, and S6.

Discussion

A nomenclature system for variable-length TPR-like motifs

The TPR sequence was originally identified as a 34-residue motif in Saccharomyces cerevisiae cell-cycle regulation machinery (Sikorski et al., 1990). The first structure of a TPR (Das et al., 1998), revealed two anti-parallel α-helices. A large number of 34 residue TPRs have since been discovered, and conform closely in sequence and structural features. However, as databases have grown, TPR-like sequences have appeared that differ from the canonical length. The nPR nomenclature introduced here (where n represents the number of residues in the repeating unit) captures this length variation.

Sequence features of the 42PR motif

We identified a new class of nPR sequences, which we term 42PRs. The repeating unit is 42 residues and shares sequence characteristics with canonical 34PRs (Figure 1). The main differences between the two sequence classes appear to be N and C-terminal extensions of the 34PR-defined A and B-helices, respectively, in 42PRs. The 42PR HMM shows a high degree of conservation near the N terminus. This conservation may reflect the sequence characteristics defining a longer A-helix. The C-terminus of the 42PR HMM shows less conservation, indicating the rules defining the extension of the B-helix may be less strict.

Conserved positions in both 42PRs and 34PRs include small and hydrophobic residues at helix interfaces and in turn regions (Figure 1). Based on these structurally restrictive environments, it is likely that the conserved nPR residues are responsible for defining the fold. These residues may act as staples, around which helical extensions can be accommodated. In this fashion, the structural registry and packing of conserved nPR residues between helices is maintained.

Structural features of nPR motifs

Due to the repetitive architecture of nPR proteins, their structures can be defined by a small number of repeating parameters: helix crossing angles, distances, and contacts. These parameters, and the tertiary structures they define, are important for function and for stability. Helix crossing angles determine the extent to which nPR arrays form a concave binding surface for target peptides and proteins (Cortajarena and Regan, 2006; Cortajarena et al., 2010; Zhu et al., 2012). The number and type of helical contacts (Ai:Bi, Bi:Ai+1, and Ai:Ai+1) likely contribute to cooperativity in folding. By comparing the structural and energetic features of consensus, naturally occurring, and highly repetitive 42- and 34PRs, we can determine which structural parameters are general to nPRs, to 42- versus 34PRs, and which are modulated by local sequence variation within families.

Aside from differences in helix lengths, many of the structural features of 42PRs are similar to those of 34PRs. Although there are small deviations from repeat to repeat, average crossing angles for the AiBi, Bi:Ai+1, and Ai:Ai+1 helices are similar (Table 4). Likewise, the helical distances and solvent accessible surface area (SASA) burial between helices are similar across nPR families (Table 4).

Table 4

Helix-helix interfaces in representative nPRs

4Y6W (Pa 42PR)3CEQ (42PR)1NA0 (c34PR)1ELW (34PR)
Average Ai:Bi helix crossing angle (°)159 ± 6.8159.3 ± 5.3162.2 ± 0.8169.6 ± 6.8
Average Bi:Ai+1 helix crossing angle (°)155 ± 2.3165.7 ± 6.2154 ± 1.4156.2 ± 6.1
Average Ai:Ai+1 helix crossing angle (°)20.5 ± 7.424.8 ± 1.430.6 ± 1.423.4 ± 5.2
Average Ai:Bi packing distance (Å)7.54 ± 0.627.37 ± 1.137.3 ± 0.016.92 ± 1.2
Average Bi:Ai+1 packing distance (Å)8.5 ± 0.037.91 ± 2.058.78 ± 0.168.69 ± 0.78
Average Ai:Ai+1 packing distance (Å)10.9 ± 0.9712.4 ± 0.4411 ± 0.229.93 ± 1.7
Average SASA burial in Ai:Bi (Å2)1308 ± 1021422 ± 1691363 ± 351260 ± 67
Average SASA burial in Bi:Ai+1 (Å2)1571± 781400 ± 1311291 ± 181252 ± 123
Average pairwise sequence ID between repeats97%48%96%22%
Total number of repeats54.53.53.5

Uncertainties represent standard errors on the mean.

In contrast, there are several notable differences in contacts within and between 42- and 34PRs (Figure 6). In addition to contacts within helices (main diagonals), nPR structures show a characteristic contact pattern, consisting mainly of contacts between successive helices (Ai:Bi and Bi:Ai+1; anti-diagonal features). Other contacts include those between successive A-helices (Ai:Ai+1; off-diagonal features), but are less frequent in 42PR structures than in 34PR structures. These Ai:Ai+1 contacts connect the regularly spaced, anti-diagonal contacts (Figure 6).

An external file that holds a picture, illustration, etc.
Object name is nihms723584f6.jpg
Contact maps of 42 and 34 residue nPRs

Contacts are defined as atom pairs from different residues that are within 2.2–4.0 Å. Points above the main diagonal represent backbone contacts. Points below the main diagonal represent backbone-side chain or side chain-side chain contacts. Protein structures are displayed in each column (cyan, 42PRs; yellow, 34PRs). Row I displays helical contacts over the entire structure. Rows II, III, and IV expand over the indicated residue range. Rows I and II display contacts within A-helices (red), B-helices (magenta), between Ai and Bi helices (blue), and between Bi and Ai+1 helices (grey). Contacts between Ai and Ai+1 helices appear as red, off-diagonal points in rows I and II. Rows III and IV show all hydrophobic and polar contacts, respectively, and are color coded with respect to distance.

The packing features describing Ai:Ai+1 helices can be visualized by alignment of AiBiAi+1 units from each motif with respect to AiBi, and are summarized in Table 4. Although the helix geometries within nPR families are similar, 42PRs have fewer Ai:Ai+1 contacts than 34PRs. For 3CEQ and 1ELW (naturally occurring 42- and 34PRs, respectively), the differences in Ai:Ai+1 contacts can be explained by local helix geometry (longer Ai:Ai+1 distances in 3CEQ). Local helix geometry in Pa 42PRs also affects the number of Ai:Ai+1 contacts, as the NA-helix kink results in more Ai:Ai+1 contacts in the N-terminal region of the contact plot, relative to the internal repeat region (Figure 6, row I). Despite these local effects, Pa 42PRs and c34PRs have overall similar Ai:Ai+1 helix distances. This suggests that sequence specific information also influences the extent of Ai:Ai+1 interaction.

To further analyze the types of interactions in the Ai:Bi, Bi:Ai+1, and Ai:Ai+1 interfaces, we sorted polar and nonpolar interactions into separate contact maps (Figure 6, rows III and IV). Sorted contact plots reveal that packing within all interfaces, both within (Ai:Bi) and between (Bi:Ai+1) repeats, is predominantly hydrophobic, although these hydrophobic contacts (Figures 5C, 5D, and Figure 6, row III) occur at greater distances than the polar contacts (Figure 6, row IV). Polar contacts are also present within the Bi:Ai+1 interfaces of Pa 42PRs, including a conserved Tyr OηH--−OεC Glu hydrogen bond on the convex side of the Pa 42PR superhelix, between adjacent 42PR motifs (Figures 1B, ​,5E).5E). Other repetitive polar interactions include a His-Ser-Gln hydrogen bond network connecting successive AiBiAi+1 helices (Figure S5) and a His-Ser A-helix N-terminal capping motif (Figure S6).

Interestingly, the 42PRs of the human kinesin light chain (3CEQ) contain a region with Bi:Bi+1 helical contacts (Figure 6, Row I), which are not seen in 34PRs. This results from a slight kink in one B-helix, allowing for enhanced hydrophobic packing, along with other contacts between polar residues on the convex face of the superhelix. These sequence and length-specific structural variations highlight the structural malleability of the nPR motif. In naturally occurring repeat proteins, sequence variation can locally tune structural features. A consensus design approach applied to 42PRs would be expected to reveal representative interactions across all 42PRs.

Folding of Pa 42PRs

The cooperative folding of the Pa 42PR arrays in this study is striking. As repeats are added, both stabilities and m-values increase. This phenomenon is characteristic of other linear repeat proteins, especially ankyrin repeats, where energetic coupling leads to highly cooperative folding (Aksel et al., 2011; Wetzel et al., 2008). In contrast, the m-values of c34PRs plateau at four repeats, consistent with a high level of partially folded states (Kajander et al., 2005; Cortajarena and Regan 2011), and decreased cooperativity compared to ankyrin repeats and the Pa 42PRs presented here.

Global fits of a 1D Ising (nearest-neighbor) model to Pa 42PR and c34PRs provide a quantitative description of cooperativity in these two repeat systems (Figure 4B). We find the cooperativity enhancement in Pa 42PRs to result from a decrease in the intrinsic stability (an increase in ΔGi) and an increase in the interfacial stability (a decrease in ΔGi,i+1; Table 2). Stability studies under the same conditions will reveal the extent to which these differences depend on differences in solution conditions, although preliminary results suggest Pa 42PRs m-values to be insensitive to salt concentration (JDM and DB, data not shown). The decreased stability of individual Pa 42PRs is surprising, as each Pa 42PR helix is longer, and has more hydrogen bonds. Moreover, the 42PR helices are predicted to be more stable, based on Agadir (2.4 and 8.28%; Muñoz and Serrano, 1994), than the helices of c34PR (0.4 and 1.74%). The best-fit intrinsic m-values (mi) for both repeat types are consistent with an increased level of denaturant sensitive helical structure in Pa 42PRs.

Possible functions of the 42PR family genes and the implications of identical repeats

Although the function of Pa_6_8860 is unclear, many homologous sequences share a common architecture of different N-terminal domains, flanked by nPRs and other repeat protein types near the C-terminus (van der Biezen and Jones, 1998). The annotated functions of these proteins range from apoptosis and cell death regulation to plant resistance. The N-terminal portion of Pa_6_8860 is predicted to contain a partial NB-ARC domain, although it lacks some of the key residues involved in binding ATP (Yan et al., 2005). The 42PRs of Pa_6_8860 show the greatest sequence identity to the human kinesin light chain (KLC) nPRs, and there is evidence that many KLC domains contain nPRs (Pernigo et al., 2013; Fischer et al., 2012; Zhu et al., 2012; Gindhart and Goldstein, 1996); a subset of these have been shown to bind to cargo. Thus, it is plausible that Pa_6_8860 functions as a kinesin light chain, and the Pa 42PRs may be involved in cargo binding in microtuble-based vesicular transport.

In the crystal lattice, there is an extensive interface between symmetry mates (Figure S3). This interface buries ~4800 Å2 of total SASA. It is possible this dimer reflects the one characterized by AUC-SV, which has a fitted KD of 1.2mM. Interestingly, the addition of a single repeat to this protein results in a ~4.5-fold tighter dimerization KD. It is therefore possible the 15 tandem Pa 42PRs in Pa_6_8860 have the potential to form even tighter interactions. If this dimerization surface overlaps the cargo binding surface, dimerization and cargo binding would likely be competitive, and thus, cargo binding may dissociate kinesin light-chains.

Due to the high ISI and a large number (15) of Pa 42PRs in Pa_6_8860, the 42PR domain is expected to have multiple identical binding sites. These identical sites would have the potential to display an avidity effect for polyvalent targets (with direct sequence and/or structural repetition). An example of direct tandem repeats binding to a repetitive target is the TALE repeats of plant pathogenic bacteria, which bind to duplex DNA (Boch et al., 2009; Deng et al., 2012; Mak et al., 2012).

EXPERIMENTAL PROCEDURES

Subcloning, protein expression, and purification

DNA sequences encoding AB, NAB, and ACB repeats (Figure 2B) were cloned from codon optimized oligonucleotides. Annealed single-repeat cassettes were ligated directly into NdeI and BglII digested pET-15b (Novagen, Madison, WI). BamHI sites were included in the AB and ACB cassettes to allow for ligation as previously described (Aksel et al, 2012). Single-site substitutions (L12M, Q17M, N19M, and I35M) were introduced using Quikchange (Stratagene, La Jolla, CA) on individual AB cassettes. c34PR constructs were created using a similar approach.

Pa 42PR constructs were expressed in Escherichia coli Rosetta R2* (DE3) cells. One liter cultures were grown in terrific broth to an OD600 of 0.8, induced by adding IPTG to 200 μM, and incubated overnight at 20° C. Bacteria were pelleted, and lysed in 50 mL 25 mM Tris-HCl, 350 mM NaCl, 25 mM imidazole, 10 mM MgCl2 pH 8, 1 mg DNase, and tagged proteins were purified from the supernatant via Ni-NTA chromatography. Purified proteins were dialyzed extensively into 25 mM Tris-HCl, 350 mM NaCl pH 8, concentrated using an Amicon stirred cell concentrator (EMD Millipore, USA), and flash frozen at -80° C. Protein concentrations were determined as previously described (Edelhoch, 1967).

To express selenium-methionine–substituted proteins, cells were pelleted at an OD600 of 0.8, and were resuspended in M9 medium containing 100 mg/L selenium-methionine (Acros Organics, USA), 500 mg/L lysine, phenylalanine, and threonine, 250 mg/mL isoleucine, leucine, and valine (inhibitory amino acids for methionine biosynthesis), and 200 μM IPTG. During purification and analysis, 5 mM TCEP was included to selenium-methionine-substituted protein samples to ensure reduction of selenium. Complete selenium-methionine incorporation was confirmed using mass spectrometry.

Circular dichroism spectroscopy

CD measurements were conducted using an Aviv Model 400 CD Spectropolarimeter (Lakewood, NJ). CD samples contained 25 mM Tris-HCl, 350 mM NaCl, pH 8 (Pa 42PRs) or 50 mM Na Phosphate, 150 mM NaCl, pH 6.8 (c34PRs). Far-UV CD spectra were recorded at 25°C using a 0.1 cm path-length quartz cuvette (Starna Cells Inc., Atascadero, CA) at protein concentrations ranging from 15–25 μM. Spectra were obtained by signal averaging every 1 nm for 30 s. Buffer spectra were subtracted prior to analysis.

Urea-induced equilibrium unfolding

Unfolding transitions were obtained by monitoring CD at 222 nm (Pa 42PRs) and 220 nm (c34PRs) in a 1 cm path-length quartz cuvette. High purity urea (Amresco, Solon, OH) was deionized by stirring with mixed-bed resin (Bio-Rad, Hercules, CA) as previously described in (Street et al., 2008). Urea concentration was determined by refractometry (Pace 1986). Titrations were performed at protein concentrations ranging from 1.0–2.5 μM using a computer-controlled Microlab titrator (Hamilton, Reno, NV). At each urea concentration, protein samples were equilibrated for 5–7 min at 25° C, and signal averaged for 30 s. Two-state analysis was performed previously described (Street et al., 2008), and shown in Figure S2.

Global analysis of equilibrium unfolding transitions of Pa 42PRs and c34PRs using 1D Ising models was performed by constructing partition functions from two-by-two transfer matrices as previously described (Aksel and Barrick 2009). These matrices contain intrinsic and interfacial free energy terms (ΔGi and ΔGi,i+1). A single ΔGi parameter is used for the three internal and capping intrinsic energies of the energies of Pa 42PR (NAB, AB, and ACB). Likewise, another ΔGi parameter is used for the two internal and capping intrinsic energies of the energies of c34PR (BA and BS). The use of single ΔGi terms for capping and internal repeats is justified by preliminary results from cap deletion studies (JDM & DB, unpublished), which show similar stability decrements for capping and internal repeats.

Denaturant sensitivities (mi=-dΔGi/d[Urea]) were ascribed to instrinsic stability through a linear relationship. Functions describing the fraction of folded repeats were generated from the partition functions, and used with baseline parameters to model urea-induced equilibrium unfolding curves. Fitting was performed using a python program written by J.D.M., importing lmfit (Newville et al., 2014) to minimize a global objective function using non-linear least-squares.

The Ising model used here has a significantly lower parameter to construct ratio than analysis using two-state models (0.75 thermodynamic parameters per construct versus 2.0 for two-state analysis; compare Tables 2 and S2). This is accompanied by increased degrees of freedom (ν=249, Pa 42PRs; ν=306, c34PRs, compared to ~20 for two-state analysis of individual constructs) in fitting. Therefore, error analysis and interpretation of the fitted parameters are more robust.

Analytical Ultracentrifugation

Analytical ultracentrifugation sedimentation velocity (AUC-SV) experiments were performed using a Beckman XL-I analytical ultracentrifuge. Prior to AUC experiments, all proteins were extensively dialyzed into CD buffer. Protein concentrations ranged from 5–100 μM.

AUC-SV cells were assembled using SedVel60K 1.2 mm meniscus-matching centerpieces (SpinAnalytical) and sapphire windows. All other cell components were purchased from Beckman Coulter. Upon sample and reference (dialysate) loading, centerpieces were aligned in a An-60Ti rotor, and menisci were matched according to (Allgood and Barrick, 2011). After remixing, the rotor was thermally equilibrated under vacuum at 25° C for at least 90 minutes. Experiment s were run for approximately 8 hours at 45–50 krpm.

Protein Crystallization and Data Collection

Crystals of native NAB(AB)3ACB were grown at room temperature (~22° C) by hanging -drop vapor diffusion. Protein solution (20 mg/mL) was mixed in either a 2:1 or 1:1 ratio with reservoir solution containing 0.1 M MES (pH 6.5) and 25–30% PEG 4K. Crystals appeared after approximately 3–7 days. Crystals were cryoprotected by transfer into a solution consisting of 0.1 M MES (pH 6.5), 35% PEG 4K, and 5–10% ethylene glycol, and then flash frozen in liquid nitrogen. The selinium-methionine-substituted Q17M variant gave rise to morphologically similar crystals under these conditions, with the addition of 5 mM TCEP.

Native and selenium derivative data sets were collected at the National Synchrotron Light Source (NSLS) beamlines X-25 and X-29 (Brookhaven National Laboratory, Brookhaven, NY)) and processed with HKL2000 (Otwinowski and Minor, 1997). Crystals belong to space group P21212 and contain one molecule per asymmetric unit.

Structure Determination and Analysis

Selenium positions were determined by SAD in ShelXC/D (Sheldrick, 2010; Schneider and Sheldrick, 2002) through HKL2MAP (Pape and Schneider 2004). Phases were calculated in SOLVE and improved by density modification in RESOLVE (Terwilliger 2003). Iterative rounds of building and refinement were performed using COOT (Emsley and Cowtan 2004), Phenix (Adams et al. 2010; Afonine et al. 2012; Afonine et al. 2009; Afonine et al. 2013; Headd et al. 2012), and Refmac (Murshudov et al., 1997). The final model was validated with the program Molprobity (Chen et al., 2010). The native structure was built from the refined Q17M structure using Phaser (McCoy et al., 2007). Structural images were generated using PyMOL Version 1.5.0.4 (Schrödinger, LLC). Helix crossing angles were calculated using helix_angles.py (R.L. Campbell, Queens University). SASA calculations were performed using MSMS (Sanner et al., 1996). Structural alignments were performed using LSQMAN (Kleywegt, 1996).

​

Highlights

  • TPR-like repeats are described with variable sequence motif lengths

  • A P anserina 42 residue TPR-like repeat (42PR) has high internal sequence ID (>91%)

  • 42PRs have extended helices compared to canonical 34-residue TPRs (34PRs)

  • P. anserina 42PRs fold more cooperatively than consensus 34PRs

Supplementary Material

supplement

Acknowledgments

We would like to thank Drs Annie Héroux and Howard Robinson of the Macromolecular Crystallography Research Resource (PXRR) at the National Synchrotron Light Source, Brookhaven, NY for their technical assistance with data collection strategy and preliminary analysis. We also thank Dr. Phil Mortimer of the JHU mass spectrometry facility and Dr. Michael Love of the JHMI X-ray facility. This work was supported by NIH grant R01 GM068462 to D.B., J.D.M. was supported by NIH training grant T32-GM008403, J.M.K. was supported by NIH grant R01 GM099231 to Daniel J. Leahy, and G.D.B. was supported by NIH grant R01 GM084192.

Footnotes

AUTHOR CONTRIBUTIONS: J.D.M. and D.B. designed the experiments. J.D.M. conducted the experiments, analyzed the results, and prepared manuscript figures. J.D.M., G.D.B., and J.M.K. interpreted and analyzed crystallographic data. J.D.M. solved the crystal structures of 4Y6W and 4Y6C. J.D.M. and D.B. wrote the manuscript. J.D.M, J.M.K., G.D.B., and D.B. edited the manuscript.

Accession numbers: The PDB accession numbers for the Pa 42PR structures NAB(AB)3ACB and Q17M reported in this paper are 4Y6W and 4Y6C, respectively.

Publisher's Disclaimer: This is a PDF file of an unedited manuscript that has been accepted for publication. As a service to our customers we are providing this early version of the manuscript. The manuscript will undergo copyediting, typesetting, and review of the resulting proof before it is published in its final citable form. Please note that during the production process errors may be discovered which could affect the content, and all legal disclaimers that apply to the journal pertain.

References

  • Adams PD, Afonine PV, Bunkóczi G, Chen VB, Davis IW, Echols N, Headd JJ, Hung LW, Kapral GJ, Grosse-Kunstleve RW, et al. PHENIX: a comprehensive Python-based system for macromolecular structure solution. Acta Crystallographica Section D Biological Crystallography. 2010;66:213–221. [PMC free article] [PubMed] [Google Scholar]
  • Afonine PV, Grosse-Kunstleve RW, Urzhumtsev A, Adams PD. Automatic multiple-zone rigid-body refinement with a large convergence radius. Journal of Applied Crystallography. 2009;42:607–615. [PMC free article] [PubMed] [Google Scholar]
  • Afonine PV, Grosse-Kunstleve RW, Echols N, Headd JJ, Moriarty NW, Mustyakimov M, Terwilliger TC, Urzhumtsev A, Zwart PH, Adams PD. Towards automated crystallographic structure refinement with phenix.refine. Acta Crystallographica Section D Biological Crystallography. 2012;68:352–367. [PMC free article] [PubMed] [Google Scholar]
  • Afonine PV, Grosse-Kunstleve RW, Adams PD, Urzhumtsev A. Bulk-solvent and overall scaling revisited: faster calculations, improved results. Acta Crystallographica Section D Biological Crystallography. 2013;69:625–634. [PMC free article] [PubMed] [Google Scholar]
  • Aksel T, Barrick D. Methods in Enzymology. Elsevier; 2009. Chapter 4 Analysis of Repeat-Protein Folding Using Nearest-Neighbor Statistical Mechanical Models; pp. 95–125. [PMC free article] [PubMed] [Google Scholar]
  • Aksel T, Majumdar A, Barrick D. The Contribution of Entropy, Enthalpy, and Hydrophobic Desolvation to Cooperativity in Repeat-Protein Folding. Structure. 2011;19:349–360. [PMC free article] [PubMed] [Google Scholar]
  • Allgood AG, Barrick D. Mapping the Deltex-Binding Surface on the Notch Ankyrin Domain Using Analytical Ultracentrifugation. Journal of Molecular Biology. 2011;414:243–259. [PMC free article] [PubMed] [Google Scholar]
  • van der Biezen EA, Jones JD. The NB-ARC domain: a novel signalling motif shared by plant resistance gene products and regulators of cell death in animals. Current Biology. 1998;8:R226–R228. [PubMed] [Google Scholar]
  • Binz HK, Stumpp MT, Forrer P, Amstutz P, Plückthun A. Designing Repeat Proteins: Well-expressed, Soluble and Stable Proteins from Combinatorial Libraries of Consensus Ankyrin Repeat Proteins. Journal of Molecular Biology. 2003;332:489–503. [PubMed] [Google Scholar]
  • Blatch GL, Lässle M. The tetratricopeptide repeat: a structural motif mediating protein-protein interactions. BioEssays. 1999;21:932–939. [PubMed] [Google Scholar]
  • Boch J, Scholze H, Schornack S, Landgraf A, Hahn S, Kay S, Lahaye T, Nickstadt A, Bonas U. Breaking the code of DNA binding specificity of TAL-Type III effectors. Science. 2009;326:1509–1512. [PubMed] [Google Scholar]
  • Chen VB, Arendall WB, Headd JJ, Keedy DA, Immormino RM, Kapral GJ, Murray LW, Richardson JS, Richardson DC. MolProbity: all-atom structure validation for macromolecular crystallography. Acta Crystallographica Section D Biological Crystallography. 2010;66:12–21. [PMC free article] [PubMed] [Google Scholar]
  • Cortajarena AL, Regan L. Ligand binding by TPR domains. Protein Science. 2006;15:1193–1198. [PMC free article] [PubMed] [Google Scholar]
  • Cortajarena AL, Regan L. Calorimetric study of a series of designed repeat proteins: Modular structure and modular folding. Protein Science. 2011;20:336–340. [PMC free article] [PubMed] [Google Scholar]
  • Cortajarena AL, Wang J, Regan L. Crystal structure of a designed tetratricopeptide repeat module in complex with its peptide ligand: Structure of designed TPR module-ligand complex. FEBS Journal. 2010;277:1058–1066. [PubMed] [Google Scholar]
  • Cyr JL, Pfister KK, Bloom GS, Slaughter CA, Brady ST. Molecular genetics of kinesin light chains: generation of isoforms by alternative splicing. Proc Natl Acad Sci USA. 1991;88:10114–10118. [PMC free article] [PubMed] [Google Scholar]
  • D’Andrea LD, Regan L. TPR proteins: the versatile helix. Trends in Biochemical Sciences. 2003;28:655–662. [PubMed] [Google Scholar]
  • Das AK, Cohen PT, Barford D. The structure of the tetratricopeptide repeats of protein phosphatase 5: implications for TPR-mediated protein–protein interactions. The EMBO Journal. 1998;17:1192–1199. [PMC free article] [PubMed] [Google Scholar]
  • DeLano WL. The PyMOL Molecular Graphics System, version 1.5.1. Schrödinger, LLC; New York: 2010. [Google Scholar]
  • Deng D, Yan C, Pan X, Mahfouz M, Wang J, Zhu JK, Shi Y, Yan N. Structural basis for sequence-specific recognition of DNA by TAL effectors. Science. 2012;335:720–723. [PMC free article] [PubMed] [Google Scholar]
  • Emsley P, Cowtan K. Coot: model-building tools for molecular graphics. Acta Crystallographica Section D Biological Crystallography. 2004;60:2126–2132. [PubMed] [Google Scholar]
  • Edelhoch H. Spectroscopic determination of tryptophan and tyrosine in proteins. Biochemistry. 1967;6:1948–1954. [PubMed] [Google Scholar]
  • Espagne E, Lespinet O, Malagnac F, Da C, Aury M, Ségurens B, Poulain J, Anthouard V, Grossetete S, Khalili H, et al. The genome sequence of the model ascomycete fungus Podospora anserina. Genome Biology. 2008;9:R77. [PMC free article] [PubMed] [Google Scholar]
  • Finn RD, Bateman A, Clements J, Coggill P, Eberhardt RY, Eddy SR, Heger A, Hetherington K, Holm L, Mistry J, et al. Pfam: the protein families database. Nucleic Acids Research. 2014;42:D222–D230. [PMC free article] [PubMed] [Google Scholar]
  • Fisher SQ, Weck M, Landers JE, Emrich J, Middleton SA, Cox J, Gentile L, Parish CA. Evidence that the kinesin light chain domain contains tetratricopeptide repeat units. Journal of Structural Biology. 2012;177:602–612. [PubMed] [Google Scholar]
  • Frith MC, Saunders NFW, Kobe B, Bailey TL. Discovering Sequence Motifs with Arbitrary Insertions and Deletions. PLoS Computational Biology. 2008;4:e1000071. [PMC free article] [PubMed] [Google Scholar]
  • Gindhart JG, Jr, Goldstein LSB. Tetratrico peptide repeats are present in the kinesin light chain. TIBS Letters. 1996;21:52–53. [PubMed] [Google Scholar]
  • Headd JJ, Echols N, Afonine PV, Grosse-Kunstleve RW, Chen VB, Moriarty NW, Richardson DC, Richardson JS, Adams PD. Use of knowledge-based restraints in phenix.refine to improve macromolecular refinement at low resolution. Acta Crystallographica Section D Biological Crystallography. 2012;68:381–390. [PMC free article] [PubMed] [Google Scholar]
  • Hunter JD. Matplotlib: A 2D Graphics Environment. Computing in Science & Engineering. 2007;9(3):90–95. [Google Scholar]
  • Johnson ML, Straume M. Comments on the analysis of sedimentation equilibrium experiments. In: Shuster TM, Laue TM, editors. Modern Analytical Ultracentrifugation. Boston: Birkhauser; 1994. pp. 37–65. [Google Scholar]
  • Kajander T, Cortajarena AL, Main ERG, Mochrie SGJ, Regan L. A New Folding Paradigm for Repeat Proteins. Journal of the American Chemical Society. 2005;127:10188–10190. [PubMed] [Google Scholar]
  • Kajava AV. What curves α-solenoids? Evidence for an α-helical toroid structure of Rpn1 and Rpn2 proteins of the 26 S proteasome. Journal of Biological Chemistry. 2002;277:49791–49798. [PubMed] [Google Scholar]
  • Karplus PA, Diederichs K. Linking crystallographic model and data quality. Science. 2012;6084:1030–1033. [PMC free article] [PubMed] [Google Scholar]
  • Kleywegt GJ. Use of non-crystallographic symmetry in protein structure refinement. Acta Crystallographica Section D Biological Crystallography. 1996;52:842–857. [PubMed] [Google Scholar]
  • Kloss E, Courtemanche N, Barrick D. Repeat-protein folding: New insights into origins of cooperativity, stability, and topology. Archives of Biochemistry and Biophysics. 2008;469:83–99. [PMC free article] [PubMed] [Google Scholar]
  • Kobe B, Kajava AV. When protein folding is simplified to protein coiling: the continuum of solenoid protein structures. Trends in Biochemical Sciences. 2000;25:509–515. [PubMed] [Google Scholar]
  • Lemaire PA, Lary J, Cole JL. Mechanism of PKR Activation: Dimerization and Kinase Activation in the Absence of Double-stranded RNA. Journal of Molecular Biology. 2005;345:81–90. [PubMed] [Google Scholar]
  • Main E, Lowe A, Mochrie S, Jackson S, Regan L. A recurring theme in protein engineering: the design, stability and folding of repeat proteins. Current Opinion in Structural Biology. 2005;15:464–471. [PubMed] [Google Scholar]
  • Main ERG, Xiong Y, Cocco MJ, D’Andrea L, Regan L. Design of Stable α-Helical Arrays from an Idealized TPR Motif. Structure. 2003;11:497–508. [PubMed] [Google Scholar]
  • Mak N-SM, Bradley P, Cernadas RA, Bogdanove AJ, Stoddard BL. The crystal structure of TAL Effector PthXo1 Bound to Its DNA Target. Science. 2012;335:716–719. [PMC free article] [PubMed] [Google Scholar]
  • McCoy AJ, Grosse-Kunstleve RW, Adams PD, Winn MD, Storoni LC, Read RJ. Phaser crystallographic software. Journal of Applied Crystallography. 2007;40:658–674. [PMC free article] [PubMed] [Google Scholar]
  • Mello CC, Barrick D. An experimentally determined protein folding energy landscape. Proc Natl Acad Sci USA. 2004;101:14102–14107. [PMC free article] [PubMed] [Google Scholar]
  • Mosavi LK, Minor DL, Peng Z. Consensus-derived structural determinants of the ankyrin repeat motif. Proc Natl Acad Sci USA. 2002;99:16029–16034. [PMC free article] [PubMed] [Google Scholar]
  • Muñoz V, Serrano L. Elucidating the folding problem of helical peptides using empirical parameters. Nat Struct Biol. 1994;1:399–409. [PubMed] [Google Scholar]
  • Murshudov GN, Vagin AA, Dodson EJ. Refinement of macromolecular structures by the maximum-likelihood method. Acta Crystallographica Section D: Biological Crystallography. 1997;53:240–255. [PubMed] [Google Scholar]
  • Newville M, Stensitzki T, Allen DB, Ingargiola A. LMFIT: Non-Linear Least-Square Minimization and Curve-Fitting for Python. Zenodo. 2014 doi: 10.5281/zenodo.11813. [CrossRef] [Google Scholar]
  • van Ooijen G, Mayr G, Kasiem MMA, Albrecht M, Cornelissen BJC, Takken FLW. Structure-function analysis of the NB-ARC domain of plant disease resistance proteins. Journal of Experimental Botany. 2008;59:1383–1397. [PubMed] [Google Scholar]
  • Otwinowski Z, Minor W. Processing of X-ray diffraction data collected in oscillation mode. Methods in Enzymol. 1997;276:307–326. [PubMed] [Google Scholar]
  • Pace CN. Determination and analysis of urea and guanidine hydrochloride denaturation curves. Methods in Enzymology. 1986;131:266–280. [PubMed] [Google Scholar]
  • Pape T, Schneider TR. HKL2MAP: a graphical user interface for macromolecular phasing with SHELX programs. Journal of Applied Crystallography. 2004;37:843–844. [Google Scholar]
  • Parmeggiani F, Pellarin R, Larsen AP, Varadamsetty G, Stumpp MT, Zerbe O, Caflisch A, Plückthun A. Designed Armadillo Repeat Proteins as General Peptide-Binding Scaffolds: Consensus Design and Computational Optimization of the Hydrophobic Core. Journal of Molecular Biology. 2008;376:1282–1304. [PubMed] [Google Scholar]
  • Pernigo S, Lamprecht A, Steiner RA, Dodding MP. Structural basis for Kinesin-1: cargo recognition. Science. 2013;340:356–359. [PMC free article] [PubMed] [Google Scholar]
  • Sanner MF, Olson AJ, Spehner JC. Reduced Surface: an Efficient Way to Compute Molecular Surfaces. Biopolymers. 1996;38:305–320. [PubMed] [Google Scholar]
  • Schuck P. Size-distribution analysis of macromolecules by sedimentation velocity ultracentrifugation and lamm equation modeling. Biophysical Journal. 2000;78:1606–1619. [PMC free article] [PubMed] [Google Scholar]
  • Schneider TR, Sheldrick GM. Substructure Solution with SHELXD. Acta Crystallographica Section D Biological Crystallography. 2002;58:1772–1779. [PubMed] [Google Scholar]
  • Sheldrick GM. Experimental phasing with SHELXC / D / E: combining chain tracing with density modification. Acta Crystallographica Section D Biological Crystallography. 2010;66:479–485. [PMC free article] [PubMed] [Google Scholar]
  • Sikorski RS, Boguski MS, Goebl M, Heiter P. A repeating amino acid motif in CDC23 defines a new family of proteins and a new relationship among genes required for mitosis and RNA synthesis. Cell. 1990;60(2):307–317. [PubMed] [Google Scholar]
  • Stafford WF, Sherwood PJ. Analysis of heterologous interacting systems by sedimentation velocity: curve fitting algorithms for estimation of sedimentation coefficients, equilibrium and kinetic constants. Biophysical Chemistry. 2004;108:231–243. [PubMed] [Google Scholar]
  • Street TO, Courtemanche N, Barrick D. Methods in Cell Biology. Elsevier; 2008. Protein Folding and Stability Using Denaturants; pp. 295–325. [PubMed] [Google Scholar]
  • Terwilliger TC. SOLVE and RESOLVE: automated structure solution and density modification. Methods Enzymol. 2003;374:22–37. [PubMed] [Google Scholar]
  • Urvoas A, Guellouz A, Valerio-Lepiniec M, Graille M, Durand D, Desravines DC, van Tilbeurgh H, Desmadril M, Minard P. Design, Production and Molecular Structure of a New Family of Artificial Alpha-helicoidal Repeat Proteins (αRep) Based on Thermostable HEAT-like Repeats. Journal of Molecular Biology. 2010;404:307–327. [PubMed] [Google Scholar]
  • Wetzel SK, Settanni G, Kenig M, Binz HK, Plückthun A. Folding and Unfolding Mechanism of Highly Stable Full-Consensus Ankyrin Repeat Proteins. Journal of Molecular Biology. 2008;376:241–257. [PubMed] [Google Scholar]
  • Wheeler TJ, Clements J, Finn RD. Skylign: a tool for creating informative, interactive logos representing sequence alignments and profile hidden Markov models. BMC Bioinformatics. 2014;15:7. [PMC free article] [PubMed] [Google Scholar]
  • Wowor AJ, Yu D, Kendall DA, Cole JL. Energetics of SecA Dimerization. Journal of Molecular Biology. 2011;408:87–98. [PMC free article] [PubMed] [Google Scholar]
  • Xu Y. Characterization of macromolecular heterogeneity by equilibrium sedimentation techniques. Biophysical Chemistry. 2004;108:141–163. [PubMed] [Google Scholar]
  • Yan N, Chai J, Lee ES, Gu L, Liu Q, He J, Wu JW, Kokel D, Li H, Hao Q, et al. Structure of the CED-4–CED-9 complex provides insights into programmed cell death in Caenorhabditis elegans. Nature. 2005;437:831–837. [PubMed] [Google Scholar]
  • Zhu H, Lee HY, Tong Y, Hong BS, Kim KP, Shen Y, Lim KJ, Mackenzie F, Tempel W, Park HW. Crystal Structures of the Tetratricopeptide Repeat Domains of Kinesin Light Chains: Insight into Cargo Recognition Mechanisms. PLoS ONE. 2012;7:e33943. [PMC free article] [PubMed] [Google Scholar]