Conceived and designed the experiments: SV. Performed the experiments: VC SV. Analyzed the data: VC SV. Wrote the paper: SV.
Protein aggregation underlies a wide range of human disorders. The polypeptides involved in these pathologies might be intrinsically unstructured or display a defined 3D-structure. Little is known about how globular proteins aggregate into toxic assemblies under physiological conditions, where they display an initially folded conformation. Protein aggregation is, however, always initiated by the establishment of anomalous protein-protein interactions. Therefore, in the present work, we have explored the extent to which protein interaction surfaces and aggregation-prone regions overlap in globular proteins associated with conformational diseases. Computational analysis of the native complexes formed by these proteins shows that aggregation-prone regions do frequently overlap with protein interfaces. The spatial coincidence of interaction sites and aggregating regions suggests that the formation of functional complexes and the aggregation of their individual subunits might compete in the cell. Accordingly, single mutations affecting complex interface or stability usually result in the formation of toxic aggregates. It is suggested that the stabilization of existing interfaces in multimeric proteins or the formation of new complexes in monomeric polypeptides might become effective strategies to prevent disease-linked aggregation of globular proteins.
The aggregation of proteins in tissues is associated with the pathogenesis of more than 40 human diseases. The polypeptides underlying disorders such as Alzheimer's and Parkinson's are devoid of any regular structure, whereas the polypeptides causing familial amyotrophic lateral sclerosis or nonneuropathic systemic amyloidosis correspond to globular proteins. Little is known about the mechanism by which globular proteins under physiological conditions aggregate from their initially folded and soluble conformations. Interestingly, several of these pathogenic proteins display quaternary structure or are bound to other proteins in their physiological context. In the present work, we show that protein-protein interaction surfaces and regions with high aggregation propensity significantly overlap in these polypeptides. This suggests that the formation of native complexes and self-aggregation reactions probably compete in the cell, explaining why point mutations affecting the interface or the stability of the protein complex lead in many cases to the formation of toxic aggregates. This study proposes general strategies to fight against diseases associated with the deposition of globular polypeptides.
The formation of insoluble amyloid protein deposits in tissues is related to the development of more than 40 different human diseases, many of which are debilitating and often fatal. The polypeptides responsible for these disorders are not related in terms of sequence or conformation
Although the study of protein aggregation from non-native states has provided a wealth of data on the physico-chemical determinants of amyloid formation, little is known about how globular proteins aggregate from their initially folded and soluble conformations under physiological conditions, where extensive unfolding is not expected to occur
Protein aggregation can be seen as an anomalous type of protein-protein interaction. In functional interactions, binding partners come together in a stable and precise orientation in seconds
The prediction of regions responsible for aggregation based on the primary sequence of a protein has been tackled by several methods, from simple considerations of the properties of amino acids to complex molecular dynamics calculations
Identification of binding sites in polypeptides is a direct computational approach to deciphering biological and biochemical function. Although sequence-based approaches to identifying protein interfaces exist, their results are often unsatisfactory. Here, we have used three different structure-based methods whose algorithms are publicly available as web servers to produce a consensus prediction of the interaction interfaces in the globular proteins under consideration (see
Two levels of prediction were considered: i) residues predicted or shown to be both in aggregation-prone regions and at interfaces and ii) residues in aggregation-prone sequences that are close in space to the interaction surface (below 3 Å). The interaction predictions were compared with the experimentally determined contacts in the quaternary structure of the proteins or in complexes of the studied proteins with other polypeptides. The regions predicted to have high aggregation propensity were compared with fragments of the analyzed proteins shown experimentally to form amyloid aggregates or to be located in the core of the mature fibrils formed by these polypeptides. We have defined a parameter called Interface Proximity Index (IPI) to evaluate the degree to which an aggregation-prone region is closer to a given interface than to the rest of the protein surface (see
Aggregation-prone regions are coloured according to their IPI values (see the scale). A) β2-microglobulin, B) transthyretin, C) immunoglobulin G heavy chain, D) SOD1 and E) immunoglobulin light chain variable domain.
Amyloidosis related to β2-Microglobulin (β2-m) is a common and serious complication in patients on long-term hemodialysis
In all panels, β2-microglobulin aggregation-prone residues at less and more than 3 Å from interaction sites are shown in red and green, respectively. Interface residues not included in aggregation-prone regions are shown in dark blue. Rest of residues are shown in light blue. A) The predicted interaction surface for monomeric β2-microglobulin is used for calculation. B) The interface between β2-microglobulin and HLA heavy chain is used for calculation (PDB ID:1DUZ). C) The interface between β2-microglobulin and HFE is used for calculation (PDB ID:1A6Z). D and E) Front (same orientation that in B) and back view of the β2-microglobulin/HLA heavy chain complex. F and G) Front (same orientation that in C) and back view of the β2-microglobulin/HFE complex.
| Predicted Aggregation segments | Fibril formers | % residues close to predicted Interface |
% residues close to real interface |
IPI | % solvent accessible residues |
|
|
|||||
| 22–31 | 21–31 | 70 (20) | 100 (18) | 0.88 | 65 |
| 21–41 | |||||
| 60–70 | 59–79 | 54 (24) | 73 (30) | 0.58 | 100 |
| 59–71 | |||||
|
|
|||||
| 22–31 | 21–31 | 70 (20) | 70 (12) | 0.83 | 65 |
| 21–41 | |||||
| 60–70 | 59–79 | 54 (24) | 81 (31) | 0.62 | 100 |
| 59–71 | |||||
|
|
|||||
| 11–19 | 10–20 | 44 (22) | 77 (36) | 0.46 | 66 |
| 26–34 | - | 0 (53) | 0 (68) | <0 | 66 |
| 92–96 | - | 40 (42) | 100 (0) | 1 | 100 |
| 105–112 | 105–115 | 62 (30) | 100 (40) | 0.6 | 100 |
| 115–121 | - | 62 (32) | 100 (0) | 1 | 100 |
|
|
|||||
| 4–8 | - | 0 (13) | 100 (3) | 0.97 | 66 |
| 100–106 | - | 43 (12) | 0 (19) | <0 | 66 |
| 111–120 | - | 70 (18) | 50 (14) | 0.72 | 100 |
| 146–153 | - | 87 (21) | 87 (2) | 0.97 | 100 |
|
|
|||||
| 25–33 | - | 33 | 0 (25) | <0 | 44 |
| 57–66 | - | 90 | 40 (32) | 0.20 | 60 |
| 76–84 | 26–123 | 56 | 56 (25) | 0.55 | 88 |
| 108–114 | - | 86 | 0 (46) | <0 | 100 |
|
|
|||||
| 19–23 | - | - | 0 (38) | <0 | 71 |
| 31–38 | - | - | 89 (24) | 0.71 | 75 |
| 46–51 | - | - | 50 (30) | 0.4 | 71 |
| 71–78 | - | - | 0 (43) | <0 | 62 |
| 84–89 | - | - | 83 (18) | 0.78 | 66 |
|
|
|||||
| 29–38 | - | - | 50 (20) | 0.60 | 80 |
| 45–52 | - | - | 75 (25) | 0.67 | 82 |
| 87–93 | - | - | 57 (24) | 0.58 | 57 |
| 100–106 | - | - | 100 (16) | 0.84 | 100 |
| 275–281 | - | - | 0 (31) | <0 | 71 |
| 289–299 | - | - | 0 (52) | <0 | 100 |
| 322–331 | - | - | 0 (37) | <0 | 80 |
| 390–396 | - | - | 63 (13) | 0.79 | 57 |
| 435–442 | - | - | 100 (16) | 0.84 | 75 |
Percentage of residues in the aggregation-prone region at less than 3 Å from a protein predicted interaction residue.
Percentage of residues in the aggregation-prone region at less than 3 Å from a residue located at the interface of the following complexes: β2-microglubulin in complex with HLA heavy chain [1DUZ] and with HFE [1A6Z]. Native tetrameric structure of transthyretin (PDB code 1TTA). Dimeric structure of SOD1 (PDB code 2C9V). Lysozyme in complex with a camelid antibody (PDB code 1OP9). Dimeric structure of Immunoglobulin LC variable domain (PDB code 2Q20). HCs and LCs of a IgG1 human immunoglobulin (PDB code 1HZH).
In brackets the percentage of residues in the aggregation-prone region close to a random surface of the same size than the considered interface.
A main interaction cluster is predicted for human β2-m (
Class I major-histocompatibility-complex (MHC) molecules (HLA molecules in humans) are ternary complexes of β2-m, an MHC heavy chain, and a bound peptide
Inside the cell, β2-m associates with the non-classical HLA class I molecule human hemochromatosis protein (HFE)
Aggregation of β2-m under physiological conditions is thought to be initiated by a cis-trans prolyl isomerization of the H31-P32 peptide bond
Because the β2-m regions likely to be involved in aggregation are already located in preformed β-strands, local fluctuations may allow anomalous intermolecular interactions between these preformed elements, leading to the formation of an aggregated β-sheet structure without extensive unfolding. In this context, the formation of β2-m complexes both inside the cell and on the cell surface might play a protective role against β2-m aggregation, either by reducing conformational fluctuations or by preventing the exposure of dangerous amyloidogenic regions, or both.
Transthyretin (TTR) constitutes the fibrillar protein found in familial amyloidotic polyneuropathy (FAP), familial amyloidotic cardiomyopathy, and central nervous system amyloidosis. Around 100 different TTR mutations have been reported, many of which are amyloidogenic
A single interaction patch is predicted for the TTR monomer (
In all panels, transthyretin (TTR) aggregation-prone residues at less and more than 3 Å from interaction sites are shown in red and green, respectively. Interface residues not included in aggregation-prone regions are shown in dark blue. Rest of residues are shown in light blue. A) The predicted interaction surface of a TTR monomer is used for calculation. B) The interface in the native tetrameric structure of TTR is used for calculation (PDB ID:1TTA). C) Dimer of TTR. D) TTR native tetrameric structure. The first dimer is twisted 90° relative to C, the second one is shown in yellow.
The crystal structure of the TTR tetramer (PDB ID: 1TTA)
Dissociation of the TTR tetramer has been reported as a prerequisite for amyloidosis. The tetrameric structure dissociates into AB and CD dimers, but they are unstable in the absence of additional quaternary interactions, explaining why TTR exists in a primarily tetramer-monomer equilibrium
Familial amyotrophic lateral sclerosis (fALS) is characterized by the presence of Copper-Zinc Superoxide Dismutase (SOD1) inclusions in spinal cords
A total of 14 residues are predicted to be at the interface of the SOD1 monomer (
In panels A, B and C SOD1 aggregation-prone residues at less and more than 3 Å from interaction sites are shown in red and green, respectively. Interface residues not included in aggregation-prone regions are shown in dark blue. Rest of residues are shown in light blue. A) The predicted interaction surface of a SOD1 monomer is used for calculation. B) The interface in the native dimeric structure of SOD1 is used for calculation (PDB ID:2C9V). C) Native dimer of SOD1, the second monomer is shown in yellow. D) Ribbon representation of the SOD1 dimer, predicted aggregation-prone regions are shown in red.
According to the crystal structure of the SOD1 dimer (PDB ID: 2C9V)
FALS has been shown to be associated with more than 100 different SOD1 mutations, which are scattered throughout the three-dimensional structure
The light chains (LCs) of immunoglobulins have been implicated in the pathogenesis of amyloidosis in patients with monoclonal B-cell proliferative disorders (AL amyloidosis)
In all panels, immunoglobulin (Ig) aggregation-prone residues at less and more than 3 Å from interaction sites are shown in red and green, respectively. Interface residues not included in aggregation-prone regions are shown in dark blue. Rest of residues are shown in light blue. A) The interface in the native structure of Ig light chain variable domain (LC) is used for calculation (PDB ID: 2Q20). B) Native homodimer of Ig LC, the second monomer is shown in yellow. C) The interface in the native structure of IgG heterotetramer is used for calculation and the Ig heavy chain (HC) represented (PDB ID: 1HZH). D) Native IgG heterotetramer. Ig LCs and the second Ig HC are indicated.
AL is distinct from other types of amyloidosis in that hypervariability yields a different set of mutations in each patient. Ramirez-Alvarado and co-workers have characterized an LC dimer isolated from an AL patient
Although AL is more frequent, in some systemic amyloidosis the amyloid deposits consist of an unusual form of IgG1 heavy chain (HC)
Using the structure of a complete human IgG1 antibody
Human lysozyme forms amyloid fibrils in individuals suffering from nonneuropathic systemic amyloidosis. The disease is always associated with non-conservative point mutations in the lysozyme gene
Two different interaction clusters are predicted for human lysozyme (
In all panels, aggregation-prone residues at less and more than 3 Å from interaction sites are shown in red and green, respectively. Interface residues not included in aggregation-prone regions are shown in dark blue. Rest of residues are shown in light blue. A) The predicted interaction surface of lysozyme is used for calculation. B) The interface between lysozyme and a camelid antibody is used for calculation (PDB ID: 1OP9). C) Lysozyme complex with a camelid antibody. D) Ribbon representation of Aß peptide. The interface between the peptide and a designed affibody is used for calculation (PDB ID: 2OTK). E) Aß peptide bound to a designed affibody.
The mechanism of lysozyme aggregation under physiological conditions probably involves thermal fluctuations that transiently expose amyloidogenic regions
A single-domain fragment of a camelid antibody has been shown to inhibit the
A nice example illustrating how new binding interfaces can effectively inhibit amyloid formation has been recently reported for the Alzheimer's Aβ peptide. Two aggregation-prone regions comprising residues 16–21 and 29–40 are consistently predicted for Aβ (
An important question to address is whether predicted interaction interfaces and aggregation-prone regions also coincide in monomeric and soluble proteins. Therefore, we have analyzed the predicted properties of four well-characterized soluble proteins: myoglobin, maltose binding protein, thioredoxin, and ubiquitin.
Human myoglobin is a compact protein not related to disease. Although after long exposure to high temperatures
In panels A–D, aggregation-prone residues at less and more than 3 Å from interaction sites are shown in red and green, respectively. Interface residues not included in aggregation-prone regions are shown in dark blue. Rest of residues are shown in light blue. In all panels the predicted interaction surface is used for calculation. A) Human myoglobin (PDB ID: 4MBN). B) Maltose Binding Protein (MBP) (PDB ID: 4MBP). C) Human thioredoxin (TRX) (PDB ID: 3TRX). Gatekeeper residues are shown in purple and active cysteines in yellow. D) Same orientation that in C, Human TRX in a mixed disulfide intermediate complex with a peptide from the transcription factor NF kappa B (PDB ID: 1MDI). E) Ribbon representation of human ubiquitin (PDB ID: 1UBQ). Aggregation-prone secondary structures near the interface are shown in red. Basic residues in the vicinity of aggregation-prone regions are shown in purple. F) Same orientation than E). Complex of human ubiquitin with a CUE ubiquitin binding domain (PDB ID: 1OTR).
Maltose binding protein (MBP) endows fused proteins with increased solubility indicating that it is by itself highly soluble
Thioredoxin A (TRX) is another tag used to increase the solubility of recombinant proteins
The question arises of why TRX does not self-assemble when it is free. It appears that evolution uses negative design to fight against protein deposition by placing amino acids that counteract aggregation at the flanks of protein sequences with high aggregation propensity
Ubiquitin is a small, soluble and highly conserved regulatory protein that is ubiquitously expressed in eukaryotes
It seems that the spatial coincidence of interfaces and sequences promoting self-assembly is not restricted to amyloidogenic proteins. To further confirm this extent, we analyzed the structure of 25 different eukaryotic proteins shown to form homodimers (
Aggregation-prone regions in which more than 85% of the residues are at less than 3 Å from the interface are highlighted in green. The PDB ID is indicated for each dimer (see also
| PDB | Protein | Source | Length | Aggregation segments | Aggregation segments close to the interfase (>50%) |
Aggregation segments close to the interfase (>85%) |
| 1F17 | Dehydrogenase | Homo sapiens | 293 | 9 | 2 | 2 |
| 1DQT | Antigen | Mus musculus | 117 | 8 | 4 | 3 |
| 1LR5 | Auxin binding protein | Zea mays | 160 | 7 | 4 | 1 |
| 1KSO | Calcium-binding protein A3 | Homo sapiens | 93 | 3 | 2 | 2 |
| 1EAJ | Coxsackie virus | Homo sapiens | 124 | 4 | 2 | 1 |
| 1PE0 | DJ-1 | Homo sapiens | 187 | 7 | 1 | 1 |
| 1JR8 | Erv2 protein mitochondrial | Saccharomyces cerevisiae | 105 | 5 | 2 | 2 |
| 1F4Q | Grancalcin | Homo sapiens | 161 | 6 | 3 | 1 |
| 1DQP | Guanidine phosphoribosyltransferase | Giardia lamblia | 230 | 10 | 3 | 2 |
| 3SDH | Hemoglobin | Scapharca inaequivalvis | 145 | 5 | 2 | 2 |
| 2HHM | Hydrolase | Homo sapiens | 272 | 11 | 4 | 3 |
| 8PRK | Inorganic pyrophosphatase | Saccharomyces cerevisiae | 282 | 8 | 3 | 2 |
| 1QMJ | Lectin | Gallus gallus | 132 | 5 | 1 | 1 |
| 1M6P | Phosphate receptor | Bos taurus | 146 | 5 | 2 | 1 |
| 1MNA | Polyketide synthase | Streptomyces venezuelae | 276 | 10 | 2 | 1 |
| 1F89 | Protein YLC351C | Saccharomy cerevisiae | 271 | 11 | 3 | 2 |
| 1LHP | Pyridoxal kinase | Ovis aries | 306 | 10 | 3 | 1 |
| 1QR2 | Quinone reductase type 2 | Homo sapiens | 230 | 9 | 5 | 1 |
| 3LYN | Sperm lysine | Haliotis fulgens | 122 | 6 | 3 | 1 |
| 1SCF | Stem cell factor | Homo sapiens | 116 | 4 | 1 | 0 |
| 1HQO | URE2 protein | Saccharomyces cerevisiae | 221 | 8 | 3 | 3 |
| 1HSS | Alpha-amylase inhibitor | Triticum aestivum | 111 | 3 | 2 | 2 |
| 1KIY | Trichodiene synthase | Fusarium sporotrichioides | 354 | 12 | 4 | 4 |
| 1MI3 | Xylose reductase | Candida tenuis | 319 | 6 | 1 | 1 |
| 1LBQ | Ferrochelatase | Saccharomyces cerevisiae | 356 | 12 | 3 | 2 |
More than 50% of the residues in the aggregation-prone region are at less than 3 Å from a residue located at the interface of the complex.
More than 85% of the residues in the aggregation-prone region are at less than 3 Å from a residue located at the interface of the complex.
During the revision of the present work, Vendruscolo and co-workers published a related study in which they used their algorithm Zyggregator to perform an extensive analysis of interfaces in protein-protein complexes
In the present work, we have used computational tools to predict aggregation-prone regions and interaction sites in globular proteins related to depositional diseases and non-pathogenic polypeptides. From the comparison of the predictions with the structural and experimental data, it appears that protein-protein interaction surfaces and regions with high aggregation propensity overlap significantly in the quaternary structure of proteins.
The proximity and coincidence of protein-protein interfaces and aggregation-prone regions suggests that the formation of native complexes and the aggregation of their monomeric subunits probably compete in the cell. This implies that the molecular machinery that performs the vast array of cellular functions and the aggregates that might interfere with these functions promoting cell stress or even cell death are sustained by similar molecular contacts. It is likely that the specificity of native protein interfaces in protein complexes has evolved to minimize anomalous interactions and therefore detrimental protein aggregation reactions. In this sense, Vendruscolo and co-workers have recently identified disulfide bonds and salt bridges as specific interactions that can stabilize aggregation-prone interfaces in their native conformations in oligomeric proteins
Overall, the present analysis provides a rational to understand how globular proteins aggregate under physiological conditions, where they posses an initially folded and cooperatively sustained conformation and extensive denaturation is not expected to occur. The data strongly suggest that the stabilization of the interface in multimeric proteins, as in the case of TTR, SOD1, or LC immunoglobulins, and/or the blocking of conformational fluctuations and exposed amyloidogenic regions through the formation of new interfaces with other protein molecules, as in the case of lysozyme or Aß peptide, might be important strategies to delay the onset or slow the progress of conformational diseases caused by globular proteins.
The observed association between the failure to attain a native interface and the build up of harmful aggregates suggests that the range of genetic human diseases which ultimately might originate from the conversion of a soluble globular protein into toxic assemblies could be much larger than previously thought. Approaches combining the prediction of aggregation-prone regions from the linear protein sequence with the analysis of real or predicted protein interfaces in the 3D-structure might provide a means to identify physiologically and therapeutically relevant amyloidogenic sequences in the proteins linked to such disorders.
Aggregation-prone regions in the studied proteins were predicted using the primary sequence as input and a consensus of the output of four different available methods. The first algorithm we used is TANGO (
Interaction residues were predicted using the monomeric three-dimensional crystal structure of each of the studied proteins as input and a consensus of the output of three different algorithms. The first approach used to predict interaction surfaces was the Optimal Docking Area (ODA) method (
To evaluate whether the proximity of an aggregation-prone region to a given real interface is specific or the sequence stretch is as close to any other patch of the same size in the protein surface, we have defined the Interface Proximity Index: IPI
Each random surface was generated by an aleatory selection of a number of solvent exposed residues equal to the number of residues constituting the real interface. One hundred random surfaces were generated for each aggregation-prone region analyzed.
Solvent-accessible and buried residues in the monomeric complex subunits where identified using the PISA server at the European Bioinformatics Institute (
An
Figures were were generated with the Swiss-PDB viewer program (
We thank Daniel Fernandez for the help with the ODA analysis and with the first draft of the present manuscript.
The authors have declared that no competing interests exist.
This work has been supported by grant BIO2007-68046 (Ministerio de Ciencia e Innovacion, Spain). The funders had no role in study design, data collection and analysis, decision to publish, or preparation of the manuscript.