This is an Open Access article distributed under the terms of the Creative Commons Attribution License (
The polypeptides involved in amyloidogenesis may be globular proteins with a defined 3D-structure or natively unfolded proteins. The first class includes polypeptides such as β2-microglobulin, lysozyme, transthyretin or the prion protein, whereas β-amyloid peptide, amylin or α-synuclein all belong to the second class. Recent studies suggest that specific regions in the proteins act as "hot spots" driving aggregation. This should be especially relevant for natively unfolded proteins or unfolded states of globular proteins as they lack significant secondary and tertiary structure and specific intra-chain interactions that can mask these aggregation-prone regions. Prediction of such sequence stretches is important since they are potential therapeutic targets.
In this study we exploited the experimental data obtained in an
The proposed method might become a useful tool for the future development of sequence-targeted anti-aggregation pharmaceuticals.
In the last decade, protein aggregation has moved beyond being a mostly ignored area of protein chemistry to become a key topic in medical sciences [
One of the major unanswered questions of protein aggregation is the specificity with which the primary sequence determines the aggregation propensity from totally or partially unfolded states. Deciphering the answer to this question will give us a chance to control the unwanted protein deposition events through specific sequence-targeted therapeutics. A first advance in this direction is the recent discovery that not all regions of a polypeptide are equally important for determining its aggregation tendency, both in natively unfolded and globular proteins. In this way, some authors, including ourselves, have proved recently that very short specific amino acid stretches can act as facilitators or inhibitors of amyloid fibril formation [
We have used a simple
The rationale behind our study is based on two recent observations in the field. First, not all the polypeptide sequence is relevant for the aggregation of a given protein, but rather there exist specific regions that drive the process [
Provided that a given polypeptide aggregates from an at least partially unstructured state, the experimental intrinsic aggregation propensities shown in Table
A good number of natural occurring mutations have been reported in proteins associated to depositional diseases. In many cases they result in changes in the global protein aggregation propensity and sometimes in the appearance of premature or acute pathological symptoms. The change in average aggregation propensity (ΔAP) between the wild type and the different mutants should predict the effect of sequence variations on the aggregation propensities, provided that they rely on changes in the intrinsic polypeptide properties.
In this section the above described analysis is applied to a set of proteins linked to depositional diseases and the obtained results are compared with the available experimental data.
As a proof of principle our approach was first tested in the molecule from which the experimental amino acid aggregation propensities were derived. Alzheimer's disease (AD) is a progressive neurodegenerative disorder characterized by the patient's memory loss and impairment of cognitive abilities. The extracellular amyloid is found in the brain and is widely believed to be involved in the progression of the disease [
A set of mutations in the CHC and adjacent positions of Aβ42 is intimately associated to early-onset familial Alzheimer diseases (FAD). The substitutions include A21G (Flemish), E22Q (Dutch) and E22G (Arctic) [
Adding to the mutations present in the population, a large set of mutations has been artificially introduced on Aβ that result in changes in its aggregation propensity. ΔAP values were also calculated for several of them and the results compared with the experimental data (Table
Type II diabetes is associated with progressive beta-cell failure manifested as a decline in insulin secretion and increasing hyperglycemia. A growing body of evidence suggests that beta-cell failure in type II diabetes correlates with the formation of pancreatic islet amyloid. Islet amyloid polypeptide (IAPP, amylin), the major component of islet amyloid, is co-secreted with insulin from beta-cells. In type II diabetes, this peptide aggregates to form amyloid fibrils that are toxic to beta-cells [
The analysis also explains the available mutational data on IAPP. Diabetes-associated IAPP amyloid occurs in primates and cats but not in rodents [
Several mechanisms have been proposed for IAPP fibril formation in type II diabetes. One widely accepted mechanism is that in type II diabetes, increased production and secretion of IAPP associated with increased demand for insulin might result in accumulation and aggregation of IAPP [
Parkinson disease is the most common neurodegenerative movement disorder and is pathologically characterized by the presence of neuronal intracytoplasmatic deposits of aggregated protein called Lewy bodies [
Several α-Synuclein mutations appear associated with familial early-onset Parkinson Disease: A30P, A53T and E46K. All they map into our predicted second "hot spot". The rates of fibril assembly of the E46K and A53T mutants have been shown to be greater than those of the wild type and A30P proteins [
β2-Microglobulin-related amyloidosis is a common and serious complication in patients on longterm hemodialysis [
In contrast to the human protein, mouse β2-m does not form fibrils even at high concentration [
Overall, our predictions on the presence and location of "hot spots" in β2-m are extremely accurate and overlap with the experimentally found relevant regions (Fig.
One of the most urgent issues in the study of amyloid fibrils is to reproduce the formation of fibrils under physiological conditions. Recently, it has been found that low concentrations of SDS around the critical micelle concentration induce the extensive growth of β2-m amyloid fibrils at physiological pH, probably through the SDS-induced conformational change of β2-m monomers [
Human lysozyme has been shown to form amyloid fibrils in individuals suffering from nonneuropathic systemic amyloidosis. The disease is always associated to point mutations in the lysozyme gene and fibrils are deposited widely in tissues [
Transthyretin (TTR) is a homotetramer of 127-amino acid subunits. TTR is found in human plasma and cerebral spinal fluid, the plasma form being the amyloidogenetic precursor. TTR constitutes the fibrillar protein found in familial amyloidotic polyneuropathy (FAP) and senile systemic amyloidosis (SSA) [
To date two different fragments of TTR have been shown to form amyloid fibrils. The peptide 105–115 can be assembled into homogeneous amyloid fibrils with favorable spectroscopic properties [
Misfolded isoforms of the naturally occurring prion protein (PrP) have been shown to be the causative agents in many mammalian neurodegenerative disorders, including Cruetzfeldt-Jakob disease (CJD) in human, scrapie in sheep, and bovine spongiform encephalopathy in cows. Prion infectivity is unique in that the pathogenic prion form (PrPSc) is involved in the conversion of the endogenous conformation (PrPC) into transformed PrPSc. The "protein-only" hypothesis [
The normal prion protein (PrPC) is a GPI-anchored glycoprotein constitutively expressed on the surface of primarily neuronal cells. It consists of two structurally different parts; a C-terminal, globular part mainly α-helical in nature (Fig.
The role of the detected aggregation-prone sequence at the N-terminus is uncertain since it is out of the protease resistant core of PrPSc. Little information exits about the role of this region, although it appears to be unnecessary both for prion transmission and aggregation. The predicted C-terminal "hot spot" includes almost all the C-terminal α-helix, named C, from the globular domain (Fig.
The central region of PrPC linking the unstructured N-terminal part with the globular C-terminal domain is believed to play a pivotal role in the PrPC conformational changes. Extensive studies on the secondary structure and fibrillogenic properties of synthetic peptides of PrP have established that the continuous segment of the prion protein spanning residues 106–147, coincident with the second "hot spot" predicted using our approach, is important for the fibrillogenic properties of the protein [
Overall, the method described here appears as a useful tool for the identification of protein regions that are especially relevant for protein aggregation and amyloidogenesis both in natively unfolded and properly folded globular proteins (Table
Nature has provided globular proteins with a reasonable conformational stability in the native state in which, as proved here, aggregation-prone sequences are buried or involved in intra-molecular interactions. This appears as a very successful evolutive strategy to avoid aggregation, since few proteins aggregate from their stable native conformation. Accordingly, amyloid-related mutations in globular proteins usually result in destabilization of the folded state allowing the exposure of previously hidden "hot spots", as those reported here. This explains the scarce success in predicting the effect of mutations in the aggregation of globular proteins (data not shown), whereas the prediction of fatal sequence changes in intrinsically unstructured proteins involved in disease is generally accurate. The effects of such mutations can be explained in most cases by intrinsic factors, as they directly result in changes on the average propensity of the full polypeptide to aggregate.
Besides providing important clues about the mechanism of protein aggregation, this study may be relevant for the therapeutics of amyloid disease, since the identified "hot spots" could be regarded as preferential targets to tackle the deleterious disorders linked to protein deposition. According to our results, different specific strategies should be employed when designing methods to avoid aggregation, depending on the disease being caused by natively unfolded or by globular proteins. In Alzheimer, type II diabetes and Parkinson diseases, shielding the already exposed aggregation-prone regions in the polypeptides by using small compounds or antibodies appears as a promising approach, whereas compounds that will stabilize the native conformation and avoid the exposure of the deleterious "hot spots" will be more effective in the case of globular proteins. Additionally, when gene therapy eventually comes to age, mutations that disrupt aggregation-prone regions in unstructured polypeptides or those which over-stabilize the native state of globular aggregation-prone proteins are expected to be useful approaches to avoid protein deposition and meliorate neurodegenerative and systemic amyloidogenic disorders.
The CHC of Aβ42 peptide was chosen as a paradigmatic aggregation-prone region for the calculation of the individual effect of each natural amino acid on protein aggregation. The specific effect on Aβ42's deposition promoted by the 20 different natural amino acids when located in the central position of this model "hot spot" were evaluated. Briefly, the wild type Aβ42 gene and its 19 mutants were inserted as a fusion protein upstream of the green fluorescence protein (GFP) and expressed individually in bacteria. In this system, the levels of GFP fluorescence in the cells depend exclusively on the
Different experimental data suggest that the aggregation of Aβ42 occurs from a mostly unfolded conformation in which the CHC is exposed to solvent [
The concept of "hot spot" of aggregation implies that the contribution of a particular residue in a protein sequence on protein aggregation is somehow modulated by its immediate neighbors. According to this, the effects of mutation on protein aggregation can not be properly calculated by a simple subtraction of the intrinsic aggregation propensities of the wild type and mutant residues. Instead, to provide a more general description of the effect of the change on the overall aggregation propensity, the individual aggregation profiles for the wild type protein and the different mutants are obtained and the differences between the areas below the corresponding profiles are calculated. The area between each profile was always normalized by the number of residues in the considered species to compare between the aggregation propensities of the complete protein and fragments coming from proteolysis, chemical synthesis or other processes. The difference between normalized areas, multiplied by a 100 factor, was designed as the change in average aggregation propensity (ΔAP). ΔAP will be positive if the mutation is predicted to increase the aggregation propensity of the polypeptide chain and negative if it is predicted to increase solubility.
AD Alzheimer's disease
Aβ Amyloid-β-protein
CHC Central hydrophobic cluster
FAD Familial Alzheimer diseases
FAP Amyloidotic polyneuropathy
GFP Green fluorescent protein
IAPP Islet amyloid polypeptide
NAC Non-Aβ component of amyloid plaques
PrP Prion protein
PrPSc Pathogenic prion form
SSA Senile systemic amyloidosis
TTR Transthyretin
ΔAP Change in average aggregation propensity
β2-m β2-Microglobulin
NSG and IP performed most of the experiments and prepared the final data and figures. FXA and JV contributed to data interpretation and manuscript redaction. SV directed the work and prepared the manuscript.
The vector expressing the Aβ42-GFP fusion was a generous gift of Michael Hecht's group. This work has been supported by Grants BIO2001-2046 and BIO2004-05879 (Ministerio de Ciencia y Tecnología, MCYT, Spain), by the Centre de Referència en Biotecnologia (Generalitat de Catalunya, Spain) and by PNL2004-40 (Universitat Autònoma de Barcelona (UAB). S.V. is supported by a "Ramón y Cajal" project awarded by the MCYT and co-financed by the UAB.
Relative experimental aggregation propensities of the 20 natural amino acids derived from the analysis of mutants in the central position of the CHC in amyloid-β-protein.
| Amino acid | |
| I | 1.822 |
| F | 1.754 |
| V | 1.594 |
| L | 1.380 |
| Y | 1.159 |
| W | 1.037 |
| M | 0.910 |
| C | 0.604 |
| A | -0.036 |
| T | -0.159 |
| S | -0.294 |
| P | -0.334 |
| G | -0.535 |
| K | -0.931 |
| H | -1.033 |
| Q | -1.231 |
| R | -1.240 |
| N | -1.302 |
| E | -1.412 |
| D | -1.836 |
Comparison of predicted and experimental changes in aggregation for Aβ variants.
| Mutation | ΔAP* | Observed aggregation‡ |
| A21G | -1.22 | - |
| E22G | +2.14 | + |
| E22Q | +0.44 | + |
| F19P | -5.09 | - |
| F19T | -4.73 | - |
| I31L | -1.07 | - |
| I32L | -1.07 | - |
| I41G | -5.76 | - |
| I41A | -4.52 | - |
| I41L | -1.075 | - |
| A42G | -1.21 | - |
| A42V | +3.95 | + |
| Δ1–4 | +4.82 | + |
| Δ1–9 | +21.86 | + |
| Δ40–42 | -8.34 | - |
| Δ41–42 | -4.26 | - |
| V12E+V18E+M35T+I41N | -18.96 | - |
| F19S+L34P | -9.18 | - |
* Change in average aggregation propensity
‡ Changes in aggregation determined experimentally.
Comparison of predicted and experimental changes in aggregation for IAPP variants, relative to the corresponding human IAPP sequence.
| Variant | ΔAP* | Observed aggregation‡ |
| (20–29) Cat | -5.12 | = |
| (20–29) Rat | -16.46 | - |
| (20–29) Hamster | -32.73 | - |
| R18H | +0.94 | + |
| L23F | +1.70 | + |
| V26I | +0.42 | + |
| R18H+L23F+V26I | +3.06 | + |
| (22–27) N22A | +31.53 | + |
| (22–27) F23A | -42.96 | - |
| (22–27) G24A | +11.96 | + |
| (22–27) I26A | -44.5872 | - |
| (22–27) L27A | -33.99 | - |
| S20G | -1.09 | + |
| ProIAPP | +31.40 | +? |
* Change in average aggregation propensity
‡ Changes in aggregation determined experimentally.
? Not yet proved experimentally.
List of the predicted "hot spots" in the different disease-linked polypeptides in this study and comparison with the available experimental data. Experimental "hot spots" refer to those protein regions shown to be involved in the aggregation process of the corresponding polypeptide. It is also noted if the predicted "hot spot" has been described as a structural element of the amyloid fibrils formed by the different peptides and proteins in the study.
|
|
|
|
|
|
|
16–21 | + | + |
| 30–36 | + | + | |
| 38–42 | + | + | |
|
|
12–18 | + | uncertain |
| 22–28 | + | uncertain | |
| 1–18 | No experimental data available | uncertain | |
|
|
27–56 | + | uncertain |
| 61–94 | + | + | |
|
|
21–31 | + | + |
| 56–69 | + | + | |
| 79–85 | + | + | |
| 87–91 | + | + | |
|
|
24–34 | - | - |
| 50–62 | + | + | |
| 76–98 | + | + | |
|
|
10–20 | + | + |
| 23–33 | No experimetal data available | uncertain | |
| 105–118 | + | + | |
|
|
1–32 | No experimetal data available | uncertain |
| 105–146 | + | + | |
| 208–252 | No experimetal data available | uncertain |