Conceived and designed the experiments: SB FN. Performed the experiments: SB. Analyzed the data: SB FN. Wrote the paper: SB.
Biological function relies on the fact that biomolecules can switch between different conformations and aggregation states. Such transitions involve a rearrangement of parts of the biomolecules involved that act as dynamic domains. The reliable identification of such domains is thus a key problem in biophysics. In this work we present a method to identify semi-rigid domains based on dynamical data that can be obtained from molecular dynamics simulations or experiments. To this end the average inter-atomic distance-deviations are computed. The resulting matrix is then clustered by a constrained quadratic optimization problem. The reliability and performance of the method are demonstrated for two artificial peptides. Furthermore we correlate the mechanical properties with biological malfunction in three variants of amyloidogenic transthyretin protein, where the method reveals that a pathological mutation destabilizes the natural dimer structure of the protein. Finally the method is used to identify functional domains of the GroEL-GroES chaperone, thus illustrating the efficiency of the method for large biomolecular machines.
The mechanical properties of biomolecules and their complexes are essential to molecular function, because many molecular processes are accompanied by conformational changes, in which domains of the molecule must be able to move with respect to each other
The first step to analysis and simulation of molecular nanomechanics is the identification of the rigid and flexible parts of biomolecules in different chemical, conformational or aggregate states considered. Conventional experimental techniques, like for example nuclear magnetic resonance (NMR), provide limited information about these processes.
One approach to identify the rigid and flexible parts in biomolecules is to partition the system into domains (also called “groups” or “clusters” in other works) that are nearly rigid. In the coarse-grained model, these domains can only move as a rigid body with six degrees of freedom (3 translation +3 rotation). Such a low dimensional model of the original high-dimensional dynamics yields itself easily to the understanding of essential mechanical properties of the molecule and how they change between conformations. Clearly, such a model only approximates the real mobility and the approximation error will depend on the number of domains considered and on the flexibility/rigidity of the molecule in the conformation considered. Consequently, such a model is better suited for describing functional transitions or aggregation than for processes involving much flexibility, such as folding.
Several methods for the identification of nearly rigid domains in biomolecules have been proposed that produce similar but not identical results. They can be categorized into model-based methods, where structural aspects such as hydrophobicity, topology, structural homology or for e.g. identical sequence motifs serve to identify the smallest building blocks
Data-based approaches in contrast define domains based on data of the flexibility of the biomolecule, such as MD simulations
Normal-mode-based techniques are limited by the fact that they only use local information of the energy landscape. PCA-based clustering methods do not suffer from this limitation, but still require all structures to be fitted to a mean or reference structure before calculating the covariance matrix. Such a fitting procedure works well as long as the structures are very similar, but if very large conformational changes are involved, then structures which are very similar to each other but very different from the reference structure may become very different after the fitting and thus produce a misleading covariance matrix. Thus, it is desirable to use a method that works with internal coordinates only. Moreover, there is a lack of the domain identification techniques that avoid ad-hoc assumptions and parameter choices that indirectly influence the number of clusters. It would be rather desirable to have an explicit control of the clustering error by adjusting the number of domains, or to have the method select the number of domains such that the clustering error is below a certain threshold.
The proposed method works by defining (i) a distance-deviation matrix between atoms based on dynamical data, (ii) formulating the clustering problem as a quadratic optimization problem that is based on this matrix and (iii) solving this clustering problem to optimality and obtaining an assignment of atoms to clusters. To illustrate the strengths and limitations of this approach a number of example systems are considered: two artificial peptides Ala5 and
The immediate use of the method is to understand dynamic processes in large macromolecules and their complexes which involve changes of molecular rigidity. This includes processes like conformational changes, ligand binding and protein aggregation
The principal objective of this work is to develop a new coarse-graining technique to partition large molecular systems
Optimal and unique molecular partitioning for given data and number of domains
Works with internal coordinates only and is thus independent of a reference structure
Can be applied to characterize models with multiple conformations without “overlooking” rarely populated conformations
Error measure for coarse-grain quality and ability to adjust the accuracy by the number of domains or the maximum acceptable clustering error
Simple applicability and robustness - no parameters other than number of domains
Model independent, so that experimental findings are easily incorporated
Efficient and simple implementation
Inter atomic distance-deviation is a common metric used for the identification of rigid domains in proteins
The analysis of local molecular rigidity is based on the distance deviation matrix
Most methods in the literature
We define the optimal partition of the molecule into domains as the one that minimized deviations within the domains:
In contrast to heuristic coarse-graining methods the minimization problem in Eq. 4 leads to an optimal partitioning of the molecule according to the number of domains chosen. Furthermore the partitioning has no bias towards equally-sized domains, i.e. it allows for domains of very different sizes if this is requested by the structure of
The minimization problem in Eq. 4 together with the normalisation condition in Eq. 2 can be written into a standard quadratic optimization problem with linear constraints that is solved here in order to identify the optimal partitioning into domains.
Because the Hessian matrix
The present quadratic optimization problem is solved using an active set method similar to that of Gill et al., described in
Besides the sparse definition of
Even though the method is robust for low
One approach to escape from local minima is applying stochastic methods such as Monte Carlo sampling, simulated annealing or genetic algorithms. Another simple approach that has shown to work well in practice is to use the solution obtained for
The error can be used as tool to choose the number of domains
Compute distance-deviation matrix,
Set
Compute optimal clustering
If
Else
The computational performance of the method was demonstrated for the examples discussed in the “
| System | no. Atoms | time in seconds | ||
| M = 2 | M = 5 | M = 81 | ||
| MR121-GSGSW | 81 | 0.015 | 0.12 | 14.46 |
| Transthyretin | 229 |
0.048 | 0.94 | 94.72 |
| Transthyretin | 2257 | 1.32 | 92.34 |
|
| GroEL-GroES | 8015 |
181.77 | 1537.01 |
|
Computation time for selected molecular systems with
To demonstrate the performance and usefulness of the method we have applied it to a series of molecular systems:
A
A
A
All molecular dynamics trajectories were generated by the molecular dynamics package Gromacs 3.3
We illustrate our approach on a number of test systems. In all cases the distance deviation matrix
Interestingly, although the method formally allows for fuzzy memberships, the optimal assignment of atoms to domains is always unique in practice, thus obtaining an exact partitioning of atoms into domains. For example, consider a hypothetical 3-atom system with the distance-deviation matrix
Now consider the case,
We conclude that no fuzzy memberships are found in macromolecules. Note, however, that the introduction of a fuzzy membership was still essential, because using this formulation we could express the optimization problem as a continuous quadratic optimization problem. The solution to this kind of problem is much easier than the solution to the integer optimization problem emerging by the priori assumption that the memberships must be integer values.
As a first example the optimization method was applied to Ala
The decrement of the clustering error is very steep for
In order to study a more complex system, the method was applied to the MR121-GSGSW peptide.
The decrement of the clustering error is very steep for
The corresponding distance-deviation matrix is shown in
The colors relate the semi-rigid regions in the distance deviation matrix to the molecular coarse-grain structure and the membership matrix.
The resulting membership matrix,
In order to test the optimality of the results, we have repeatedly solved the clustering problem for the MR121-GSGSW peptide using different initial conditions: (i) for given
When using a random assignment to clusters for the first step
The transport protein transthyretin (TTR) is primarily synthesized in liver, choroid plexus, and the retina. The primary function is the transport of thyroxine and retinol binding protein (RBP). Both molecules can bind to the homo-tetrameric structure of TTR, which is found at a physiological pH of
The obtained coarse-grain structures separate the dimer into two monomers (A) and (B) for
Transthyretin is one of the human proteins known to be associated with local amyloidosis. Amyloid fibrils are the polymerized form of the protein, their internal structure mainly consists of cross
Transthyretin aggregation to amyloid fibers has been the subject of many studies
To date, a large number of TTR variants have been associated with amyloid formation
In
To demonstrate the applicability to experimental data we used the method to partition molecular structures of transthyretin obtained by x-ray crystallography. Besides the wild type structure (PDB code 1DVQ), which was also used in the molecular dynamics simulation, five related structures of transthyretin complexed with resveratrol, diclofenac, flurbiprofen, DDBF, oFLU, and PHENOX (PDB codes 1DVS, 1DVT, 1DVU, 1DVX, 1DVY, 1DVZ)
As for molecular dynamics data the obtained coarse-grain structures separate the dimer into two monomers (A) and (B) for
The distance-deviation matrices in
The mean row value of each matrix indicates flexible regions around reference residues 11, 77, 105, 191 and 220. The corresponding structures are color coded according to the average row value of
The structural modification induced by the amyloidogenic variants (TTR-58Arg and TTR-58His) contribute to an de/increase in rigidity in some regions of the structure (see increasing red regions in the variants compared to the native TTR in
However in addition to the overall stability, increased atomic motion of specific regions in the dimer may influence the stability of the dimer and favor transient dissociation. The local flexibility/rigidity of atoms is reflected by
The existence of semi-rigid domains and their relative dynamics are essential for the functionality of large macromolecular machines. Here, we analyse the dynamics of the GroEL-GroES chaperone complex (see
The
Computational studies have provided important insights into the allosteric mechanism of the chaperonin GroEL-GroES. Protein folding within the complex involves binding, encapsulation, and release of the substrate protein
The cartoon representation of three important structures (
Due to the size of the system only a short molecular dynamics trajectory with duration of
To study GroEL in more detail, we have further performed domain identification on the subunit of the
The black lines indicate the identified domains boundaries between the apical (A - red), intermediate (I - green) and equatorial (E - blue) domain in the GroEL subunit. The mean value of each row and the color assignment are shown on the right.
The coarse-graining algorithm developed in this paper is an optimal and systematic approach to decompose ensembles of molecular structures into semi-rigid domains. It consists of three steps: (i) obtaining an ensemble containing the atomic fluctuations, e.g. using molecular dynamics simulation, (ii) computation of the pair distance-deviation matrix and (iii) definition of semi-rigid domains by a quadratic optimization method, to distinguish and to quantify the rigid and flexible domains within the protein structure. The method identifies rigid regions that can vary in size and shape. The objective function minimized in the procedure is a direct measure of the clustering error and thus the within-cluster flexibility neglected by assigning the atoms into domains. We have been able to study the rigidity of proteins in systems involving 8,015 residues on a normal desktop computer.
In contrast to other methods the algorithm does not require the choice of any parameters other than the number of domains. Being able to fix the number of domains is an advantage, since it gives the user a tool to decide how much flexibility he wants to resolve and to control the magnitude of the clustering error. A straightforward automatic way to select the number of domains is by requiring the clustering error to be below a specified threshold.
The coarse-graining algorithm has been applied to a number of benchmark problems. First, the consistency and error dependence of the method was demonstrated on two short peptides, by systematically increasing the number of domains
Since the clustering method proposed here is a data-based method, its result will depend on the quality of that data. In principle, the result will only be globally converged, if the underlying simulations have visited all relevant conformations within the data set according to the Boltzmann probability. However the GroEL-GroES results and other studies on very large systems such as viruses
Besides the robustness and reliability the method is easy to implement, efficient and useful in obtaining the essential nanomechanical properties of the molecule, we expect it to become a useful tool for the analysis of large-scale molecular systems.
As an outlook the method presented here can be used as a first step to generate a simulation model for large molecules or aggregates that can for example be simulated with Brownian dynamics. In addition to the identification of mobile domains this requires also the estimation of interaction forces and diffusion constants from either simulation or experimental data. This task is a subject of ongoing work.
The authors like to thank Bert de Groot for helpful suggestions and Klaus Altland for valuable discussions about transthyretin.