This is an Open Access article distributed under the terms of the Creative Commons Attribution License (
Efficient communication between distant sites within a protein is essential for cooperative biological response. Although often associated with large allosteric movements, more subtle changes in protein dynamics can also induce long-range correlations. However, an appropriate formalism that directly relates protein structural dynamics to information exchange between functional sites is still lacking.
Here we introduce a method to analyze protein dynamics within the framework of information theory and show that signal transduction within proteins can be considered as a particular instance of communication over a noisy channel. In particular, we analyze the conformational correlations between protein residues and apply the concept of mutual information to quantify information exchange. Mapping out changes of mutual information on the protein structure then allows visualizing how distal communication is achieved. We illustrate the approach by analyzing information transfer by the SH2 domain of Fyn tyrosine kinase, obtained from Monte Carlo dynamics simulations. Our analysis reveals that the Fyn SH2 domain forms a noisy communication channel that couples residues located in the phosphopeptide and specificity binding sites and a number of residues at the other side of the domain near the linkers that connect the SH2 domain to the SH3 and kinase domains. We find that for this particular domain, communication is affected by a series of contiguous residues that connect distal sites by crossing the core of the SH2 domain.
As a result, our method provides a means to directly map the exchange of biological information on the structure of protein domains, making it clear how binding triggers conformational changes in the protein structure. As such it provides a structural road, next to the existing attempts at sequence level, to predict long-range interactions within protein structures.
Cooperative protein response and thus cooperative network behavior requires information transfer between distal sites in a protein or protein complex. Protein structures often achieve such long range communication by allosteric movement [
Now that the principles relating correlated structural dynamics to signal transduction mechanisms within proteins are becoming apparent, we here present a method to identify and quantify these structural fluctuations in terms of information allowing computation of the information transfer between active sites. In particular, we use information theory [
To illustrate our information theoretic approach, we map the residue-based information transfer network of the SH2 domain of the Fyn tyrosine kinase [
The SH2 domain has a typical length of approximately 100 amino acids and displays a simple fold (shown in Figure
Experimental work on the role of the linker between the SH3 and SH2 domain of Fyn has shown that the nature of the linker, i.e. the residue pattern, determines both orientation and coupling between the two domains [
In Shannon's theory of information transduction along noisy channels, where the level of noise corresponds to the fraction of the input characters that may be wrongly translated to output characters, information is quantified as the degree of certainty obtainable about the output signal, once a particular input is given [
So when do two sidechains exchange information? Two linked residues are dependent when knowing one conformation will convey information about the conformation of the other. In this case, the two residues share information. For instance, consider two amino acid residues in a protein, each having some conformational variation in their side chains. The mutual information shared by these residues is defined by the entropy reduction that is observed at the second residue when the conformation of the first is fixed in an arbitrary configuration (and vice versa). Thus, when ligand binding conformationally restricts the first residue then mutual information will quantify how much of the signal of ligand binding is received at the second residue since its conformational flexibility will also be restricted. Mutual information between two residues is thus zero when the conformational state of one residue does not provide any information about the conformational state of the other residue. On the other hand, mutual information is maximal when each conformational state of a residue's sidechain uniquely defines the conformational state of the other residue's sidechain. In accordance with information theory, the conformational space of each residue corresponds to the residue's alphabet.
The calculation of mutual information between any two amino acids in a protein requires reliable statistics of the conformational dynamics of all amino acids in the native structure of the protein (see Methods for details). The method relies on computing both the probability of residue i to adopt an arbitrary conformation independently and the joint probability of finding residue i and residue j in an arbitray combined combination. This is achieved via effectively sampling the conformational states of each amino acid in a protein structure so that all possible combinations of conformations of residues i and j can be explored. It is clear that the computational size of such a sampling problem can only be addressed by reducing the number of possible conformations that each amino acid can adopt to a relatively small number of discrete states. To achieve this, we treat the sidechain and backbone flexibility of each amino acid as separate components of its overall conformational dynamics and consider only a finite number of states for each.
In order to keep the number of sidechain conformations computationally tractable, it is common to employ a discrete, finite-size alphabet for each residue, called rotamer library. Such rotamer libraries are constructed by extracting from a database of high resolution protein structures the most frequently occurring states and thus the degree of coarse graining that is employed to record the statistics imposes a resolution on the data. For the purpose of detecting conformational dependencies between neighboring residues in a protein structure, we require a finer resolution than is provided by common rotamer libraries used for homology modeling and we thus constructed a database of backbone dependent sidechain conformations with a 10 degree resolution on the sidechain dihedral (Chi-) angles (see Methods). Given the backbone dihedral angles of a certain residue in the protein structure, a list of possible sidechain conformations and their probabilities can be retrieved from this database. The number of chi angles in the residue determines the size of this list, meaning that residues with small sidechains tend to have shorter lists than residues with long sidechains. The conformational state of the sidechain in the protein structure serves as a starting point for generating these lists.
Information concerning the flexibility of the backbone of a protein can be obtained from either molecular dynamics simulations (MD) or NMR data. Exploration of the conformational space accessible to the protein backbone using MD is a computationally expensive method. Moreover, Vendruscolo and co-workers have recently shown that MD simulations yield more realistic results when restricted by experimental data in the form of residue-residue distances derived from Nuclear Overhauser Effects (NOEs) in a Nuclear Magnetic Resonance (NMR) experiment [
The sidechain sampling results for a single backbone structure obviously introduce a strong bias in the apparent information flux towards residues whose sidechains arbitrarily happen to be strongly coupled in this particular backbone conformation. When information is obtained from an ensemble of related backbone conformations however, each backbone introduces slight variations in the pattern of residue-residue couplings. As a result, the combination of the sidechain sampling on the entire ensemble of backbone structures acts as a filter that removes sporadic couplings while accumulating consistent couplings, thereby revealing the true network of information exchange between all residues. We have here taken the entire ensemble of backbone structures in the NMR datasets of the Fyn kinase SH2 domain as an adequate sample and have not systematically explored if a reduced number of backbones could be employed to the same effect.
Since the Monte Carlo sampling directly yields the equilibrium distribution of sidechain conformations observed at each position along the protein backbone, the quantity of mutual information between all residue pairs in the protein can be calculated using probability and entropy calculations (see Methods and [
In order to obtain the network of residue-residue couplings in the Fyn SH2 domain, we employed a FoldX based sidechain sampling on the backbone structures of this SH2 domain as determined by NMR on the protein domain in isolation and bound to its phosphopeptide ligand (pdb identifiers 1AOU and 1AOT [
To capture how the information exchange within the SH2 domain changes due to ligand binding we require NMR structural data on both the bound and unbound state of the protein. Since there is no structural data on the unbound state of Fyn SH2 available publically, we assume here that the backbone flexibility for bound and unbound state can be derived from the ensemble stored in the 1AOU dataset by energy minimization (see Methods). This simplification has as result that large conformational changes in the backbone of the SH2 domain are not taken into account. Yet since domains do normally not experience large structural changes, this simplification may still provide viable results.
Figure
To visualise the mechanistic implications for our model system a clustering algorithm (see Methods) was applied on the mutual information matrix in Figure
Thus, our analysis reveals a strong coupling between the peptide binding site of Fyn SH2 and its SH3-SH2 and SH2-Kinase linkers. These first results show that the sampling approach discussed here provides a direct way to identify and quantify highly coupled groups of residues. Moreover, it confirms the notion that subtle changes in structural dynamics can effectively couple residues at distal locations in the structure. Note that the clustering does not show how this information exchange actually occurs, i.e. it does not provide a causal explanation. It provides an identification of the residues that may be involved with signaling. Yet, the idea is that when, by lowering the threshold, a coherent collection is obtained, possibly identifying a consecutive path that links the binding region to other parts of the domain structure.
To understand the key elements of information exchange in proteins we compared the change in mutual information of both silent and informative residues with the rest of the protein. We will here illustrate our analysis by two examples that recapitulate the key features of both informative as well as silent residues. Ser23 will serve as an instance of the collection of silent residues whereas HisβD4 represents a prominent member of the informative residues. Specifically, we mapped the change in mutual information of these residues with all the other residues of the SH2 domain onto the structure. Strong changes (ΔI > 1.0) are colored red, weak changes (ΔI < 0.3) blue and intermediate changes are white. In Figure
As argued in previous sections, our information theoretical approach provides a method to quantify the change in exchange of mutual information between all pairs of residues of the Fyn SH2 domain as a consequence of ligand binding. Our analysis revealed that for the majority of residues the conformational coupling to the rest of the protein is not altered upon peptide binding and as a consequence they can be considered silent in terms of signal transduction. A small fraction of residues, however, experience a significant change in conformational coupling and are thus information-rich. Clustering only the most informative of these residues (ΔI ≥ 1.1 bits) revealed a strong coupling between the peptide binding site and the region harboring the linkers connecting the Fyn SH2 to the other domains of the Fyn kinase. It remains of course to be explained by which structural mechanism these distal sites become coupled. As argued earlier, the current framework can only identify and quantify the residues that are involved in the interaction. To identify the causal relationships a different analysis is required.
Several mechanisms for signal transduction have been described in the literature [
As our model only displays increase in conformational coupling upon ligand binding, suggesting a pathway model, we here extract the group of residues that define the pathway through clustering (see Methods) of informative residues that have a mutual affinity higher than a particular noise level (0.5 bits). In Figure
As discussed above only a subset of residues of the SH2 domain is involved in signal transduction, whereas the other residues seem to be conformationally isolated from ligand binding. Importantly, signal transduction is achieved throughout the core of the SH2 domain and involves the main secondary structure elements of the structure. Hence, we show that the Fyn SH2 domain acts as indivisible information transmission unit propagating information from its binding site over the tertiary structures to the linker regions of the domain.
Driven by the availability of genomics and proteomics data, biological research is increasingly focused on cellular networks rather than on individual proteins. Accordingly, a major effort in systems biology is devoted to understanding how topology shapes macromolecular networks into functional units performing information processing tasks such as logical gating, band-pass filtering or signal modulation. Biological networks are generally represented by graphs of interconnected protein nodes. In reality, however, these networks form macromolecular complexes whose assembly is determined by intermolecular atomic interactions. As shown by several experimental and theoretical studies, ligand binding may induce changes in the structural dynamics of proteins. These perturbations propagate cooperatively throughout the protein, coupling distal locations and thereby effectively transferring information. Hence, proteins domains are the elementary units of information processing in cellular organisms. As such proteins can themselves be represented as networks of interconnected residues exchanging information. Protein complexes can by extension be modeled in a similar way. The method introduced here allows to quantitatively translate protein structural dynamics in terms of Shannon's information theory. Moreover, mutual information permits to determine how much information is exchanged between each pair of residues within a protein structure. Transposing structural dynamics in terms of Shannon information has the advantage of providing a description level that matches the relevant functional features – i.e.biological information,- that we aim to extract from signal transduction pathways. Moreover, the analysis does not make any assumptions on the information transmission model. It strictly focuses on identifying the residues whose coupling increases or decreases as a consequence of peptide binding. Further, the method is applicable in identical manner to residue networks at atomic scale as well as to macroscopic descriptions of cellular networks.
As explained in Results, the mutual information exchanged between residues in protein residue networks is simply a statistical thermodynamics description of the conformational coupling between these residues. By discretising conformational space at the residue level, we easily find the average Shannon information for each residue by calculating the average entropy of the conformational probability distribution for each residue. The mutual information exchanged between two residues is the difference of the Shannon entropy from one residue and the conditional Shannon entropy of that residue given knowledge of the probability distribution of the other residue. In other words, it represents the amount of coupling between two protein residues. Once the mutual information is calculated for each pair of residues, clustering can be applied to determine the structural mechanism by which information is transferred between functional sites. In effect we hereby consider a protein as an instance of a noisy coding channel linking a probability distribution of input signals at one binding site with a probability distribution of output signals at another binding site. For the Fyn SH2 domain, we find that information between functional sites is conveyed by a contiguous pathway of residues that cross the hydrophobic core to connect both sites. In particular our results suggest a clear information flux between both binding sites and residues near the linkers of the domain towards the other domains of the Fyn kinase. As a result it is likely that removal of the phosphopeptide from the SH2 binding site uncouples the SH2 domain from both the SH3 domain as well as from the kinase domain, thereby facilitating the activation of the Fyn kinase as already suggested in the literature [
Using experimental methods such as NMR in combination with protein engineering approaches together with our method for information quantification should allow to test and explicit signal transduction mechanisms for a variety of model systems, which would be highly valuable especially for modular domains such as SH2, SH3 or PDZ domains. Especially for the latter domain a number of experimental studies have already shown the presence of long-range communication [
In this work an NMR ensemble and average minimized structure of the Fyn SH2 domain was used [
The energetic properties of all structures in the NMR ensemble were determined using the FoldX forcefield [
Since no data is publically available on the unbound domain structure of Fyn SH2, we derived an ensemble from the existing 1AOU data by first removing the peptide and then energy minimizing the structures using the Yamber2 force field. Figure
Average RMSD difference between bound and unbound structures of the Fyn SH2.
Mutual information expresses the amount of information that the output conveys about the input (and vice versa). It is formally expressed in terms of entropy:
All entropy values can be easily derived from the probabilities related to the conformational states of the residues. For instance, if X corresponds to a particular residue, then P(X = x) correspond to the probability if finding the residue's sidechain in conformation x and P(X = x, Y = y) corresponds to the probability of finding residues X and Y in conformation x and y, respectively. This probabilistic information is gathered by the Monte Carlo sampling discussed later.
Since the number of conformational states depends on the size of the amino acid sidechain, the actual mutual information values can differ. This kind of bias towards residues with large absolute values, which are predominantly the larger, partially solvent exposed sidechains, needs to be removed. To normalize the results, a measure called redundancy is used. Redundancy is calculated as
We derived a statistical rotamer database based on conditional statistics of dihedral angles. All statistics were derived from the WHAT IF dataset [
A set of n random rotamers can be asked from the database and will be derived from the probability distribution previously calculated. The first dihedral angle, Chi1, will be derived first depending on the backbone conformation of the actual residue to mutate (Phi, Psi angles). If the position in the structure is on a poorly populated area of the Ramachandran plot (according to the WHAT IF dataset), we use statistics derived from neighbor bins (+/- 20° on each dihedral angles, keeping the most populated bin) and a small fraction (10%) of the sidechains will also be generated using independent probabilities (P(Chi1)). Chi2 angles are generated according to the single conditional probability distribution P(Chi2|Chi1) based on the previously determined dihedral angle. For longer sidechains, i.e. for Chi3 and Chi4 angles, the random dihedrals are generated according to the probability distributions P(Chii | Chii-1, Chii-2).
In the context of the work presented here, the aim is not to create a rotamer library per se, but more importantly to represent the conformational space of a residue X as a discrete alphabet or state space. The approach used here allows enumeration of states with greater resolution than classical rotamer libraries. A similar set of higher resolution sidechain conformers were used by Honig and co-workers with great success for sidechain reconstruction [
Before starting the sampling process, the alphabet of possible sidechain conformations for each residue is determined using the method discussed in the previous section. The process iterates over all backbones in the NMR data, checks whether the conformation in the data is present and adds it if it does not yet exist, meaning it has a 10 degrees difference with the conformations already in the alphabet. Once this is determined the same alphabet is used for the sampling of every backbone in the NMR ensemble.
One run off the sampling process, which uses the FoldX force field to determine the energy of each state and a Monte Carlo algorithm to perform the conformational sampling, is given one backbone from the NMR data, commences with an energy minimization phase, to ensure that the sampling starts from a state in the energy well and then start collecting data of the residue states. In order to collect many data points this process is run many times in parallel for every backbone separately.
Details and concept of the FoldX force field are discussed at length by Guerois et al. [
Note that FoldX is used to score the complete conformation including the packing interactions.
In the core of the protein changes in conformational dynamics are smaller in absolute value since the alphabet for smaller hydrophobic residues is limited and the conformational restrictions from the neighbouring residues is much more pronounced. Nevertheless, small absolute changes in such positions may still represent large relative changes. Thus, a relative scoring of the change in information exchange will separate meaningful signals from meaningless signals in a better way. Using a logarithm of the ratio of bound mutual information versus unbound mutual information is hence a reasonable way to represent this relative change in mutual information.
The matrix of change in mutual information between all residue pairs is clustered using the CAST algorithm [
TL contributed in the conception and development of the principles of this work, developed the computational framework and carried out the computational analysis. JFB, FS, LS, JS and FR conceived the principles of the work and assisted in the development of the computational framework. All authors drafted, edited and approved the manuscript.
Discussions with and technical assistance of Sebastian Maurer-Stroh, Joke Reumers and Joost Van Durme are gratefully acknowledged.