This is an Open Access article distributed under the terms of the Creative Commons Attribution License (
Protein tertiary structure prediction is a fundamental problem in computational biology and identifying the most native-like model from a set of predicted models is a key sub-problem. Consensus methods work well when the redundant models in the set are the most native-like, but fail when the most native-like model is unique. In contrast, structure-based methods score models independently and can be applied to model sets of any size and redundancy level. Additionally, structure-based methods have a variety of important applications including analogous fold recognition, refinement of sequence-structure alignments, and de novo prediction. The purpose of this work was to develop a structure-based model selection method based on predicted structural features that could be applied successfully to any set of models.
Here we introduce SELECTpro, a novel structure-based model selection method derived from an energy function comprising physical, statistical, and predicted structural terms. Novel and unique energy terms include predicted secondary structure, predicted solvent accessibility, predicted contact map, β-strand pairing, and side-chain hydrogen bonding.
SELECTpro participated in the new model quality assessment (QA) category in CASP7, submitting predictions for all 95 targets and achieved top results. The average difference in GDT-TS between models ranked first by SELECTpro and the most native-like model was 5.07. This GDT-TS difference was less than 1% of the GDT-TS of the most native-like model for 18 targets, and less than 10% for 66 targets. SELECTpro also ranked the single most native-like first for 15 targets, in the top five for 39 targets, and in the top ten for 53 targets, more often than any other method. Because the ranking metric is skewed by model redundancy and ignores poor models with a better ranking than the most native-like model, the BLUNDER metric is introduced to overcome these limitations. SELECTpro is also evaluated on a recent benchmark set of 16 small proteins with large decoy sets of 12500 to 20000 models for each protein, where it outperforms the benchmarked method (I-TASSER).
SELECTpro is an effective model selection method that scores models independently and is appropriate for use on any model set. SELECTpro is available for download as a stand alone application at:
Selecting the most native-like model from a set of possible models is a crucial task in protein structure prediction. A variety of Model Quality Assessment Programs (MQAPs) have been developed that assign numeric scores to models in a set, and then use the scores to rank the models and ultimately select a single model. MQAP methods can be divided roughly into three categories based on the type of information they use: evolutionary methods use sequence or profile similarity between target sequence and template, consensus methods use similarity between models, and structure-based methods use model coordinates [
Evolutionary methods can provide quality scores that have been shown to correlate with structural similarity to native [
Consensus methods take advantage of the observation that similar models produced by different predictors tend to be more accurate than those that are structural outliers. In practice, consensus methods outperform the methods they draw from, and they rarely pick a very poor model. The disadvantage, however, is that when the best model is a structural outlier it will be overlooked for lack of popularity [
While consensus methods depend on similarity between models, structure-based methods calculate scores on each model independently. For this reason, structure-based methods can be applied to model sets of any size and diversity, and will produce the same score for a model regardless of the other models in the set. Structure-based methods can also be used for template-free modeling [
Here we describe SELECTpro, a novel structure-based MQAP that combines high and low resolution energy terms into a model selection method that is effective on model sets of variable size, diversity, and target difficulty. Most of our assessment is calculated from the CASP7 model quality assessment category (QA) results published online [
We analyze the CASP7 quality assessment category predictions with a focus on the quality of the model ranked first by each predictor and the recovery of the most native-like model in the set. Only
Quality of Model Ranked First (MQA1) Relative to Most Native-Like Model (Mmax)
|
|
|
|||||||||
|
|
|
Δ |
Δ |
Δ |
|
Δ |
Δ |
Δ |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 713_1 | 95 (124) | 7 | 11 | 63 | 5.44 |
|
|
|
|
2.5E-01 |
| 634_1 | 95 (124) | 7 | 15 | 53 | 7.75 |
|
|
|
|
|
| 704_1 | 95 (124) | 5 | 8 | 49 | 7.76 |
|
|
|
|
|
| 178_1 | 95 (124) | 8 | 12 | 59 | 8.44 |
|
|
|
|
|
| 633_1 | 95 (124) | 6 | 9 | 52 | 10.12 |
|
|
|
|
|
| 692_1 | 95 (124) | 6 | 9 | 52 | 10.16 |
|
|
|
|
|
| 657_1 | 95 (124) | 1 | 5 | 40 | 12.71 |
|
|
|
|
|
| 691_1 | 95 (124) | 0 | 1 | 24 | 15.10 |
|
|
|
|
|
| 091_1 | 94 (123) | 11 | 18 | 61 | 7.93 |
|
|
|
|
|
| 026_1 | 94 (123) | 1 | 2 | 40 | 9.30 |
|
|
|
|
|
| 338_5 | 93 (122) | 2 | 3 | 37 | 15.10 |
|
|
|
|
|
| 556_1 | 93 (121) | 10 | 15 | 51 | 6.83 |
|
|
|
|
|
| 734_1 | 92 (120) | 4 | 4 | 36 | 16.16 |
|
|
|
|
|
| 718_1 | 92 (119) | 1 | 3 | 32 | 14.04 |
|
|
|
|
|
| 717_1 | 87 (112) | 3 | 7 | 36 | 10.15 |
|
|
|
|
|
| 016_1 | 86 (111) | 5 | 9 | 49 | 7.93 |
|
|
|
|
|
| 038_1 | 85 (108) | 3 | 7 |
|
5.75 |
|
|
|
|
1.2E-01 |
| 276_1 | 80 (104) | 5 | 5 | 39 | 8.94 |
|
|
|
|
|
| 013_1 | 78 (100) | 4 | 6 | 41 | 9.86 |
|
|
|
|
|
| 703_1 | 69 (86) | 3 | 6 | 35 | 8.74 |
|
|
|
|
|
| 191_1 | 61 (78) | 2 | 5 | 32 | 9.35 |
|
|
|
|
|
| 066_1 | 55 (72) | 1 | 2 | 14 | 23.19 |
|
|
|
|
|
a The number of targets where the QA group made a valid prediction (
* SELECTpro (699_1) results appear in bold face and all results that are better than SELECTpro are underlined. Statistically significant p-values (p < .05) are also in bold.
The assessment of the recovery of the most native-like model, is performed on both
Recovery of Top GDT-TS Model (Mmax)
|
|
|
||||||||||||
|
|
|
|
|
||||||||||
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|||||||||||||
|
|
|
|
|
|
|
- |
|
|
|
|
- | - | - |
| 704_1 | 95 (124) | 46.5 | 17.8 |
|
|
|
633_1 | 95 (124) | 20.7 | 11.8 |
|
|
|
| 178_1 | 95 (124) | 42.3 | 19.6 |
|
|
|
634_1 | 95 (124) | 29.5 | 12.7 |
|
|
5.7E-02 |
| 657_1 | 95 (124) | 78.5 | 37.0 |
|
|
|
704_1 | 95 (124) | 24.1 | 13.1 |
|
|
|
| 634_1 | 94 (121) | 52.0 | 16.5 |
|
|
|
178_1 | 95 (124) | 24.1 | 13.7 |
|
|
|
| 091_1 | 94 (123) |
|
17.4 |
|
|
|
657_1 | 95 (124) | 53.5 | 32.0 |
|
|
|
| 633_1 | 94 (121) | 39.0 | 20.6 |
|
|
|
713_1 | 94 (122) | 18.3 | 10.9 |
|
|
2.0E-01 |
| 026_1 | 94 (123) | 55.9 | 22.7 |
|
|
|
692_1 | 94 (122) | 20.6 | 11.6 |
|
|
6.7E-02 |
| 556_1 | 93 (121) | 33.8 |
|
|
|
* | 091_1 | 94 (123) |
|
12.3 |
|
|
|
| 692_1 | 93 (119) | 38.7 | 20.6 |
|
|
|
026_1 | 94 (123) | 37.3 | 18.3 |
|
|
|
| 691_1 | 93 (120) | 98.1 | 28.6 |
|
|
|
691_1 | 94 (123) | 54.4 | 22.2 |
|
|
|
| 338_2 | 93 (122) | 60.4 | 30.2 |
|
|
|
556_1 | 93 (121) | 21.2 | 10.3 |
|
|
4.9E-01 |
| 713_1 | 92 (116) |
|
12.8 |
|
|
3.2E-01 | 338_2 | 93 (122) | 28.2 | 16.8 |
|
|
|
| 734_1 | 89 (116) | 55.2 | 31.5 |
|
|
|
734_1 | 88 (115) | 28.9 | 18.1 |
|
|
|
| 718_1 | 83 (105) | 81.6 | 31.9 |
|
|
|
718_1 | 83 (105) | 46.4 | 26.9 |
|
|
|
| 717_1 | 78 (98) | 46.8 | 22.8 |
|
|
|
717_1 | 78 (98) | 28.4 | 16.4 |
|
|
|
| 013_1 | 78 (100) | 60.1 | 27.5 |
|
|
|
013_1 | 78 (100) | 32.4 | 17.6 |
|
|
|
| 276_1 | 78 (102) | 52.9 | 28.9 |
|
|
|
276_1 | 78 (102) | 29.0 | 18.7 |
|
|
|
| 038_1 | 70 (87) |
|
11.9 |
|
|
4.6E-01 | 038_1 | 74 (95) | 19.8 | 10.7 |
|
|
3.6E-01 |
| 703_1 | 69 (86) | 37.2 | 20.6 |
|
|
|
703_1 | 69 (86) | 20.6 | 14.5 |
|
|
|
| 191_1 | 61 (78) | 45.5 | 21.9 |
|
|
|
191_1 | 61 (78) | 27.6 | 15.2 |
|
|
|
| 066_1 | 55 (72) | 91.1 | 54.6 |
|
|
|
066_1 | 55 (72) | 48.0 | 46.5 |
|
|
|
| 016_1 | 53 (72) |
|
20.0 |
|
|
|
016_1 | 53 (70) | 18.2 | 18.5 |
|
|
|
a The number of targets where the QA predictor scored Mmax (
b In the CASP7 submission SELECTpro did not have a score for Mmax of target T0356 due to a processing error. We added in the score for this analysis in order to make complete common subset comparisons.
* SELECTpro (699_1) results appear in bold face and all results that are better than SELECTpro are underlined. Statistically significant p-values (p < .05) are also in bold.
Correlation of Selected Groups
|
|
|
|
|
Δ |
| 634_1 (Pcons) a | 95 | 0.811 | 0.847 | 0.036 |
| 713_1 (Circle-QA) b | 95 | 0.765 | 0.823 | 0.058 |
| 633_1 (ProQ) b | 95 | 0.716 | 0.781 | 0.064 |
| 699_1 (SELECTpro) b | 95 | 0.676 | 0.763 | 0.087 |
| 556_1 (LEE) c | 93 | 0.814 | 0.792 | -0.023 |
a Consensus method.
b Structure based method.
c QA scores calculated as GDT-TS similarity to human predictor of LEE.
The utility of SELECTpro for selecting the best model from a small set is demonstrated by selecting from the five models submitted for each target by the top automated predictors. These small set selection results are calculated using
To make fair comparisons to groups participating on only a subset of targets, common subset comparisons between SELECTpro and each of these groups are included in Tables
For multiple domain targets, the sum of GDT-TS over all domains is used as the GDT-TS of the model. Since the QA predictions correspond to the entire structures, it is impossible to fairly assess the domains independently.
To assess the significance of the summary statistics compared in Table
The following notations are used throughout the results section:
• Mmax: The model with the highest GDT-TS among all server models.
• M
•
•
The recovery of Mmax by a QA predictor can only be evaluated if Mmaxwas scored by the predictor. In most cases QA predictors did not provide scores for all available server models, and frequently there is no score for Mmax. For example, predictor 016_1 (AMBER/PB) made submissions on 86 targets, but Mmax is only scored for 53 of these targets – so only these targets (
In this section on the assessment of the model ranked first, and the corresponding Table
• Δ
•
• Δ
The columns of Table
Another way to assess the quality of M
How well does a QA predictor recover Mmax? The traditional metric to assess Mmax recovery is the
• M
• Δ
•
• Δ
Figure
The assessor evaluation of the quality assessment category [
Incomplete models present a challenge to SELECTpro and other structure-based methods because the scores for each model are only comparable when calculated on coordinates for the same set of residues. Another issue is that some complete models have severe chain-breaks, severe steric clashes, or significant portions modeled only as extended chains. These local problems can overwhelm the energy of what may otherwise be a good model. Consensus methods do not suffer from these local structure problems. Given this rationale, one would expect structure-based methods to see the most improvement in terms of average Pearson Correlation on
Predictors in CASP may submit up to five models, but CASP evaluation focuses on the model designated as Model 1. Clearly, the selection of Model 1 is critical in the CASP setting and for protein structure prediction in general. Figure
Here we analyze SELECTpro's model selection capability on the large decoy sets for 16 small proteins from a recent I-TASSER benchmark set [
On the benchmark set SELECTpro has an average GDT-TS of 63.7, while I-TASSER has an average GDT-TS of 62.1. SELECTpro's average Δ
A MQAP that can select the most native-like model from a set of possibilities has a variety of applications in protein structure prediction. The new quality assessment category introduced in CASP7 allows for the unbiased assessment of MQAPs on the models produced by automated predictors. This category allows researchers to focus on the model scoring aspect of protein structure prediction.
The results presented in this work demonstrate that SELECTpro, a structure-based model selection method, consistently selects one of the best models from the large diverse sets of models produced by automated predictors, across all levels of target difficulty. On these large diverse sets of models, SELECTpro also recovers the single most native-like model well compared to other methods. On the small sets of five models submitted for each target by the top automated predictors, in most cases SELECTpro selects better models than the predictors themselves.
Since SELECTpro and other structure-based methods score models independently, they can be incorporated into the model selection pipelines of individual protein structure prediction servers. For this reason, it may help predictors if the CASP organizers distinguished methods that score models independently from those that do not.
Consensus and structure-based methods can be combined to achieve improved results. For example, the meta-server method Pmodeller [
SELECTpro has been made publicly available as a server, where users may submit from 2 to 100 models for evaluation. In addition to the global confidence scores, the scores of individual energy terms are also returned to the user by email for each model submitted. SELECTpro is one of several protein structure tools in the SCRATCH suite of predictors [
All of the comparative analysis in this work is performed on the server models and quality assessment predictions submitted in the CASP7 [
The
The scores produced by SELECTpro are comparable on complete models of the same sequence. There is no standard for the handling of incomplete models and we assume that participating groups took a variety of approaches. Using only complete models ensures that the MQAP scores are calculated from the same coordinates. Thus, the models retained in
Structure-based MQAPs are susceptible to local structural irregularities in models, and will tend to score such models poorly. This is why methods developed to select near-native models from sets of decoys remove such models from consideration [
The Cα-Cα clash model filter enforces a squared difference penalty for Cα-Cα distances less than 3.6 Ǻ. The distance between the Cα atoms of residue
The Cα-Cα chain break model filter enforces a squared difference penalty for
The expanded termini filter removes models where a large portion of the structure is modeled as expanded chain with no non-local interactions. The screening procedure is: scan from the N-terminus until three consecutive residues have a contact number of at least 10, and repeat from the C-terminus. The contact number of a residue is defined here as the number of other Cβ atoms within 10 Ǻ of the residue's Cβ [
In the reduced representation the heavy backbone atoms, carbonyl oxygen, amide hydrogen (N, Cα, C, O, H), and Cβ are represented explicitly. For glycine residues a pseudo Cβ is calculated. The side-chain atoms are represented by a single united point (centroid) [
In the all heavy-atom representation the centroid is removed and the heavy side chain atoms are represented explicitly. The side-chains are initially placed onto the backbone of the reduced representation in their most likely conformation according to the SCWRL backbone-dependent rotamer library [
The parameter weights were determined by repeatedly varying individual weights and maximizing the sum of the GDT-TS of the lowest
Throughout this work the convention of all capital letters referring to global energy and all lower case referring to local energy is used. For instance,
v
u
Ω
Ω
Ω
The details of how the novel reduced representation energy terms are calculated are presented in this section. The predicted structural terms
The predicted structural feature predictions used in
The predicted secondary structure term
The definition of
The solvent accessibility predictor ACCpro predicts the percent of solvent accessibility in 5% increments for each residue. Using 25% exposure as a binary threshold the accuracy of the predictor is ~77% [
The contact map predictor CMAPpro predicts the probability of contact or non-contact between Cα atoms, with a contact threshold of 12 Å. The strategy utilized to infer predicted contacts from the probability matrix [
The predicted contact map can help identify the highest GDT-TS models in the set, even when they are not highly similar to native. A good example of this is CASP7 target T0304 is a 122 residue α/β protein where the highest GDT-TS model in the set is Zhang-Server_TS1 (GDT-TS = 45.55). Most secondary structure predictors (including SSpro) failed to predict the first two strands making this target especially difficult. No QA method ranked the highest GDT-TS model first; however, SELECTpro ranked it second and the model ranked first by SELECTpro (T0304.Zhang-Server_TS4) has the second highest GDT-TS. These models have the lowest
The formation of hydrogen bonds between the residues of
In the equations for
Between two anti-parallel strand partners, only every other pair of residues is hydrogen bonded. For the pairs that are not hydrogen bonded, a pseudo-bonding calculation is used. The hydrogen bonding energy and pseudo-bonding energy are both calculated and the minimum of the two is used in
If residues
Φ(
The penalty for the observed value (
The all-atom energy terms depend on atom-atom interactions when all heavy atoms are included in the model. In the all-atoms energy equations
In the interest of completeness and reproducibility we include the details of the energy terms that are adapted from previous work.
This term penalizes steric clashes between non-bonded atoms explicitly represented in the reduced representation. The penalty for overlapping atoms is the overlap distance squared as defined here:
A centroid-centroid repulsive term is used to reduce the overcrowding of side-chains in the reduced representation. The minimum distance between two centroids in the calculation is the minimum observed for each pair of residue types –
The motivation for this term is to model the hydrophobic effect. The level of burial for each residue in the model is estimated by the number of other Cβ atoms within 10 Ǻ (the contact number
This context independent pair-wise potential comes from Equation 6 of [
This context specific pair-wise potential is from [
The radius of gyration is a simple measure of the global compactness of a domain.
A fundamental characteristic of native globular protein structures is their efficient steric packing of atoms in the protein core. A Lennard-Jones 12-6 potential with damped repulsion (
Solvation energy is calculated using the implicit solvation model described in [
Electrostatic interactions between charged atoms are treated by simple repulsion and attraction according to inverse distance squared. The use of distance squared rather than linear distance encourages the formation of salt bridges in the models. There is a correction for atom-atom distance below the minimum realistic value. The ideal distance between oppositely charged atoms is
•
•
•
•
•
AR and PB designed the novel energy terms. AR implemented the methods and carried out the experiments. AR and PB authored the manuscript. Both authors approved the manuscript.
Work supported by NIH grant LM-07443-01, NSF grants EIA-0321390 and IIS-0513376, and a Microsoft Faculty Research Award to PFB.