To fully understand how pathogens infect their host and hijack key biological processes, systematic mapping of intra-pathogenic and pathogen–host protein–protein interactions (PPIs) is crucial. Due to the relatively small size of viral genomes (usually around 10–100 proteins), generation of comprehensive host–virus PPI maps using different experimental platforms, including affinity tag purification-mass spectrometry (AP-MS) and yeast two-hybrid (Y2H) approaches, can be achieved. Global maps such as these provide unbiased insight into the molecular mechanisms of viral entry, replication and assembly. However, to date, only two-hybrid methodology has been used in a systematic fashion to characterize viral–host protein–protein interactions, although a deluge of data exists in databases that manually curate from the literature individual host–pathogen PPIs. We will summarize this work and also describe an AP-MS platform that can be used to characterize viral-human protein complexes and discuss its application for the HIV genome.
Protein–protein interaction networks have been generated using AP-MS and Y2H targeting various organisms, ranging from bacteria to humans. Yeast two-hybrid screenings consist of testing all pair wise combinations of proteins, which generates a collection of binary interactions. High-throughput Y2H maps have been generated for Saccharomyces cerevisiae
These unbiased approaches have also been used to study PPIs between proteins that are derived from a specific virus. For example, an intraviral Hepatitis C (HCV) Y2H interaction map was built using a limited set of predefined coding segments, which revealed the functional interactions between the proteins in the viral life cycle when a cell culture system is absent
These pair-wise interaction studies have also been extended into studying the interaction landscape between viral proteins and host factors. For example, Lotteau and colleagues published a proteome-wide, Y2H-based mapping of interactions among HCV and human proteins. They reported 314 interactions (in addition to 170 literature curated interactions) and discovered that HCV CORE protein was a major perturbator of the insulin, Jak/STAT and TGFß pathways
Collectively, the global network properties of human proteins targeted by pathogens, including bacteria and viruses, were recently studied Literature derived HIV–human protein–protein interaction map. A network displaying HIV-human interactions derived from the National Institute of Allergy and Infectious Diseases Division of AIDS (NIAID) HIV-1 Human Protein Interaction Database. HIV proteins correspond to red nodes whereas yellow nodes represent host factors. In total, 1785 unique HIV-human interactions among 1175 human and 15 HIV proteins are presented. Further work will be required to determine which interactions are direct or functionally relevant.
Since the HIV-human interactions are mostly literature-curated
In the following section, we describe: (1) different strategies that can be employed to affinity tag HIV proteins, (2) purification protocols, and (3) ways in which the resulting data can be analyzed and integrated with other types of information.
Characterization of HIV–human protein–protein interactions during infection would arguably result in a dataset that would be most physiologically relevant. In order to accomplish this, however, one would have to tag the proteins in the context of the viral genome, infect the appropriate cells with these genetically altered viruses and then purify and identify the complexes. So far Integrase, Vif and Vpr have been successfully tagged within the provirus while maintaining infectivity
A more amenable AP-MS approach for comprehensively characterizing HIV–human protein–protein interactions is to individually clone each of the factors into an appropriate tagging construct, separately express these factors in a human cell line, and then purify and characterize the resulting complexes. Although simpler, one could argue that the resulting data may be less relevant since some of the viral proteins need other HIV factors for proper function or localization (e.g. MA, IN, Vpr and RT as part of the preintegration complex) and the tagged proteins may be significantly overexpressed when compared to levels during infection.
It is also worth pointing out that the late genes of HIV-1 are expressed from intron-containing mRNAs and depend on the Rev protein for nuclear export and translation. Removing inhibitory sequences and adapting the codon usage for mammals can achieve efficient expression of these ORFs in the absence of Rev. Such codon-optimized versions have been generated of all HIV-1 late genes by different labs
A variation on the approach described in Section
A further step would be to carry out the purifications in the presence of HIV infection. Even though in this scenario, two copies of the viral proteins essential for virus production and infection would be present (i.e. tagged and untagged), this approach would potentially identify intraviral interactions and those that are dependent on other HIV proteins or the viral RNA.
Some HIV proteins are likely to interact with several host protein complexes in order to perform multiple functions during the viral replication cycle. Information about the composition of different HIV–host complexes can be gained using the “split-tag” approach The double pull-down approach to characterize HIV–human protein complexes. In this strategy, cell lines are expressing two proteins, one viral and one host, each with a different affinity tag. The first purification step enriches for the viral protein, and presumably all the complexes it is associated with whereas the second purification step targets a host protein and enriches for a specific and stoichiometric viral-host protein complex.
In this way, one could systematically create lines dually expressing each viral-host protein pair that was derived from the single HIV purification experiments and subject the extract to this “double pull-down” strategy. This approach would: (1) verify the relevance of individual interactions derived from single purification experiments, (2) place host factors into their respective complexes via the co-enrichment patterns and (3) identify more physiological, stoichiometric HIV-human protein complexes, which would more likely be used for subsequent functional assays or even structural studies. Of course, an additional, complementary strategy would be to carry out single, reciprocal purifications of the tagged human proteins, alone and also in the presence of the appropriate viral protein, which would help functionally verify the host-pathogen protein–protein interactions and expand the network to include more host proteins and complexes.
The HIV-1 open reading frames encoding precursor proteins (GagPol, Gag, Pol, gp160), subunits (MA, CA, NC, p6, PR, RT, IN, gp120, gp41) and accessory proteins (Vif, Vpr, Vpu, Nef, Tat, Rev) are PCR amplified from either a proviral vector or codon-optimized templates and ligated into the vector pcDNA4/TO (Invitrogen) carrying either a 5’ 3xFlag2xStrep (FS) or a 3’ 2xStrep3xFlag (SF) tag
A convenient cell line for large-scale AP-MS experiments is HEK293 since these cells are easy to culture and transfect. HEK293 cells are maintained in DMEM high glucose, 10% FBS and antibiotics at 37 °C in 5% CO2. For transient transfections, 2.5 × 106 cells are seeded per 150 cm2 dish and transfected the following day with 5 μg plasmid using standard calcium phosphate precipitation
Since T cells are the natural target cells for HIV, expression and purification of HIV proteins from a T cell line is desirable. Since large-scale T cell transfections are difficult, stably transfected cell lines should be generated. A regulatable expression system should be used since expression of several HIV proteins affect cell viability. Jurkat TRex cells (Invitrogen), for example, stably express the tetracylin repressor protein so that genes from Tet operon containing plasmids like pcDNA4/TO are only expressed upon tetracyclin induction. Jurkat TRex cells are maintained in RPMI, 10% FCS, 10 μg/ml blasticidin, 1% PenStrep. For generation of stable clones, 1 × 106 cells are transfected with 2 ug linearized pcDNA4/TO plasmid by electroporation according to the manufacturer’s instructions (Amaxa). Stably transfected cells are selected in 300 μg/ml Zeocin for several weeks and single clones are isolated by limited dilution. Expression of HIV proteins is induced by 1 μg/ml Doxycylin for 12–24 h. Expression levels of proteins with a short half-life (e.g. Vif) can be increased by addition of 500 nM protease inhibitor MG132 (Calbiochem) during induction.
In order to obtain two independent datasets for each protein and cell line, both Flag and Strep affinity purifications can be performed ( The purification-mass spectrometry strategy to characterize HIV–human protein complexes. Cloned viral genes are inserted into a construct that fuses a 2XStrep3XFlag dual affinity tag on the C-terminus of each factor. These constructs can be used for transient transfection in HEK293 cells or for generating stably expressing Jurkat cells. After lysis, the extract is subjected to either Anti-Flag or Strep-Tactin IP beads where an aliquot of the beads, as well as the elution, is subjected to trypsin digestion and the material is analyzed using the OrbiTrap mass spectrometer. In both cases, a portion of the eluate is also subjected to SDS–PAGE analysis, where the gel is stained, bands excised, proteins extracted and the protein is digested with trypsin and analyzed using a Q-Star Elite mass spectrometer. The data obtained from both sets of purifications and multiple points during the isolation are then integrated together and subjected to an algorithm to derive quantitative viral-host protein–protein interactions. See text for a more detailed description.
Insoluble material is then pelleted for 20 min at 2800
For purification of individual proteins, 30–50 ul IP beads (anti-Flag M2 Affinity Gel, SIGMA or Strep-Tactin Sepharose, IBA) is added to the precleared lysate and the immunprecipitation is performed in batch on an overhead shaker for at least 1 h at 4 °C. Beads are then washed extensively either in columns (Poly-Prep, BioRad) or in batch (2 ml dolphin tubes, BLD Science) with cold 0.1% NP40, 50 mM Tris–HCl pH 7.4, 150 mM NaCl, 1 mM EDTA. In case of purification of proteins with RNA binding motifs (Gag, NC, Tat, Rev), unspecific association of RNA associated proteins like splicing factors, RNA helicases, etc. can be reduced by incubation of the beads with 1500 U RNase A (Fermentas) for 30 min on ice, followed by washing. The last washing step is performed with detergent-free buffer to avoid interference with MS analysis. 10 μl of the beads are directly analyzed by MS using on-bead trypsin digest, while the rest of the beads is eluted with 30–50 μl of either 100 μg/ml 3xFLAG peptide (Elim Biopharmaceuticals) or
Subcellular compartments may have to be enriched prior to purification to identify functional HIV–human protein–protein interactions. Nuclear proteins like Integrase, Rev, Tat, and Vpr can be extracted from the nuclear fraction using high salt buffer according to the standard protocol
Membrane-associated proteins like Gag, Vpu, Nef and Env can be affinity purified from membrane fractions enriched by flotation in a discontinuous iodixanol gradient. To this end, cells are hypotonically lysed and disrupted by dounce homogenization. The lysate is adjusted to 40% OptiPrep (SIGMA), overlaid with 28% Optiprep and TNE buffer (50 mM Tris, pH 7.4, 150 mM NaCl, 5 mM EDTA) and centrifuged at 165,000
For the double pull-down experiment, 10 × 150 cm2 plates HEK293 cells are co-transfected with vectors coding for the strep-tagged viral protein and one or more Flag-tagged host proteins. The affinity purifications are performed as described above, with the first step being scaled up accordingly. The eluates after both steps are compared by SDS–PAGE and mass spectrometry.
An inherent problem of AP-MS experiments is the high number of unspecific interactions that can be detected. Usually the data obtained from the purification of the affinity tagged protein of interest is compared to a negative control using untagged protein, untransfected cell lysate, tagged GFP or preimmune serum. While this helps to identify a limited set of unspecific binders, proteins often have distinct background interactions depending on their localization and nature, e.g. nuclear proteins have a different set of background interactors than membrane proteins. In general, the more unrelated proteins with similar characteristics are analyzed, the easier it is to identify specific, and therefore physiologically relevant interactions.
Since AP-MS studies reveal little information about whether the association is direct or indirect, interactions should be confirmed using different methodologies, including
Once pull-down samples are acquired and analyzed by mass spectrometry, proteomic information concerning host–pathogen interactions can be ascertained. MS identification of a host protein as a putative interactor allows for further investigation into the host–pathogen interaction, employing methods such as yeast two-hybrid, biochemical assays, or viral infectivity assays, to verify and establish the biological relevance of the viral-host interaction. In this regard, proper quantitative analysis of mass spectrometry data derived from affinity purified viral or host proteins is essential in identifying biologically meaningful PPIs in order to minimize time and resources expended on false-positive MS identified interactors. However, one major caveat of reported PPIs obtained from AP-MS experiments is often little information relating to protein abundance or specificity is directly revealed. Also, the interactor abundance may not even be the best indicator of the interaction reliability, especially since some protein abundance strongly depend on interaction affinity as well as its concentration in the cell or in the final experimental sample. If the purified material from a single affinity purification of one tagged protein bait is analyzed, even when contrasted with data from a non-tagged control, it remains incredibly difficult to ascertain specificity and reproducibility with respect to putative interactors. Furthermore, shotgun sequencing approaches to protein identification suffer from poor sampling of IP proteins independent of precision of sample purification replicates. For example, shotgun sequencing of the same sample may only result in 30–40% overlap with respect to proteins identified and may require 5–10 separate runs to obtain 95% coverage
To identify physiologically relevant PPIs, it helps to have information pertaining to protein abundance in the sample, some metric of bait-interactor specificity, and some metric of interactor reproducibility. As an example concerning specificity, RNA binding proteins (e.g. Tat and Rev) will often be identified with ribosomal protein subunits due to binding to an RNA molecule and not a direct PPI occurring
Each pull-down experiment can be represented as a vector of abundance scores (defined below) for all of the unique interactors found in our approach. When the specific interactor is not pulled down with particular bait, its abundance score is set to zero. The vectors of different affinity purification experiments can then be organized into a two-dimensional matrix, and the vectors of the replicated experiments extending further into the third dimension. In order to quantify abundance of identified proteins within a sample and to allow normalization of this data across different samples such that values between multiple datasets can be accurately compared, one can use the label-free SI
Viruses, like HIV, are incredibly complicated, resourceful organisms that are involved in many diverse functions during infection. However, their genomes are surprisingly small considering the tasks they must carry out, and therefore they rely very heavily on the cellular machinery in the host cells they infect. Based on this, one might expect that one viral protein would be involved in multiple processes and therefore would hijack several host complexes during infection. Using an approach like AP-MS, therefore, would be a powerful way to identify these relationships, especially when it is conducted in an unbiased and systematic way. Overlaying the PPI network with genetic information derived from global RNAi screens
We thank members of the Krogan lab for helpful comments. This work was funded by