Evidence of horizontal gene transfer from virus to host, or host to virus, can be found in a variety of genomes1,2,3,4. Examples of viral endogenization are the CRISPR spacer sequences found in bacterial host genomes4. Recently, we have found evidence of host protein sequences embedded in the nonstructural polyproteins of (+)ssRNA Group IV viruses. These sequences within the coding regions of the viral genome can be propagated generationally. The short stretches of homologous host-pathogen protein sequences (SSHHPS) are found in the virus and host5,6. SSHHPS are the conserved cleavage site motif sequences recognized by viral proteases that have homology to specific host proteins. These sequences direct the destruction of specific host proteins.
In our previous publication6, we compiled a list of all of the host proteins that were targeted by viral proteases and found that the list of targets was non-random (Table 1). Two trends were apparent. First, the majority of the viral proteases that cut host proteins belonged to Group IV viruses (24 of 25 cases involved Group IV viral proteases), and one protease belonged to the (+)ssRNA Group VI retroviruses (HIV, human immunodeficiency virus)7. Second, the host protein targets being cut by the viral proteases were generally involved in generating the innate immune responses suggesting that the cleavages were intended to antagonize the host's immune responses. Half of the host proteins targeted by the viral proteases were known components of signaling cascades that generate interferon (IFN) and proinflammatory cytokines (Table 1). Others were involved in host cell transcription8,9,10 or translation11. Interestingly, Shmakov et al.4 have shown that many CRISPR protospacer sequences correspond to genes involved in plasmid conjugation or replication4.
Group IV includes, among others, Flaviviridae, Picornaviridae, Coronaviridae, Calciviridae, and Togaviridae. Several new and emerging pathogens belong to Group IV such as the Zika virus (ZIKV), West Nile (WNV), Chikungunya (CHIKV), severe acute respiratory syndrome virus (SARS) and Middle East respiratory syndrome virus (MERS). The (+)ssRNA genome is essentially a piece of mRNA. To produce the enzymes necessary for genome replication, the (+)ssRNA genome first must be translated. In alphaviruses and other Group IV viruses, the enzymes necessary for replication are produced in a single polyprotein (i.e., nsP1234 for VEEV). The nonstructural polyprotein (nsP) is proteolytically processed (nsP1234 nsP1, nsP2, nsP3, nsP4) by the nsP2 protease to produce active enzymes12 (Figure 1). Cleavage of the polyprotein by the nsP2 protease is essential for viral replication; this has been demonstrated by deletion and site-directed mutagenesis of the active site cysteine of the nsP2 protease13,14. Notably, the translation of viral proteins precedes genome replication events. For example, nsP4 contains the RNA-dependent RNA polymerase needed to replicate the (+)ssRNA genome. Genome replication can produce dsRNA intermediates; these intermediates can trigger the host's innate immune responses. Thus, these viruses may cleave host innate immune response proteins early in infection in order to suppress their effects15,16,17.
Silencing can occur at the level of DNA, RNA, and protein. What is common to each of the silencing mechanisms shown in Figure 1 is that short foreign DNA, RNA, or protein sequences are used to guide the destruction of specific targets to antagonize their function. The silencing mechanisms are analogous to "search and delete" programs that have been written in three different languages. The short cleavage site sequence is analogous to a "keyword". Each program has an enzyme that recognizes the match between the short sequence (the "keyword") and a word in the "file" that is to be deleted. Once a match is found, the enzyme cuts ("deletes") the larger target sequence. The three mechanisms shown in Figure 1 are used to defend the host from viruses, or to defend a virus from a host's immune system.
Viral proteases recognize short cleavage site motif sequences between ~2-11 amino acids; in nucleotides, this would correspond to 6-33 bases. For comparison, CRISPR spacer sequences are ~26-72 nucleotides and RNAi are ~20-22 nucleotides18,19. While these sequences are relatively short, they can be recognized specifically. Given the higher diversity of amino acids, the probability of a random cleavage event is relatively low for a viral protease recognizing protein sequences of 6-8 amino acids or longer. The prediction of SSHHPS in host proteins will largely depend upon the specificity of the viral protease being examined. If the protease has strict sequence specificity requirements the chance of finding a cleavage site sequence is 1/206 = 1 in 64 million or 1/208 = 1 in 25.6 billion; however, most proteases have variable subsite tolerances (e.g., R or K may be tolerated at the S1 site). Consequently, there is no requirement for sequence identity between the sequences found in the host versus the virus. For viral proteases that have looser sequence requirements (such as those belonging to Picornaviridae) the probability of finding a cleavage site in a host protein may be higher. Many of the entries in Table 1 are from the Picornaviridae family.
Schechter & Berger notation20 is commonly used to describe the residues in a protease substrate and the subsites to which they bind, we utilize this notation throughout. The residues in the substrate that are N-terminal of the scissile bond are denoted as P3-P2-P1 while those that are C-terminal are denoted as P1'-P2'-P3'. The corresponding subsites in the protease that bind these amino acid residues are S3-S2-S1 and S1'-S2'-S3', respectively.
To determine which host proteins are being targeted, we can identify SSHHPS in the viral polyprotein cleavage sites and search for the host proteins that contain them. Herein, we outline procedures for identifying SSHHPS using known viral protease cleavage site sequences. The bioinformatic methods, protease assays, and in silico methods described are intended to be used in conjunction with cell-based assays.
Sequence alignments of the host proteins targeted by viral proteases have revealed species-specific differences within these short cleavage site sequences. For example, the Venezuelan equine encephalitis virus (VEEV) nsP2 protease was found to cut human TRIM14, a tripartite motif (TRIM) protein6. Some TRIM proteins are viral restriction factors (e.g., TRIM5α21), most are thought to be ubiquitin E3 ligases. TRIM14 lacks a RING (really interesting new gene) domain and is not thought to be an E3 ligase22. TRIM14 has been proposed to be an adaptor in the mitochondrial antiviral signalosome (MAVS)22, but may have other antiviral functions23. Alignment of TRIM14 sequences from various species shows that equine lack the cleavage site and harbor a truncated version of TRIM14 that is missing the C-terminal PRY/SPRY domain. This domain contains a polyubiquitination site (Figure 2). In equine, these viruses are highly lethal (~20-80% mortality) whereas in humans only ~1% die from VEEV infections24. Cleavage of the PRY/SPRY domain may transiently short circuit the MAVS signaling cascade. This cascade can be triggered by dsRNA and leads to the production of interferon and pro-inflammatory cytokines. Thus, the presence of the SSHHPS may be useful for predicting which species have defense systems against specific Group IV viruses.
In Group IV viruses, IFN antagonism mechanisms are thought to be multiply redundant25. Host protein cleavage may be transient during infection and concentrations may recover over time. We found in cells that TRIM14 cleavage products could be detected very early after transfection (6 h) with a plasmid encoding the protease (cytomegalovirus promoter). However, at longer periods, the cleavage products were not detected. In virus-infected cells, the kinetics were different and cleavage products could be detected between 6-48 h6. Others have reported the appearance of host protein cleavage products as early as 3-6 h post infection9,11.
Proteolytic activity in cells is often difficult to catch; the cleavage products can vary in their solubility, concentration, stability, and lifetime. In cell-based assays, it cannot be assumed that cleavage products will accumulate in a cell or that the band intensities of cut and uncut protein will show compensatory increases and decreases as the cut protein may be degraded very quickly and may not be detectable in a Western blot at an expected molecular weight (MW) (e.g., the region containing the epitope could be cleaved by other host proteases or could be ubiquitinated). If the substrate of the viral protease is an innate immune response protein, its concentration may vary during infection. For example, some innate immune response proteins are present prior to viral infection and are induced further by interferon26. The concentration of the target protein may therefore fluctuate during infection and comparison of uninfected vs. infected cell lysates may be difficult to interpret. Additionally, all cells may not be uniformly transfected or infected. In vitro protease assays using purified proteins from E. coli on the other hand have fewer variables for which to control and such assays can be done using SDS-PAGE rather than immunoblots. Contaminating proteases can be inhibited in the early steps of the protein purification of the CFP/YFP substrate, and mutated viral proteases can be purified and tested as controls to determine if the cleavage is due to the viral protease or a contaminating bacterial protease.
One limitation of in vitro protease assays is that they lack the complexity of a mammalian cell. For an enzyme to cut its substrate, the two must be co-localized. Group IV viral proteases differ in structure and localization. For example, the ZIKV protease is embedded in the endoplasmic reticulum (ER) membrane and faces the cytosol, whereas the VEEV nsP2 protease is a soluble protein in the cytoplasm and nucleus27. Some of the cleavage site sequences found in the ZIKV SSHHPS analysis were in signal peptides suggesting that cleavage might occur co-translationally for some targets. Thus, the location of the protease and the substrate in the cell also needs to be considered in these analyses.
Cell-based assays can be valuable for establishing a role for the identified host protein(s) in infection. Methods that aim to halt viral protease cleavage of host proteins such as the addition of a protease inhibitor6 or a mutation in the host target16 can be used to examine their effects on viral replication. Overexpression of the targeted protein also may affect viral replication28. Plaque assays or other methods can be used to quantify viral replication.