A subscription to JoVE is required to view this content. Sign in or start your free trial.

Method Article

Application of I TASSER, trRosetta, UCSF Chimera, HADDOCK server, and HEX loria for De Novo and In Silico Design of Proteins

1.6K views

DOI:

10.3791/68003

July 8th, 2025

In This Article

Summary

Protein design involves the construction of amino acid sequences and the incorporation of specific motifs to create functional variants. This approach is critical for the development of antimicrobial peptides (AMPs) to combat antibiotic-resistant pathogens. This paper presents a procedure for protein construction using various bioinformatics tools.

Abstract

The frontiers of protein design lie in the construction of amino acid sequences and the incorporation of specific motifs, such as binding sites, into receptors, ion channels, or other proteins. This approach allows the production of protein variants with specific functions and purposes. An exemplary application of these bioinformatics resources is the improvement of antimicrobial peptides (AMPs) or the design of synthetic AMPs. This development is significant because of the recent increase in emerging diseases, often caused by pathogens that are difficult to eradicate and exhibit antibiotic resistance. The consequences of this trend include the global extinction of animal species and human deaths. The rapid synthesis of novel and effective AMPs is, therefore, essential. These servers and programs can facilitate the rapid development of vaccines, AMPs, or effective treatments against these pathogens. The aim of this work was to present and propose a procedure for protein construction from sequences and to elucidate the process of using different tools to build these molecules step by step.

Introduction

The incidence of animal diseases is increasing, and this phenomenon is influenced by climate change. The rise in global temperature has led to the proliferation of microorganisms and the emergence of opportunistic pathogens, contributing to the development of virulent and drug-resistant diseases1. Climate change affects the soil, water, skin, and gut microbiome, resulting in persistent stress on animal physiology and environmental adaptation. This phenomenon has been demonstrated to compromise immune responses, rendering animals more susceptible to emerging pathogens, which can have lethal consequences and potentially lead to the extinction of certain species2,3,4,5. A salient example is white-nose syndrome (WNS), a disease caused by the fungus Pseudogymnoascus destructans, affecting hibernating bats. This disease and others were not identified until a substantial mortality event occurred in the American continent in 20066. To mitigate the damage caused by such pathogens, researchers have turned to the administration of antimycotics, such as itraconazole and ketoconazole, and probiotics, including inoculation with the beneficial microbiota Pseudomonas fluorescens7. Another example of a disease caused by a fungus is that of Ophidiomyces ophiodiicola, a fungus that causes dermatitis and multisystemic diseases in both captive and wild snakes4,8.

These studies underscore the necessity of heightened awareness and understanding of fungal diseases in wild animals to implement practical solutions expeditiously. Fungal infections have been reported in amphibians, and one prevalent pathogen that affects these animals is the aquatic fungus Batrachochytrium dendrobatidis (Bd), which causes chytridiomycosis. This disease causes severe skin lesions and disrupts normal skin function in susceptible animals, potentially leading to mortality9. This pathogen has been disseminated extensively since the 20th century, contributing to the extinction of 90 amphibian species. The prevalence of these infections poses a substantial challenge in preventing them and identifying effective treatments, as they necessitate consideration of various ecological, physiological, microbiome, and other factors10.

It is a well-documented phenomenon that pathogenic microorganisms possess the capacity to expeditiously evolve resistance mechanisms through genetic coding, a process that can be exemplified by the production of external efflux pumps for drugs and antibiotics11. Furthermore, these microorganisms have been observed to modify the membrane lipid composition with the objective of counteracting the effects of specific molecules12. The field of protein sequences has garnered significant interest, particularly in the context of its integration into treatment approaches and the construction of protein structure models. These models are designed to optimize laboratory time and resources prior to the initiation of protein synthesis13. The prevention of infections in these cases is a complex task involving multiple ecological, physiological, and microbiome factors14. This approach facilitates more effective management of time and resources in the laboratory and emphasizes the complexity of the task.

In the interest of enhancing the potential for practical applications of existing compounds, as well as creating new compounds with more significant potential for such applications, it is imperative to employ expertise from various disciplines, including chemistry, molecular biology, physics, and programming. To this end, several tools are available, including I-TASSER (Iterative Threading ASSembly Refinement)15,16 and trRosetta17,18 (Protein structure prediction by TRansform-restrained Rosetta). structural protein prediction and UCSF Chimera (a visualization system for exploratory research and analysis from the University of California, San Francisco)19 and HEX loria (an interactive protein docking and molecular superposition program)20 for molecular docking and PDB file editing.

I-TASSER is an automated online server that predicts structure and function. This program is accessible via a web browser or as a download. On the submission page, users can input the sequence in FASTA format or download an archive in .txt format. The server process involves identifying structural templates with a degree of homology, as described in Local Meta-Threading Server LOMETS20,21,22,23,24,25. Following the processing of a million potential results by the server, a filtration process is initiated, and fragments are selected based on the Root Median Square Deviation (RMSD), TM (a metric used to evaluate structural similarity between two models), and C-score (a metric used to evaluate the confidence of protein structural models generated by tools). Users can customize the results based on predetermined parameters provided in the program or input their own values depending on their specific sequence or structural expectations20,21,22,23,24,25.

trRosetta is a highly effective program for protein structure prediction. On this server, users can submit sequences in an archive format, TXT, FASTA, input sequence, or Multiple Alignment Sequences (MSA). This program allows users to select between templates or execute trRosetta-single for sequence folding. These servers utilize a deep-learning network to generate structural predictions. The sequence prediction process involves an initial step of contact and distance map prediction, followed by the selection of the top five templates for the final model and a step to minimize energy for folding and refinement17.

The High Ambiguity Protein-Protein Docking (HADDOCK) server is a computational tool that facilitates the flexible docking of two or more PDB files, encompassing protein, peptide, ligand, and receptor structures. This server employs an approach that incorporates AIRs (Ambiguous Interaction Restraints), which encode information from predicted or identified proteins during the docking process. AIRs encompass amino acid restrictions, categorizing them as passive or active amino acids. The incorporation of active amino acids is of particular significance, as they are indicative of potential binding or interaction sites between protein receptors. The program's design incorporates solvent exposure into the model26.

HEX Loria is an online interactive server that facilitates the analysis of protein complexes and molecular docking. This server can read coordinate archives in the Protein Data Bank (PDB) format and rapidly generate docking calculations using NVIDIA or CUDA graphic targets. The program can produce up to 1,000 docking predictions and utilizes the Correlation Fourier for the overall orientations of two molecules27. The HEX Loria interface is designed to be user-friendly, enabling researchers to seamlessly upload their protein structures and configure parameters for docking simulations. The server's capacity to utilize NVIDIA or CUDA graphics processing units has been shown to result in substantial acceleration of computation time, thereby allowing researchers to obtain results with greater expediency. Furthermore, HEX Loria offers a range of visualization options for docking results, empowering users to undertake detailed analysis of predicted protein-protein interactions.

The utilization of the UCSF Chimera is imperative to ensure the accuracy of research findings. This software necessitates installation for the execution of the present protocol, and the ensuing functionalities will be employed: measurement of interatomic distances and various other analytical tasks. The software boasts features such as labeling, coloration, surface rendering, and exportation of figures in JPG or PNG formats. The incorporation of labels and colors facilitates the identification of amino acids with a high probability of contact between the receptor and the designed antimicrobial protein. Additionally, it enables the precise measurement of the distance between these amino acids.

This protocol aims to identify alternative treatment options for the emerging health concern afflicting amphibians. Chytridiomycosis, a fungal infection of the skin, has been observed in various stages of amphibian development. The primary source of infection among these animals is the water bodies where they typically reside and engage in various activities. The development of an effective antifungal agent necessitates the identification of at least three molecular targets: interaction with the cell wall using charges, binding to the plasma membrane, and finally, binding and interaction with a receptor protein at the membrane level, which plays a vital role in the fungus Bd.

Access restricted. Please log in or start a trial to view this content.

Protocol

1. Identify molecular target sequences in Bd

NOTE: The selection of amino acids in the sequence is based on recommendations from the literature. Key amino acids include K, W, G, A, V, I, and L. To find out which programs to use, check the Table of Materials.

  1. Retrieve target sequences from the NCBI database.
    1. Access the National Center for Biotechnology Information (NCBI) database at https://www.ncbi.nlm.nih.gov/.
    2. Select All Databases and choose Nucleotide as the search category.
    3. Enter Batrachochytrium dendrobatidis in the search bar and filter results by strain and species.
    4. Download and save the sequence in a text editor or FASTA format for further analysis.

2. Analyze the sequence by BLAST or SMART BLAST

  1. Perform a BLAST search.
    1. Open a web browser and go to BLAST (Basic Local Alignment Search Tool).
    2. Select Protein BLAST to compare the query sequence against protein databases https://blast.ncbi.nlm.nih.gov/Blast.cgi.  
  2. Verify target specificity.
    1. Ensure the selected molecular target is not present in other species. If the target appears in multiple species, repeat sequence selection until specificity is confirmed (see Figure 1).
      NOTE: This step ensures that no undesirable secondary molecular targets interfere with AMP or antifungal protein function
  3. Check for homology in protein sequences.
    1. Verify that the designed amino acid sequence does not match homologous proteins in other species.
      NOTE: This step confirms the antifungal function is specific to the intended targets

3. Design the antimicrobial protein sequence

NOTE: The amino acid content and sequence should be tailored to the desired protein function, specifically antifungal activity. The designed sequence should encode a protein with antimycotic effects against Bd. Check this on Antifp predict server28.

  1. Literature review and sequence design
    1. Conduct a comprehensive review of recent literature to identify key characteristics of Bd and potential antimicrobial mechanisms.
      NOTE: Previous studies describe antimicrobial peptides (AMPs) against Bd, synthesized by various animals and microorganisms. Shai's formula29provides a framework for designing synthetic antifungal peptides. This method should be used to optimize protein sequences for maximal antifungal activity.

4. Predict 3D structure using I-TASSER

NOTE: This protocol applies to the molecular target and the designed antimicrobial sequence.

  1. To access and submit sequences, visit the I-TASSER server for protein structure and function prediction https://zhanggroup.org/I-TASSER/. 
  2. Register using an institutional email to receive a username and password for tool access.
  3. Submit the molecular target or designed sequence in FASTA format, .txt file, or copy-paste directly into the input field.
  4. Assign a unique name to the sequence.
  5. Click Submit, and let the program analyze the sequence, providing an estimated processing time.
    NOTE: Pause the protocol until results are received. Processing time typically ranges from 24 to 72 h for small proteins (≤50 amino acids) and may take longer for larger proteins (up to 3,000 amino acids).
    Guidance on tool selection: Use I-TASSER for small proteins, such as designed antimicrobial peptides. For large proteins, use trRosetta, which is optimized for complex structural modeling.

5. Predict 3D structure using trRosetta

  1. Visit the trRosetta web server: https://yanglab.qd.sdu.edu.cn/trRosetta/.
  2. Enter the target receptor sequence in FASTA, .txt, or MSA format, or paste it directly into the server.
  3. Assign a name to the structurally predicted protein.
  4. Click Submit, ensuring the option to exclude templates is selected. Choose trRosettaX-single homologs to initiate the prediction process.
    NOTE: Pause the protocol until results are received.
  5. Wait for a confirmation email upon successful submission.
  6. Analyze the prediction results.
    1. Verify that the TM-score is a measure of model quality.
    2. Examine contact maps, distance maps per amino acid, and rotation maps for alpha and beta carbons at angles ω (omega), θ (theta), and φ (phi).
  7. Download results in ZIP format for further analysis.

6. Molecular docking and interaction analysis (using HADDOCK)

  1. Access the HADDOCK docking server: https://wenmr.science.uu.nl/haddock2.4/.
  2. Register an account using an institutional email, confirm the account, and log in.
  3. Return to the home page, select New Job, and start a submission.
  4. On the HADDOCK submission page, fill in the required information: job name, molecule number (in this case, 2). Select all.
  5. Upload Molecules for Docking. Upload PDB structures for the receptor and ligand. Select protein-ligand docking for both.
  6. Repeat the steps for ligand submission, select all chains, and leave the remaining settings as default.
  7. Click Next for the subsequent submission step.
    NOTE: In sections 1 and 2, add the list of amino acids involved in receptor-ligand interactions and mark them as active residues. Passive residues remain predefined. The active sites are the interaction sites and must be chosen according to the designed protein's characteristics and the receptor protein's exposed amino acids.
  8. Finalize and run the docking process.
    1. Click Next to keep the default docking parameters unchanged.
    2. Click Submit to begin processing.
      NOTE: Expected processing time ranges from 24 to 72 h.
  9. Review the docking results.
    1. The server automatically selects and ranks the best docking results. Analyze amino acid interactions using structural visualization software (e.g., PyMOL or UCSF Chimera).
  10. Download the final docked structures in PDB format for further study.
    NOTE: Refer to Supplemental File 1: Supplemental Figure S1, Supplemental Figure S2, Supplemental Figure S3, Supplemental Figure S4, and Supplemental Figure S5.

7. Measure distances (OPTIONAL) [see Supplemental File 1: Supplemental Figure S6, Supplemental Figure S7, and Supplemental Figure S8.]

  1. Analyze docking results using a structure visualization program (PDB). Open the molecular docking PDB file in UCSF Chimera and evaluate structural proximity and potential binding sites (measured distances).
  2. Double-click the Chimera icon to open PDB. Select File, Open, and choose a file from the document directory window.
  3. Identify atoms. Go to actions, atoms/ bonds, and select show.
  4. Select atoms closest to ligand and receptor. Select atom 1 with Ctrl+click and atom 2 with Ctrl+Shift+Click.
  5. Go to tools, structure analysis, and select distances.
  6. Click Create (verify atom selection). The window displays amino acids and locations of atoms 1 and 2.
    NOTE: The distance column shows angles between atoms.
  7. Download the table and save it on the local computer.
    NOTE: Refer to Supplemental File 1 for additional information.

Access restricted. Please log in or start a trial to view this content.

Results

The sequence under development utilizes the following formula: K(X)3KW(X)2K(X)2K, where (X) can be G, A, V, I, or L28. The search results obtained included protein sequences of more than 300 amino acids, which were selected due to their status as complete protein sequences, as opposed to partial sequences. It should be noted that complete sequences are essential for the study. The functions of each receptor/membrane channel were also searched for within the sequenc...

Access restricted. Please log in or start a trial to view this content.

Discussion

The bioinformatics approach employed in protein structural prediction has inherent limitations due to its reliance on computational programs. These programs attempt to approximate real protein structures but are constrained by factors such as temperature, solubility, and solvent presence, which may not fully replicate in vivo conditions. Additionally, the method depends on the accurate identification of key conserved proteins through multiple alignments and curated databases, making the reliability of these databases cru...

Access restricted. Please log in or start a trial to view this content.

Disclosures

The authors have no conflicts of interest to disclose.

Acknowledgements

Jimena R. Villarreal is a doctoral student from the Program of Doctorado en Ciencias Biológicas, Universidad Autónoma de Querétaro (UAQ), and has received a Consejo Nacional de Humanidades, Ciencias y Tecnologías (CONAHCyT) fellowship (1003112).

Access restricted. Please log in or start a trial to view this content.

Materials

List of materials used in this article
NameCompanyCatalog NumberComments
BLASTN/AN/AOnline server
ComputerAMDN/ASO: XP, W7, W8, W10, ubuntu, linux
HADDOCK serverBONVINLABN/AOnline server
HEX 5.2BBSRC-INRIAN/ADownload-Online server
I-TASSERZHANG LABN/ADownload-Online server
MUSCLEN/AN/AOnline server
NCBIN/AN/AOnline server
PYMOLSCHRÖDINGERN/ADownload-XP, W7, W8, W10, W11, LINUX
SWISS-MODELBIOZENTRUMN/AOnline server
trRosetta serverYANG LABN/AOnline server
USCF CHIMERARBVIN/ADownload-XP, W7, W8, W10, W11, LINUX

References

  1. McArthur, D. B. Emerging infectious diseases. Nurs Clin North Am. 54 (2), 297-311 (2019).
  2. Schilliger, L., Paillusseau, C., Francois, C., Bonwitt, J. Major emerging fungal diseases of reptiles and amphibians. Pathogens. 12 (3), 429(2023).
  3. Campbell, L. J., Garner, T. W. J., Hopkins, K., Griffiths, A. G. F., Harrison, X. A. Outbreaks of an emerging viral disease covary with differences in the composition of the skin microbiome of a wild United Kingdom amphibian. Front Microbiol. 10, 1245(2019).
  4. Lorch, J. M., et al. Snake fungal disease: An emerging threat to wild snakes. Philos Trans Royal Soc Lond BBiol Sci. 371 (1709), 20150457(2016).
  5. Costa, S., Lopes, I. Saprolegniosis in amphibians: An integrated overview of a fluffy killer disease. Jo F BaselSwitz. 8 (5), 537(2022).
  6. Lemieux-Labonté, V., Dorville, N. A. S. -Y., Willis, C. K. R., Lapointe, F. -J. Antifungal potential of the skin microbiota of hibernating big brown bats (Eptesicusfuscus) infected with the causal agent of White-Nose Syndrome. Front Microbiol. 11, 1776(2020).
  7. Hoyt, J. R., et al. Field trial of a probiotic bacteria to protect bats from White-Nose Syndrome. Sci Rep. 9 (1), 9158(2019).
  8. Origgi, F. C., et al. Ophiodimyces ophiodiicola, etiologic agent of snake fungal disease, in Europe since late 1950s. Emerg Infect Dis. 28 (10), 2064-2068 (2022).
  9. Van Rooijk, P., Martel, A., Haesebrouck, F., Pasmans, F. Amphibian chytridiomycosis: a review with a focus on fungus host interactions. Vet Res. 46, 137(2015).
  10. Fisher, M. C., Garner, T. W. J., Walker, S. F. Global emergence of Batrachochytrium dendrobatidis and amphibian Chytridiomycosis in space, time, and host. Ann Rev Microbiol. 63 (1), 291-310 (2009).
  11. Hancock, R., Brinkman, F. S. L. Function of Pseudomonas porins in uptake and efflux. Ann Rev Microbiol. 56 (1), 17-38 (2002).
  12. Carfrae, L., et al. Inhibiting fatty acid synthesis overcomes colistin resistance. Nat Microbiol. 8 (6), 1026-1038 (2023).
  13. Robson, B. Denovo protein folding on computers. Benefits and challenges. Comp Biol Med. 143, 105292(2022).
  14. Munita, J. M., Arias, C. A. Mechanisms of antibiotic resistance. Microbiol Spect. 4 (2), (2016).
  15. Zheng, W., et al. Folding non-homologous proteins by coupling deep-learning contact maps with I-TASSER assembly simulations. Cell Rep Met. 1 (3), 100014(2021).
  16. Yang, J., Zhang, Y. I-TASSER server: New development for protein structure and function predictions. Nuc Acid Res. 43 (W1), W174-W181 (2015).
  17. Du, Z., et al. The trRosetta server for fast and accurate protein structure prediction. Nat Prot. 16 (12), 5634-5651 (2021).
  18. Wang, W., Peng, Z., Yang, J. Single-sequence protein structure prediction using supervised transformer protein language models. Nat Comp Sci. 2 (12), 804-814 (2022).
  19. Pettersen, E., et al. UCSF Chimera-A visualization system for exploratory research and analysis. J Comp Chem. 25 (13), 1605-1612 (2004).
  20. Ghoorah, A. W., Devignes, M. -D., Smail-Tabbone, M., Ritchie, D. W. Protein docking using case-based reasoning. Proteins: Struct Funct Bioinf. 81 (12), 2150-2158 (2013).
  21. Battey, J., et al. Automated server predictions in CASP7. Proteins: Struct Funct Bioinf. 69 (S8), 68-82 (2007).
  22. Kryshtafovych, A., et al. Evaluation of the template-based modeling in CASP12. Proteins: Struct Funct Bioinf. 86 (S1), 321-334 (2018).
  23. Anishchenko, I., et al. Protein tertiary structure prediction and refinement using deep learning and Rosetta in CASP14. Proteins: Struct Funct Bioinf. 89 (12), 1722-1733 (2021).
  24. Kinch, L. N., Schaeffer, R. D., Kryshtafovych, A., Grishin, N. V. Target classification in the 14th round of the critical assessment of protein structure prediction (CASP14). Proteins: Struct Funct Bioinf. 89 (12), 1618-1632 (2021).
  25. Honorato, R. V., et al. Structural biology in the clouds: The WeNMR-EOSC ecosystem. FrontMolBiosci. 8, 729513(2021).
  26. Van Zundert,, et al. The HADDOCK2.2 Web Server: User-friendly integrative modeling of biomolecular complexes. JMolBiol. 428 (4), 720-725 (2016).
  27. Macindoe, G., Mavridis, L., Venkatraman, V., Devignes, M. -D., Ritchie, D. W. HexServer: An FFT-based protein docking server powered by graphics processors. Nucl Acid Res. 28 (Web Server), W445-W449 (2010).
  28. Agrawal, P., et al. In silico approach for prediction of antifungal peptides. Front Microbiol. 9, 323(2018).
  29. Shai, Y. Mode of action of membrane active antimicrobial peptides. Biopolymers. 66 (4), 236-248 (2002).

Access restricted. Please log in or start a trial to view this content.

Reprints and Permissions

Tags

Protein DesignIn Silico ProteinAntimicrobial PeptidesProtein Structure PredictionMolecular DockingI-TASSER ServertrRosetta PredictionHADDOCK DockingSynthetic AMPs