Method Article

Machine Learning Prediction of 3D Domain Swapping Proteins in Medicinal Plants

DOI:

10.3791/68519

August 15th, 2025

In This Article

Summary

Loading...
$$\rightleftharpoonup{xx}$$ $$\longleftharp{xx}$$, $$\longrightharp{xx}$$,

This study examines proteins involved in or predicted to be involved in 3D domain swapping from various medicinal plant genomes. It employs machine learning models to accurately predict 3D domain-swapping proteins and anticipate their functions and relevance to secondary metabolite production.

Abstract

Loading...
$$\rightleftharpoonup{xx}$$ $$\longleftharp{xx}$$, $$\longrightharp{xx}$$,

3D domain swapping is a protein structural phenomenon in which two or more protein subunits exchange identical structural subunits and form oligomers. Proteins that exhibit 3D domain swapping play a crucial role in various biological functions, such as secondary metabolite biosynthesis, and in coping with several biotic and abiotic stresses in medicinal plants. This study investigates the ability to predict 3D domain swapping patterns among the genomes of medicinal plants using random forest and K-nearest neighbor classifiers models, demonstrating accuracies of 91.6% and 88.7%, respectively. A total of 420 (31%) of sequences were predicted as being putatively involved in 3D domain swapping. An enrichment investigation was also carried out on the predicted 3D domain-swapped protein sequences from various medicinal plants for function annotation based on Gene Ontology (GO) terms, Kyoto Encyclopedia of Genes and Genomes (KEGG) pathways analysis, and their domain distribution in secondary metabolites biosynthesis pathways. Functional annotation of predicted sequences infers that 3D domain swapped sequences were involved in diverse molecular functions such as photosynthetic electron transport in photo system II and electron transporters, transferring electrons within the cyclic electron transport pathway of photosynthesis activity, oxidative phosphorylation, and gene regulation of environmental stresses (biotic and abiotic) by synthesizing secondary metabolites (terpenoids, alkaloids, and polyamines). These findings underscore the ability of machine learning to predict the involvement of proteins in the 3D domain-swapping phenomenon, their respective function, and their potential to facilitate drug discovery and bioengineering initiatives.

Introduction

Loading...
$$\rightleftharpoonup{xx}$$ $$\longleftharp{xx}$$, $$\longrightharp{xx}$$,

Computational methods have revolutionized protein research by enabling detailed analysis and predictions about protein structures, functions, and interactions. The precise identification and annotation of protein functions are essential for unraveling molecular mechanisms of life and carry profound importance for advancements in medicine and drug development. However, the inherent complexity and expense of experimental methods limit their scalability to accommodate the large sequence data. As a result, the development of computational methods for protein function prediction has emerged as a pivotal area in computational and molecular biology, addressing this gap through innovative large-scale approaches1.

3D domain swapping2 is a structural phenomenon in proteins where segments of a shared structure are exchanged between individual chains. In 1994, the introductory documentation about the mechanism of 3D domain swapping was found in the diphtheria toxin dimer3. However, the foundational principles of 3D domain swapping can be traced back four decades. Bovine pancreatic ribonuclease A (RNase A) has been observed to form dimers during lyophilization in acetic acid, through sophisticated chemical modification experiments4. In protein oligomerization, two or more protein chains interchange identical structural elements through flexible hinge regions. The portion of the protein exchanged between monomeric subunits is referred to as the swapped domain, which may consist of an entire globular domain, a loop, or a secondary structural element in certain proteins. Conversely, the regions that remain unchanged in their original positions within the monomers are termed non-swapped domains5. A coupling between non-swapped domains is defined as a non-swapped domain interface (NSDI) shown in Figure 1. In 3D domain swapping, the relationship between the non-swapped domain of one protein subunit and the already swapped domain of another subunit is referred to as the swapped domain interface (SDI). Another significant aspect of this phenomenon is the hinge region, a flexible linker segment that connects the non-swapped and swapped domains. This hinge region plays an essential function in helping the movement of 3D domain swapping and serves as a conformational switch, enabling the structural reconfiguration necessary for domain swapping to occur. 3D domain swapping has been implicated in various biological processes, including protein assembly and functional regulation5. It is also associated with certain protein misfolding diseases, where aberrant swapping can lead to the formation of aggregates or amyloid fibrils.

Domain swapping diagram showing protein monomers and interfaces in molecular biology process.
Figure 1: Structural representation of 3D domain-swapping. The representation illustrates the exchange of identical structural elements between two protein monomers via a flexible hinge region, resulting in the formation of a domain-swapped dimeric or oligomeric assembly Please click here to view a larger version of this figure.

Different types of 3D domain swapping have been identified based on the nature of the swapped domains and the resulting oligomeric structures6. In the year 2002, Eisenberg and their colleague identified three types of 3D domain swapping: Bonafide domain swapping (BDS), quasi-domain swapping (QDS), and candidate for 3D domain swapping (CDS)7. The most common protein class is bona fide domain swapping. It refers to a state where both dimer and monomer molecules are present in a stable form, where the dimer is supposed to adopt a domain-swapped configuration while the monomer is supposed to adopt a closed configuration. In quasi domain swapping, a protein is known to be present in the oligomeric state, but its homologous structure is known to be present in the monomeric state. In CDS, it only confirms protein classification under the domain-swapped categories, whereas structural information of involved monomers or their monomeric homologues is not present8. In this procedure, monomers or their monomeric homologues are absent; rather, heterologous protein molecules are present. Table 1 shows an example of these three categories9.

Table 1: Types of 3D domain swapping with example. Please click here to download this Table.

Subsequent studies revealed numerous domain-swapped structures, paving the way for the understanding of the 3D domain swapping concept. The earliest structural evidence supporting this phenomenon was observed in the molecule of the Cro repressor protein from bacteriophage λ, which forms a dimeric structure through the exchange of its C-terminal strands. In 1996 the researcher found monomeric Cro molecule was involved in 3D domainswapping10.Other structure11, such as βB2-crystallin12, human CksHs213, beef liver catalase14, and recombinant human interleukin-515, were also reported as possible 3D domain-swapped structures. Based on the swapped domain position within the protein molecules, 3D domain swapping is categorized into three types: C-terminal domain swapping, N-terminal domain swapping, and the relatively rare central domain swapping. Authors use a complete human genome to predict a domain-swapped case. They use Random Forest and Support Vector Machine as binary classifiers with 81.7% and 73.9 % accuracies, respectively. Nearly 44% of the protein sequence was predicted as 3D domain-swapped in the human genome6. Enrichment analysis was carried out on predicted cases for their domain distribution, disease distribution, and functional annotation based on Gene Ontology (GO). Another approach investigates the complete Genome-Wide Analysis of Ocimum tenuiforum by using the Random Forest approach and found nearly 25% of the protein sequences in Ocimum tenuiforum are predicted to be involved in the 3D domain-swapping. Researchers also perform functional annotation by using GO term association and their protein domain family association and found only 1158 sequences were involved in abiotic stress16.

Plants demonstrate a variety of distinctive protein families, with specific proteins recognized for having the ability to undergo structural rearrangements, including 3D domain swapping. 3D domain swapping may influence biotic and abiotic stress conditions with the production of secondary metabolites or other biologically relevant pathways that hold several pharmacological applications; therefore, they provide a significant focus for this investigation. The crystal structure of Wal1 demonstrated a domain-swapped dimeric configuration, featuring two dimers within the asymmetric unit and the structural rearrangement found in cystatin C. Phytocystatins play a significant role in abiotic stress, and they also improve crop resistance. A total of 20,121 3D domain-swapped predicted proteins has been identified from available literature and from various relevant repositories, from which 17,552 are of plant origin and 2569 are from non-plant organisms17,18.

In the literature, various other examples of 3D domain swapping prediction of different plants using a machine learning approach have been reported. Such as Arabidopsis thaliana showing 33.7% (4058 out of 12,033 reviewed sequences), Medicago truncatula 20.9% (39 out of 186 reviewed sequences), Solanum tuberosum 36.5% (146 out of 400 reviewed sequences), Solanum lycopersicum 25.5% (108 out of 423 reviewed sequences), and Ocimum tenuiforum 15.5% (5706 out of 36841 reviewed sequences)6. Computational methods have become invaluable in exploring the structural and functional aspects of 3D domain-swapped proteins. These approaches facilitate the prediction of sequences on the basis of the best possible features19, annotation, and enrichment analysis, offering insights into the fundamental molecular mechanisms. The prediction of 3D domain swapping in various protein sequences was accomplished by using a support vector machine (SVM)-based classifier. This approach was developed by integrating sequence and structural features, yielding prediction accuracies of 76.33% on the training dataset and 73.81% on the test dataset, signifying its potential in identifying domain swapping tendencies in proteins20. The main reason for choosing KNN over SVM is that KNN has a much simpler training process. It only needs one main parameter, K (it is the number of nearest neighbors considered when making a prediction), to be set, while SVM requires careful tuning of several parameters like the kernel type, C, and gamma. Additionally, KNN handles multi-class classification more easily, whereas SVM usually needs more complex strategies, like one-vs-one strategies.

Machine learning for protein structural prediction
The prediction of protein structure involves the inference of a protein's three-dimensional shape from its FASTA sequence. Recent advancements in this field have been significantly driven by the application of various machine learning techniques to evolutionary data21,22. Early approaches for extracting information from co-evolutionary data relied on machine learning methods. However, more recent strategies, particularly those utilizing deep residual networks, have demonstrated superior performance in predicting potential target23. Alphafold is a deep learning-based database, whereas Rosetta is a physics-based energy function tool. These tools are used to predict the full 3D atomic structure that can be approximated to experimental structures. They only provide atomic resolution structural coordinates, but these tools do not specify the 3D domain swapping events. In contrast, the suggested approach is not a full 3D-structure predictor instead, it only predicts whether a protein is likely to undergo 3D domain swapping or not24,25.

Medicinal plants have been utilized for centuries as natural resources for the prevention and treatment of various diseases, attributed to their bioactive compounds. They are essential in traditional medicine and contribute to the development of modern pharmaceutical drugs. However, there is no significant work has been carried out in the field of protein domain swapping of medicinal plants for drug discovery. Algorithms, including machine learning and their ensemble models, recognize patterns and relationships in sequence data to predict 3D domain swapping. The reason behind choosing the medicinal plants for this study is that it can detect functional diversity of a protein in plants involved in 3D domain swapping, resulting in the understanding of stress response, pathogen defense, and metabolic biosynthesis. The technical challenges associated with determining protein 3D domain swapping in large numbers and complex, advanced oligomeric conformations by utilizing NMR or crystallography techniques highlight the necessity for developing advanced computational approaches.

The overall goal of the proposed research work is to employ a computational methodology by using Random Forest (RF)26 and K-nearest neighbor (KNN)27 algorithms for their protein sequence structure prediction task and complete genome-wide analysis of swapped cases in various medicinal plant sequences. The Random Forest (RF) is a binary classifier; it is a robust and versatile machine learning algorithm widely used in protein sequence analysis. It operates by constructing an ensemble of decision trees (DT) and attains quite a high accuracy in both the training and testing datasets. For protein sequence tasks, RF can analyze a variety of features, including physicochemical properties, sequence composition, secondary structure, and evolutionary information derived from sequence alignments or profiles. It excels in handling large, noisy datasets and provides feature importance metrics28. The K-Nearest Neighbors (KNN) classifier is a simple yet effective machine learning algorithm widely applied in protein sequence analysis. It works by classifying an input sequence based on the majority class of its k closest neighbors in the feature space. In protein sequence analysis, KNN can be used to predict functional categories, structural properties, and sub-cellular localization. Features for KNN classification often include amino acid composition, sequence motifs, evolutionary profiles, or physicochemical attributes. KNN is a popular choice for protein analysis because of its simplicity; interpretability and effectiveness make it a valuable tool for protein sequence analysis29. Together, these algorithms provide complementary strengths, enabling a comprehensive computational framework for the study of 3D domain swapping in various medicinal plant sequences. These two models only predict possible cases that may undergo domain-swapping; it does not predict the full 3D structural coordinates or which part or residue is particularly involved in this process of swapping. This is a matter of further research, as identifying specific regions or residues that are involved in swapping to clarify the mechanism and functional aspects of 3D domain-swapping events.3D domain-swapping comprises a variety of structurally distinct phenomena, such as closed-loop configurations, open-ended swaps, and other differences that include different hinge regions and domain architectures. In light of this, the suggested approach only categorizes all kinds of 3D domain-swapping events under a unified classification. This selection was initially driven by the limited availability of annotated data that can specify a particular class of 3D domain-swapping. This is the first kind of attempts on various medicinal plant-based datasets, and we feel that distinguishing between different swapping mechanisms could enhance the predictive accuracy of the model.

An Enrichment analysis in the medicinal plants helps identify key genes, proteins, and pathways involved in bioactive compound synthesis and stress response. It provides insights into molecular mechanisms. These analyses in medicinal plants investigate essential genes, proteins, and biological processes associated with the synthesis of bioactive chemicals, stress resilience, and disease resistance. Functional annotation of selected proteins, such as Gene Ontology (GO) and KEGG pathway analysis, uncovers molecular mechanisms underlying secondary metabolite production and environmental stress responses. This approach substantially aids in drug discovery, improves crop resilience, and elucidates plant metabolic pathways for sustainable agriculture and medicinal advancements.

Novel contributions of the study
This study highlights the capabilities of machine learning to promote more precise and efficient models to predict protein structural patterns in biological datasets.

This research integrates novel features into machine learning algorithms, improving their capacity to predict protein functions with greater clarity. Such features provide an enhanced understanding of proteins' structure-function relationships, aiding in more precise protein design and optimization in protein engineering.

Enrichment analysis of predicted proteins helps to find out the key biological pathways, cellular process and molecular functions providing valuable targets for biomarker discovery, drug design and understanding disease mechanism. This analysis enhances the precision in pathways-based interventions and therapeutic target identifications.

Access restricted. Please log in or start a trial to view this content.

Protocol

Loading...
$$\rightleftharpoonup{xx}$$ $$\longleftharp{xx}$$, $$\longrightharp{xx}$$,

NOTE: This segment provides a comprehensive outline of the suggested methodology, which encompasses six primary steps: (a) data collection, (b) feature selection, (c) data pre-processing, (d) data post-processing (e) model development (f) result evaluation of proposed model and, (g) enrichment analysis at different level of positively predicted sequences. The Core i5-10500 is a 10th-generation processor that has been used in this research work. It has six cores and 12 threads, with a base frequency of 3.10 GHz and a maximum turbo speed of 4.50 GHz. It comes equipped with 12 MB of Intel Smart Cache, supports DDR4-2666 memory, and includes UHD Graphics 630 for integrated visuals. Built for efficient multitasking and productivity, it is compatible with the LGA 1200 socket and operates at a 65W TDP.

1. Data collection

  1. Use3d3dswap-pred (https://caps.ncbs.res.in/3dswap-pred/3dswap-pred.html) and 3DSwap+ (https://caps.ncbs.res.in/3Dswapplus/index.html)3030 Supplementary Table 1, Supplementary Table 2, Supplementary File 1, Supplementary File 2). Use a total of 573 manually curated PDB entries of 3D domain-swapped molecules, primarily from medicinal plants.
  2. Employ the Best Representative Profile (BRP) method to construct the negative dataset by assigning a Best Representative Sequence (BRS) to each Pfam31 protein family (http://pfam.xfam.org/). Search a total of 10,112 structural sequences against all Pfam BRPs using HMMER32 (https://www.ebi.ac.uk/Tools/hmmer/) with an E-value threshold of 0.001. Process the resulting sequence using DIAL to identify structural domains(https://bioinformaticshome.com/db/tool/DIAL). Include 575 PDB entries as a negative dataset (training).
  3. Include 314 protein sequences derived through the BRP method and 261 manually curated non-3D domain-swapped sequences identified by 3DSwap+ as a negative dataset.
  4. Retrieve a total of 1,355 reviewed protein sequence entries from various medicinal plants from UniProt33 (https://www.uniprot.org/). Include 13 medicinal plants in this study (testing/prediction dataset): Citrus sinensis (RS-102), Vitis vinifera (RS-226), Mentha (RS-28), Cannabis sativa (RS-21), Acorus clamus (RS-511), Coffea arabica (RS-104), Papaver somniferum (RS-58), Hibiscus (RS-88), Jasmine (RS-89), Thyme (RS-30), Thymus (RS-14), Illicium oligandrum (RS-72), and Citrus limon (RS-12). Phylogenetic tree is given in Supplementary Figure 1.

2. Using the feature for model creation

  1. Utilize a comprehensive feature set comprising 453 features for predicting protein sequences in medicinal plants. Include 439 established features and 16 novel features, carefully selected based on a thorough literature review.
  2. Use the AAindex database34(https://www.genome.jp/aaindex/) to determine amino acid physicochemical properties. Apply the WEKA35 machine learning platform (https://www.waikato.ac.nz/int/research/institutes-centres-entities/institutes/artificial-intelligence-institute/research/software/) for feature selection. Refer to Table 2 for the newly incorporated features.

Table 2: List of newly added features along with their descriptions. This table summarizes the novel features engineered during the feature extraction phase to enhance model performance in predicting 3D domain swapping. Each feature captures specific sequence-based, physicochemical, or structural properties of the proteins that are hypothesized to influence their propensity for domain swapping. Detailed descriptions are provided to clarify the biological relevance and computational derivation of each feature. Please click here to download this Table.

3. Data pre- and post-processing

NOTE: Data preprocessing. Data preprocessing is an essential step in machine learning that involves preparing raw data for analysis. This includes cleaning the data (e.g., removing duplicates, handling missing values, etc.).

  1. Use Supplementary Coding File 1 for encoding categorical variables into numeric form, which includes four major steps: (1) Define score dicts, (2) process fasta sequence by seq I0. parse, (3) Create data frame, (4) numerical analysis of 3D domain-swapped features.
  2. Refine and interpret the model's output through data post-processing. Recalibrate predictions, aggregate results, and apply a threshold for classification. Classify the bimodal 3D domain swapping dataset using a threshold value β = 0.5 (range 0-1), as supported by previous studies36,37,38.
  3. Partition the dataset with a 70:30 train-test split, allocating 70% for model training and 30% for independent evaluation.
  4. Standardize the training data using the Standard Scaler method to adjust feature values to a mean of 0 and a standard deviation of 1, ensuring better compatibility with machine learning estimators. Conduct feature selection to identify the most impactful variables.
    NOTE: Standardizing a dataset is a widely used prerequisite for many machine learning algorithms to ensure optimal performance. Data that is far from normal distribution can negatively affect the performance of machine learning models39.
  5. Do validation. To ensure robust and unbiased model evaluation, K-fold cross-validation was also implemented Supplementary Table 3.
  6. Partition the dataset into K subsets. Train and evaluate the model iteratively on K=5 subsets while using the remaining subset for independent testing.

4. Creation and implementation of the model

  1. Implement and fine-tune Random Forest (RF) and K-Nearest Neighbors (KNN) algorithms to optimize predictive performance. Train, validate, and test the models using 573 and 575 protein sequences for training, and 1,355 protein sequences for testing.
  2. Use a Python script (Supplementary Coding File 1) to extract numerical feature values based on the selected protein sequence. Save these numerical values for both datasets in CSV files and submit them to the RF and KNN classifiers for binary classification model generation.
  3. Use these models to distinguish between 3D domain-swapped and non-3D domain-swapped proteins. Refer to Figure 2 and Figure 3 for the generalized framework for predicting 3D domain swapping.
  4. Apply hyper parameters such as n_estimators=20, max_depth=4, and random_state=42 for RF model.
  5. Apply the hyper parameter neighbors=5 for the KNN classifier model.
    NOTE: These parameter settings were fine-tuned to enhance the models' performance on the manually curated medicinal plant dataset.

Ensemble learning diagram showing decision trees, classification, and KNN data processing flowchart.
Figure 2: Illustrative schematic representation. The representation shows the Random Forest and K-Nearest Neighbor (KNN) machine learning models, showcasing their algorithmic structure and functional workflow in the context of binary classification tasks Please click here to view a larger version of this figure.

Data processing workflow diagram including dataset acquisition, feature extraction, model creation.
Figure 3: Machine learning-based workflow. The diagram illustrates a machine learning-based workflow for predicting 3D domain-swapping in medicinal plant proteins. The process begins with dataset acquisition, combining the benchmark and meta datasets. Feature extraction involves deriving both existing and novel features from protein sequences. In dataset pre-processing, FASTA sequences are converted into numerical values using sequence parsing and dictionary mapping to create structured data frames, including specific 3D domain-swap features. Post-processing includes assigning domain swap status using a threshold (β = 0.5), splitting data into training and testing sets (70-30 ratio), standardizing features, and performing feature selection with k-fold validation. Two machine learning models, Random Forest (RF) and K-Nearest Neighbors (KNN), are trained, and their performance is evaluated using model metrics along with AUROC and AUPRC to assess predictive accuracy and robustness Please click here to view a larger version of this figure.

5. Static assessment and model evaluations of ML classifiers

  1. Implement and fine-tune Random Forest (RF) and K-Nearest Neighbors (KNN) algorithms to optimize predictive performance. Train, validate, and test the models using 573 and 575 protein sequences for training, and 1,355 protein sequences for testing.
    Sensitivity formula X_sen=TP/(TP+FN) equation for statistical analysis.    (1)
    Specificity formula, X_spe = TN/(TN+FP), mathematical equation in statistical analysis.    (2)
    MCC formula illustrating statistical analysis of binary classification results.    (3)
    Accuracy calculation formula, ACC, mathematical equation for data analysis in research studies.    (4)
    Precision formula, TP/(TP+FP), statistical analysis method in scientific research.    (5)
    F1 score equation; 2×Precision×Recall/(Precision+Recall); formula diagram; performance metric.    (6)
    NOTE: Define TP (True Positive) as the percentage of sequences correctly predicted as domain-swapped by the models. Define TN (True Negative) as the percentage of sequences correctly predicted as non-domain-swapped. Define False Positive (FP) as the cases where non-domain-swapped proteins are incorrectly predicted as domain-swapped. Define False Negative (FN) as the cases where domain swapped proteins are incorrectly predicted as non-domain-swapped.
  2. Calculate Xsen as the ratio of true positives accurately identified, and Xspe as the ratio of true negatives correctly identified by the model.
    NOTE: Use MCC to evaluate the quality of binary classifications by analysing True Positives, True Negatives, False Positives, and False Negatives from the confusion matrix.
  3. Calculate Accuracy (ACC) to measure the proportion of correct predictions among the total predictions.
    NOTE: Precision evaluates a proportion of correctly identified positives among all the predicted positives, whereas the F1-score is the harmonic mean of Precision and Sensitivity, equilibrating False Positives and False Negatives.
  4. Use AUC to evaluate the performance of classification models in binary classification tasks. Calculate AUC (Area Under the Curve) using the Scikit-learn Python library40, based on positively and negatively classified proteins.
  5. Use the test data to evaluate these parameters.

6. Implementation of model for the prediction of 3D domain swapping on various medicinal plant species

  1. Apply RF and KNN on a total of 1,355 reviewed sequences (prediction dataset) from 13 different medicinal plants, including Citrus sinensis, Mentha, Vitis vinifera, Thyme, Thymus vulgaris, Jasmine, Hibiscus, Papaver somniferum, Illicium oligandrum, Citrus limon, Acorus calamus, Cannabis sativa, and Coffea arabica.
    NOTE: These proteins are associated with various functions like Citrus sinensis, Mentha, thyme helps indigestion. Vitis vinifera, Jasmine, Hibiscus is a very rich source of antioxidants and they also improve skin health. Illicium oligandrum is known for its antifungal and antibacterial properties etc.

7. Enrichment investigation of positively predicted 3D domain swapped proteins from medicinal plant

  1. Map the accession codes of these protein sequences to secondary metabolite categories using data from UniProt.
  2. Extract gene ID from predicted sequences.
  3. Open KEGG online webserver41 (https://www.genome.jp/kegg/) to conduct KEGG pathway enrichment analysis by pasting the individual gene ID or protein name to check their respective pathways.
  4. Perform a comparative analysis through by contrasting the predicted 3D domain-swapping proteins against a prediction dataset to detect statistically significant overrepresentation of specific gene ontology (GO) terms and biological pathways. Paste the gene ID of the predicted sequence into ShinyGO v0.742 (https://bioinformatics.sdstate.edu/go74/) and select the desired function or pathway category (biological function, Cellular component, molecular function, KEGG, etc.).
    NOTE: Visualize these annotations using an online cloud-based platform43 that allows users to write and execute Python code.

Access restricted. Please log in or start a trial to view this content.

Results

Loading...
$$\rightleftharpoonup{xx}$$ $$\longleftharp{xx}$$, $$\longrightharp{xx}$$,

Proteins exhibiting 3D-domain swapping are frequently linked to a variety of biological functions. However, a systematic, genome-wide analysis of protein sequences that were involved in 3D domain swapping remains largely unexplored. In this research, we conducted an initial investigation to putatively predict 3D domain-swapped proteins from 13 medicinal plants, focusing on their association with secondary metabolite domains, KEGG pathways, and Gene Ontology (GO) terms. A comprehensive analysis of the accessible literatur...

Access restricted. Please log in or start a trial to view this content.

Discussion

Loading...
$$\rightleftharpoonup{xx}$$ $$\longleftharp{xx}$$, $$\longrightharp{xx}$$,

The understanding of 3D domain swapping in proteins across their complete genome holds significant importance in various fields. Machine learning approaches, especially random forest and K-nearest neighbor26,27, have been employed to predict 3D domain swapping directly from protein sequence data at the genome level. These two ML models include various steps such as dataset collection, feature selection, data pre/post processing, implementation of customized model...

Access restricted. Please log in or start a trial to view this content.

Disclosures

Loading...
$$\rightleftharpoonup{xx}$$ $$\longleftharp{xx}$$, $$\longrightharp{xx}$$,

The authors declare that they have no known competing financial interests or personal relationships that could have appeared to influence the work reported in this paper.

Acknowledgements

Loading...
$$\rightleftharpoonup{xx}$$ $$\longleftharp{xx}$$, $$\longrightharp{xx}$$,

No funding was received for this research.

Access restricted. Please log in or start a trial to view this content.

Materials

List of materials used in this article
NameCompanyCatalog NumberComments
3DSwap+https://caps.ncbs.res.in/3Dswapplus/index.html 
PfamEMBL-EBIhttp://pfam.xfam.org/: 
HMMEREMBL-EBIhttps://www.ebi.ac.uk/Tools/hmmer/:  
Aaindexhttps://www.genome.jp/aaindex/:
DIALBioinformaticshome.comhttps://bioinformaticshome.com/db/tool/DIAL: 
WEKAThe University of Wekatohttps://www.waikato.ac.nz/int/research/institutes-centres-entities/institutes/artificial-intelligence-institute/research/software/
 KEGGhttps://www.genome.jp/kegg/ :
ShinyGOhttps://bioinformatics.sdstate.edu/go/ :
uniprotUniprothttps://www.uniprot.org/ :

References

Loading...
$$\rightleftharpoonup{xx}$$ $$\longleftharp{xx}$$, $$\longrightharp{xx}$$,
  1. Radivojac, P., et al. A large-scale evaluation of computational protein function prediction. Nat Methods. 10 (3), 221-227 (2013).
  2. Bennett, M. J., et al. 3D domain swapping: a mechanism for oligomer assembly. Protein Sci. 4, 2455-2468 (1995).
  3. Bennett, M. J., et al. Domain swapping: entangling alliances between proteins. Proc Natl Acad Sci U S A. 91, 3127-3131 (1994).
  4. Liu, Y., et al. The crystal structure of a 3D domain-swapped dimer of RNase A at a 2.1-Å resolution. Proc Natl Acad Sci U S A. 95 (7), 3437-3442 (1998).
  5. Straub, J. E., et al. Principles governing oligomer formation in amyloidogenic peptides. Curr Opin Struct Biol. 20, 187-195 (2010).
  6. Upadhyay, A. K., et al. Genome-wide prediction and analysis of 3D domain-swapped proteins in the human genome from sequence information. PLoS One. 11 (7), e0159627(2016).
  7. Liu, Y., et al. 3D domain swapping: as domains continue to swap. Protein Sci. 11, 1285-1299 (2002).
  8. Schlunegger, M. P., et al. Oligomer formation by 3D domain swapping: a model for protein assembly and misassembly. Adv Protein Chem. 50, 61-122 (1997).
  9. Rousseau, F., et al. Implications of 3D domain swapping for protein folding, misfolding and function. Adv Exp Med Biol. 747, 137-152 (2012).
  10. Anderson, W. F., et al. Structure of the Cro repressor from bacteriophage λ and its interaction with DNA. Nature. 290 (5809), 754-758 (1981).
  11. Albright, R. A., et al. High-resolution structure of an engineered Cro monomer shows changes in conformation relative to the native dimer. Biochemistry. 35, 735-742 (1996).
  12. Bax, B., et al. X-ray analysis of beta B2-crystallin and evolution of oligomeric lens proteins. Nature. 347, 776-780 (1990).
  13. Parge, H. E., et al. Human CksHs2 atomic structure: A role for its hexameric assembly in cell cycle control. Science. 262, 387-395 (1993).
  14. Fita, I., et al. The NADPH binding site on beef liver catalase. Proc Natl Acad Sci U S A. 82 (6), 1604-1608 (1985).
  15. Milburn, M. V., et al. A novel dimer configuration revealed by the crystal structure at 2.4 Å resolution of human interleukin-5. J Biol Chem. 268, 2117-2120 (1993).
  16. Upadhyay, A. K., et al. Genome-wide analysis of domain-swap predicted products in the genome of anti-stress medicinal plant: Ocimum tenuiflorum. Bioinformat Biol Insights. 13, 1177932218821362(2019).
  17. Simpson, G. A., et al. Crystal structure and interconversion of monomers and domain-swapped dimers of the walnut tree phytocystatin. Biochim Biophys Acta Prot Proteom. 1872 (2), 140975(2024).
  18. Tran, L. H., et al. 3D domain swapping dimerization of the receiver domain of cytokinin receptor CRE1 from Arabidopsis thaliana and Medicago truncatula. Front Plant Sci. 24 (12), 756341(2021).
  19. Shameer, K., et al. Insights into protein sequence and structure-derived features mediating 3D domain swapping mechanism using support vector machine-based approach. J Bioinform Comput Biol. 4, 33-42 (2010).
  20. Shameer, K., et al. 3dswap-pred: prediction of 3D domain swapping from protein sequence using Random Forest approach. Protein Pept Lett. 19, 1010-1020 (2012).
  21. Tasnim, F. Protein sequence classification through deep learning and encoding strategies. Proc Comp Sci. 238, 876-881 (2024).
  22. Xu, Y., et al. Deep dive into machine learning models for protein engineering. J Chem Info Modeling. 60 (6), 2773-2790 (2020).
  23. Lilhore, U. K., et al. Optimizing protein sequence classification: integrating deep learning models with Bayesian optimization for enhanced biological analysis. BMC Med Inform Decis Mak. 24, 236(2024).
  24. Jumper, J., et al. Highly accurate protein structure prediction with AlphaFold. Nature. 596 (7873), 583-589 (2021).
  25. Du, Z., et al. The trRosetta server for fast and accurate protein structure prediction. Nat Protoc. 16 (12), 5634-5651 (2021).
  26. Genuer, R., et al. Random forests: Some methodological insights. arXiv. , (2008).
  27. Guo, G., et al. KNN model-based approach in classification. Lect Notes Comput Sci. 2888, 986-996 (2003).
  28. Kathuria, C., et al. Predicting the protein structure using random forest approach. Procedia Comput Sci. 132, 1654-1662 (2018).
  29. Gui, Y., et al. Application of K-nearest neighbors in protein-protein interaction prediction. Highlights Sci Eng Technol. 2, 125-131 (2022).
  30. Joseph, A. P., Shingate, P., Upadhyay, A. K., Sowdhamini, R. 3PFDB+: improved search protocol and update for the identification of representatives of protein sequence domain families. Database. 2014, bau026(2014).
  31. Finn, R. D., et al. The Pfam protein families database. Nucl Acids Res. 41, D211-D222 (2013).
  32. Mistry, J., et al. Challenges in homology search: HMMER3 and convergent evolution of coiled-coil regions. Nucl Acid Res. 41, e121(2013).
  33. Apweiler, A. UniProt: The universal protein knowledgebase. Nucl Acids Res. 46 (5), 2699-2699 (2018).
  34. Kawashima, S., et al. AAindex: amino acid index database, progress report 2008. Nucl Acids Res. 36, D202-D205 (2008).
  35. Frank, E., et al. Data mining in bioinformatics using WEKA. Bioinformatics. 20, 2479-2481 (2004).
  36. Pedregosa, F., et al. Scikit-learn: Machine Learning in Python. J Mach Learn Res. 12, 2825-2830 (2011).
  37. Wu, C., et al. Prediction of DNA methylation site status based on fusion deep learning algorithm. IEEE AEMCSE Proc. 5, 180-183 (2022).
  38. Wilhelm, T., et al. Phenotype prediction based on genome-wide DNA methylation data. BMC Bioinform. 15, 1-15 (2014).
  39. Uddin, S., et al. Dataset meta-level and statistical features affect machine learning performance. Sci Rep. 14 (1), 1F670(2024).
  40. Pedregosa, F., et al. Scikit-learn: Machine Learning in Python. J Mach Learn Res. 12, 2825-2830 (2011).
  41. Kanehisa, M., et al. Kyoto encyclopedia of genes and genomes. Nucl Acid Res. 28 (1), 27-30 (2000).
  42. Ge, S. X., et al. ShinyGO: a graphical gene-set enrichment tool for animals and plants. Bioinformatics. 36 (8), 2628-2629 (2020).
  43. Bisong, E. Google Colaboratory. Building Machine Learning and Deep Learning Models on Google Cloud Platform. , Apress. Berkeley, CA. (2019).
  44. Kirby, G. W. Biosynthesis of the morphine alkaloids. Science. 155 (3759), 170-173 (1967).
  45. Toffolatti, S. L., et al. Role of terpenes in plant defence to biotic stress. Biocont Agents Sec Metabol. , 401-417 (2021).
  46. Boncan, D. A. T., et al. Terpenes and terpenoids in plants: interactions with environment and insects. Int J Mol Sci. 21 (19), 7382(2020).
  47. Mansouri, H., et al. The response of terpenoids to exogenous gibberellic acid in Cannabis sativa L. at vegetative stage. Acta Physiol. Plant.33, 1085-1091 (2011).
  48. Pál, M., et al. Speculation: Polyamines are important in abiotic stress signalling. Plant Sci. 237, 16-23 (2015).
  49. Mattoo, A. K., et al. Higher polyamines restore and enhance metabolic memory in ripening fruit. Plant Sci. 174 (4), 386-393 (2008).

Access restricted. Please log in or start a trial to view this content.

Reprints and Permissions

Request permission to reuse the text or figures of this JoVE article

Request Permission

Tags

3D Domain SwappingMachine Learning PredictionMedicinal PlantsProtein OligomerizationRandom ForestK Nearest NeighborFunctional AnnotationSecondary Metabolite BiosynthesisGene OntologyKEGG Pathways
Video Coming Soon

Related Articles