Review Article

Multi-Omics and Integrative Analytics in Natural Products Discovery

DOI:

10.3791/69458

November 28th, 2025

In This Article

Summary

Loading...
$$\rightleftharpoonup{xx}$$ $$\longleftharp{xx}$$, $$\longrightharp{xx}$$,

This review emphasizes cutting-edge multi-omics methodologies, including metabolomics, genomics, transcriptomics, and proteomics, which are integrated with artificial intelligence and machine learning, as well as network analysis, to enhance the discovery and functional validation of new natural products.

Abstract

Loading...
$$\rightleftharpoonup{xx}$$ $$\longleftharp{xx}$$, $$\longrightharp{xx}$$,

Natural products (NPs) have long been an essential source of new bioactive compounds for drug discovery; however, traditional methods for screening and isolating these compounds can be slow and often yield diminishing returns. Fortunately, advanced multi-omics and computational approaches present powerful solutions to these challenges. This review highlights innovative methodologies that integrate metabolomics, genomics, transcriptomics, and proteomics with bioinformatics and analytical chemistry to accelerate NP discovery. For instance, untargeted metabolomics platforms like high-resolution liquid chromatography-tandem mass spectrometry (LC-MS/MS) and Global Natural Products Social (GNPS) molecular networking allow for comprehensive profiling of new compounds, while targeted isotope-labeling strategies enhance this process. Additionally, genome and metagenome mining tools such as antibiotics and secondary metabolite analysis shell (antiSMASH), Deep Biosynthetic Gene Cluster (DeepBGC), and Pipeline for Reconstructing Integrated Syntheses of Metabolites (PRISM) quickly identify biosynthetic gene clusters (BGCs) in both cultured and uncultured organisms, often using heterologous expression to validate products. Transcriptomic analyses, including RNA sequencing (RNA-seq), co-expression networks, and fluxomics, help clarify how pathways are regulated, while quantitative proteomics techniques like tandem mass tags/isobaric tags for relative and absolute quantitation (TMT/iTRAQ) and label-free methods, along with chemoproteomics approaches such as cellular thermal shift assay and thermal proteome profiling (TPP), uncover molecular targets and their mechanisms of action. This review also places significant emphasis on the role of artificial intelligence (AI) and machine learning (ML) in integrating multi-omics data, spanning activities from constructing gene-metabolite correlation networks to leveraging knowledge graphs and graph neural networks for data fusion and functional prediction. Finally, this review concludes by discussing the synergistic benefits of multi-omics for natural-product discovery, addressing current technical challenges, and exploring future directions toward high-throughput, intelligent data integration for next-generation NP research.

Introduction

Loading...
$$\rightleftharpoonup{xx}$$ $$\longleftharp{xx}$$, $$\longrightharp{xx}$$,

Natural products, which are molecules produced by various organisms, including microbes and plants, have long served as a vital source of medicines, ranging from antibiotics like penicillin to anticancer drugs such as Taxol. A significant portion of our current antibiotics and chemotherapy agents is derived from these natural compounds1. However, the traditional methods used to discover new NPs, such as culturing microbes and employing bioassay-guided isolation, have become increasingly difficult. In the late 20th century, pharmaceutical interest in NPs declined as companies faced diminishing returns; many compounds that were easy to find had already been rediscovered, and the technical challenges associated with screening and characterizing these compounds made the process labor-intensive. Fortunately, this trend is now reversing. Recent advancements in omics and analytical technologies are revitalizing the search for NPs2,3. For instance, high-throughput DNA sequencing enables researchers to explore genomes for hidden biosynthetic gene clusters (BGCs), while ultra-sensitive mass spectrometry (MS) can identify tiny amounts of novel metabolites within complex mixtures. These technological innovations arrive at a crucial time, as there is an urgent global demand for new therapeutics, particularly antibiotics that can combat drug-resistant bacteria4. Modern research in NPs is rising to meet this challenge by integrating traditional expertise in NPs with state-of-the-art genomic and chemical analysis techniques.

Multi-omics integration is central to the current advancements in biological research. The concept of "multi-omics" involves the simultaneous analysis of various biological data layers, such as an organism's DNA (genome), mRNA transcripts, proteins, and metabolites. The application of omics approaches has transformed NPs research, enabling high-throughput screening, rapid identification of novel compounds, and deeper exploration of the intricate networks that shape biological systems5. By integrating these multi-layered datasets, researchers can directly link genes to the metabolites they encode, thereby constructing a more comprehensive understanding of NPs biosynthesis6. For example, genome mining algorithms can pinpoint a cluster of genes that may encode a new antibiotic, while transcriptomic data can indicate whether those genes are actively expressed. Additionally, metabolomic profiling can identify the corresponding antibiotic molecule in the organism's culture extract7,8,9. This integrated approach significantly speeds up the discovery of novel compounds and the understanding of their biosynthetic pathways. More importantly, drug-target identification is increasingly relying on integrated multi-omics approaches, which are gradually supplanting traditional single-omics strategies10. Recent advancements in analytical chemistry, including high-resolution liquid chromatography-MS (LC-MS) and nuclear magnetic resonance (NMR), allow researchers to detect and structurally analyze new NPs from complex biological samples11,12. These techniques produce vast amounts of data, which is where bioinformatics and machine learning (ML) become essential. Artificial intelligence (AI) algorithms and network-based analyses are increasingly utilized to analyze multi-omics data, uncovering significant patterns and predictions that would be difficult to discern manually13,14. By merging cutting-edge omics technologies with AI-driven data mining, researchers can now establish a robust pipeline for the faster and more systematic discovery and analysis of NPs than ever before.

While multi-omics strategies have significantly advanced the discovery of NPs, their successful implementation relies heavily on meticulous experimental planning. For instance, the type and quality of samples are crucial; it is essential to collect well-preserved microbial cultures, environmental isolates, or clinical materials under conditions that minimize the degradation of nucleic acids and metabolites15. The depth of sequencing also plays a vital role in data resolution: whole-genome sequencing generally requires a coverage of at least 30× for accurate assembly, while transcriptomic profiling often necessitates tens of millions of reads to identify low-abundance transcripts, as recommended by Genohub16. Additionally, for metabolomics, appropriate extraction solvents and methods are needed to capture diverse chemical classes; one should consciously choose extraction protocols that minimize bias and maximize reproducibility17. However, these methodologies come with their own set of challenges. Integrating large, heterogeneous datasets can be computationally intensive, and the instability or low abundance of metabolites can complicate subsequent analyses18. Furthermore, the high costs or limited sample sizes may restrict the design of studies19. Contamination and bias in sequencing data, particularly in low-input or diluted samples, also pose significant risks, as highlighted by Lusk et al. regarding contamination in high-throughput sequencing20. By recognizing these practical and technical limitations, researchers can make informed decisions about suitable multi-omics combinations and interpret their results with the necessary caution.

This review highlights the innovative methods that are revolutionizing the discovery of NPs through the use of multi-omics and integrative analytics. It also explores the growing significance of ML and AI in linking these various data streams, including the construction of gene-metabolite correlation networks and the identification of gene clusters that are likely to produce new bioactive compounds. Additionally, it examines the significance and future potential of this integrated approach. In conclusion, the combination of different omics technologies with advanced analytics is not only speeding up the discovery of new NPs today but is also paving the way for the development of next-generation drugs that are urgently needed by society.

Access restricted. Please log in or start a trial to view this content.

Review and Perspective

Loading...
$$\rightleftharpoonup{xx}$$ $$\longleftharp{xx}$$, $$\longrightharp{xx}$$,

To compile this review, we performed a comprehensive literature search across multiple databases and indexing platforms, including PubMed, Scopus, Web of Science, ScienceDirect, and Google Scholar. The search covered publications from January 2015 through July 2025. Keywords and Boolean combinations included: "natural product discovery," "omics," "metabolomics," "genomics," "transcriptomics," "proteomics," "machine learning," "artificial intelligence," "biosynthetic gene clusters," and "multi-omics integration." Reference lists of key articles and recent reviews were also screened to identify additional relevant publications. In addition to peer-reviewed studies, a limited number of preprint articles from platforms such as bioRxiv and medRxiv were consulted when they provided timely and relevant insights not yet available in published literature. All preprints are clearly labeled to maintain transparency. This section reviews the latest developments and future outlooks in the key areas of omics-based NPs discovery, specifically focusing on metabolomics, genomics, transcriptomics, and proteomics. As well as their integration via AI/ML. The primary tools and resources discussed across these fields are compiled for reference in Table 1.

1. Metabolomics in NPs discovery

  1. Analytical platforms for metabolomics
    Metabolomics is crucial for discovering NPs as it allows for a detailed analysis of the small molecules generated by various organisms. The primary analytical tools used for characterizing the metabolomes of NPs include high-resolution MS techniques, such as LC-MS/MS and gas chromatography-MS (GC-MS), along with NMR spectroscopy21,22. Untargeted LC-MS/MS workflows, like ultra-high-performance LC-MS paired with MS/MS, enable the identification of a wide range of metabolites in complex mixtures, while GC-MS excels in profiling volatile compounds. Additionally, new methods such as direct infusion-tandem MS and MS imaging are enhancing the metabolomics toolkit for NPs23,24. Advanced mass spectrometric techniques, including multimodal MS, which combines positive and negative ion modes or employs multi-stage MSn fragmentation, and hyphenated methods like LC-MS coupled with NMR, offer deeper structural insights into the metabolites that are detected25,26.
  2. Data processing and computational tools
    Specialized software tools play a crucial role in the processing and interpretation of metabolomics data. Programs like XCMS, MZmine, and OpenMS are commonly used to detect and align chromatographic features, or peaks, while meticulous data preprocessing steps such as peak alignment, normalization, and filtering are essential for achieving reproducible metabolic profiles27. The identification of compounds is supported by MS/MS spectral libraries, including NIST and mzCloud, as well as in silico fragmentation tools like SIRIUS and MetFrag28. The Global Natural Products Social (GNPS) platform has emerged as a key resource for molecular networking, which organizes related MS/MS spectra to emphasize chemical families and facilitate the dereplication of known compounds29. For instance, an analysis using GNPS on LC-MS/MS data from Streptomyces cultures successfully identified known antibiotics like spiramycin and actinomycin, alongside previously unreported analogues30,31. The molecular networking technique effectively grouped related unknown spectra, and the addition of ion-identity networking enhanced the connectivity of features. These network-based analyses, when combined with multivariate statistical tools such as principal component analysis and orthogonal partial least squares discriminant analysis, as well as correlation network analysis, can link patterns of metabolite production to specific biological or environmental conditions. In one particular study, GNPS molecular networking of fermentation extracts from a plant-associated Streptomyces demonstrated distinct metabolite profiles that varied depending on the growth mediator used. Known metabolites, including spiramycin, actinomycin, and a carcinogenic compound called lyngbyatoxin, were detected; however, many network nodes did not match any entries in spectral databases32. This finding suggests the presence of genuinely novel metabolites, highlighting how untargeted metabolomics, in conjunction with GNPS, can identify potentially new NPs for further exploration.
    Combining metabolomic screening with functional bioassays facilitates the swift prioritization of strains or extracts that yield bioactive compounds. Additionally, emerging community resources like the Natural Products Atlas (NPAtlas), along with integrated metabolome-genome analysis pipelines, are instrumental in associating detected metabolites with their BGCs in the source organisms33,34. This integrative strategy enables researchers to correlate metabolite signals with genomic data, thus enhancing the efficiency of identifying NPs and tracing their origins.

2. Genomics in NPs discovery

Over the past decade, microbial NPs research has entered a transformative "deep-mining era" fueled by advances such as high-throughput sequencing, ultra-sensitive high-resolution MS, and integrated multi-omics approaches35. These tools have shifted discovery efforts from serendipitous isolation toward data-driven, targeted exploration. Genome mining pipelines, including antiSMASH 7.0 and DeepBGC, now enable systematic identification of cryptic BGCs, particularly in under-explored taxa36,37.

  1. High-throughput sequencing and hybrid assembly
    High-throughput DNA sequencing technologies, such as short-read high-throughput DNA sequencing combined with long-read sequencing platforms, have significantly transformed the genomic study of organisms that produce NPs. By employing hybrid assembly strategies that integrate both long and short reads, researchers can achieve high-contiguity genome assemblies, which are essential for capturing entire BGCs within a genome38. Additionally, metagenomic sequencing, which includes the recovery of metagenome-assembled genomes, broadens the discovery of BGCs to uncultured microbes, enabling scientists to explore environmental DNA for its biosynthetic potential39. In instances where conventional genome assembly methods fall short in resolving large or repetitive gene clusters, library-based techniques, such as cosmid or bacterial artificial chromosome libraries, can be utilized to isolate and clone substantial genomic fragments. This approach facilitates the characterization of complete gene clusters through targeted sequencing or heterologous expression, thereby enhancing our understanding of the genetic basis for NPs biosynthesis40,41.
  2. Genome mining tools
    Specialized bioinformatic tools are essential for analyzing genomic data to identify the distinct signatures of secondary metabolite pathways. Programs like antiSMASH, PRISM, and DeepBGC play a crucial role in this process by automatically detecting candidate BGCs. They do this by recognizing specific biosynthetic domains, such as those associated with nonribosomal peptide synthetases, polyketide synthases, ribosomally synthesized and post-translationally modified peptide (RiPP) pathways, and terpene cyclases, among others. To aid in dereplication, databases like MIBiG, which compiles known BGCs and their products, and NPAtlas, which catalogs known NPs, help researchers match new sequences or compounds to established examples37,42,43. Clustering algorithms, such as BiG-SCAPE, along with tools like CORASON, organize related BGCs into families, offering a network view of biosynthetic diversity and facilitating the identification of novel cluster families through comparisons with known ones44. Additional bioinformatic analyses, including BLAST searches and hidden Markov model scans, can predict the potential functions of enzymes encoded within a BGC, while comparative genomics approaches, such as synteny analysis and the co-evolution of genes, help refine these functional predictions45. Furthermore, genomic data enable the annotation of regulatory elements linked to NPs biosynthesis, including pathway-specific promoters and transcriptional regulators46. These insights provide a foundation for pathway engineering; for example, synthetic biology techniques like transformation-associated recombination cloning and Red/ET recombineering empower researchers to capture and express silent or cryptic BGCs in a heterologous host organism, thereby activating the production of their metabolites47.
    Genome mining efforts have successfully led to the discovery of new NPs by concentrating on specific genetic signals. For instance, a systematic search for genes linked to glycopeptide antibiotic resistance revealed several unusual BGCs predicted to encode new antibiotics. This method resulted in the identification of two novel compounds, compstatin and cobramycin, both of which exhibit unique mechanisms of action by targeting bacterial autolysins48. In a similar vein, exploring the human microbiome for biosynthetic gene content yielded a lipopeptide antibiotic candidate known as "humicin," recognized for its potential effectiveness against Staphylococcus and Streptococcus species49. In these instances, the compounds were initially identified through computational methods, with their chemical structures predicted based on genomic data. Subsequently, these predicted molecules were chemically synthesized and demonstrated the expected antibiotic activities, thereby validating the genome-based discovery strategy. On a broader scale, large-scale sequencing of various bacteria, including many rare actinomycetes, is uncovering a wealth of "cryptic" BGCs-gene clusters whose products remain unknown. For example, investigations into Streptomyces genomes have revealed thousands of uncharacterized BGCs50. An emerging approach to linking these orphan clusters to actual metabolites involves integrating metabolomic analyses with genomic surveys. By correlating the presence or expression of a gene cluster with the detection of unique metabolite features through platforms like GNPS or mass spectral databases, researchers can start to assign tentative chemical products to the numerous new BGCs, thus facilitating the discovery of truly novel NPs from genomic data.

3. Transcriptomics in natural product discovery

Transcriptomic analyses offer a dynamic perspective on the activity of BGCs by measuring mRNA expression under various conditions. Specifically, RNA-seq experiments can uncover which BGCs are actively transcribed and how their expression levels fluctuate in response to environmental or experimental changes. By comparing transcript levels across different conditions, differential expression analysis, utilizing tools such as DESeq2 or edgeR, can pinpoint BGCs that are either up-regulated or down-regulated, thereby highlighting gene clusters that may be inducible or remain silent under standard conditions51. Although conventional RNA-seq conducted on bulk cultured cells is the typical approach for investigating BGC expression, more advanced techniques like single-cell RNA-seq are starting to be utilized in complex microbiomes. This shift allows for the observation of gene expression in environmental organisms that are difficult to culture52,53. A standard RNA-seq workflow generally includes mapping reads to a reference genome using aligners such as STAR or HISAT2, quantifying gene expression with tools like featureCounts or Kallisto, and normalizing the data to identify significant expression changes. Following this, co-expression network analysis, often performed with the WGCNA package, can cluster genes exhibiting similar expression patterns. When metabolite data is also available, these networks can be further expanded to incorporate metabolites, thus linking specific compounds with co-expressed genes54.

Transcriptomic data have proven to be valuable in identifying conditionally active biosynthetic pathways. A notable example of this approach is found in a detailed comparison of four Salinispora strains51. In this study, it was revealed that more than half of the BGCs in each strain were expressed to some degree, with many clusters that were inactive in one strain showing expression in another. By analyzing gene expression profiles across these strains, the researchers were able to pinpoint several regulatory factors that seemed to either silence or activate specific clusters. This integrative method led to the identification of a previously unknown family of NPs, known as salinipostins. Transcriptomic data indicated a cryptic gene cluster as a potential source for an unidentified compound; when this cluster was induced to express (through regulatory changes), it resulted in the production of salinipostin molecules, which were subsequently isolated and structurally characterized. This study exemplifies how transcriptomics can facilitate the discovery of NPs: by analyzing differential gene expression, researchers can highlight candidate BGCs, and the correlation between gene expression and metabolite profiles provides supporting evidence for gene-metabolite relationships that can be validated through targeted experiments.

4. Proteomics in NPs discovery

  1. Shotgun and targeted proteomics approaches
    Proteomics, which involves the large-scale study of proteins, provides a vital perspective in NPs research by directly measuring the enzymes and proteins that play roles in biosynthetic pathways. Generally, proteomic approaches can be categorized into two main types: shotgun (untargeted) proteomics and targeted proteomics. In shotgun proteomics, proteins extracted from microbial or plant samples undergo enzymatic digestion, typically with trypsin, and the resulting peptides are analyzed using high-resolution LC-MS/MS, often employing advanced mass spectrometers such as quadrupole time-of-flight (QTOF) or high-resolution orbital trap mass spectrometers55. The MS data collected, known as MS/MS spectra, are then matched to peptide sequences by searching against a protein database derived from the organism's genome or a closely related species, utilizing software tools such as MaxQuant or Mascot56,57. This process allows for the quantification of changes in protein abundance under different conditions, which can be achieved either through label-free methods or by employing isotope labeling techniques, such as stable isotope labeling by amino acids in cell culture (SILAC) or isobaric tagging with TMT/iTRAQ reagents, to determine which biosynthetic enzymes are up-regulated or down-regulated58,59. In contrast, targeted proteomics methods, including selected reaction monitoring (SRM) or parallel reaction monitoring (PRM), concentrate on a specific set of peptide ions, often unique peptides from key biosynthetic enzymes or regulatory proteins, enabling precise and sensitive quantification of these peptides across multiple samples60,61.
  2. Innovative approaches in proteomics
    One significant challenge in the proteomic analysis of secondary metabolism is the detection of large biosynthetic enzymes, such as nonribosomal peptide synthetases and polyketide synthases, which often exceed 100 kDa in size and yield few detectable peptides with standard digestion protocols, making them difficult to observe. To tackle this issue, Bumpus et al. developed a method known as PrISM, which enhances the detection of NRPS and PKS proteins by utilizing a specific cofactor found on these enzymes62. PrISM integrates conventional shotgun proteomics with a phosphopantetheinyl (Ppant) ejection assay. Notably, NRPS and PKS enzymes possess a covalently attached Ppant arm on their acyl carrier or peptidyl carrier domains; during MS/MS fragmentation, this prosthetic group frequently generates a unique fragment ion, referred to as the "Ppant ejection" ion. By computationally filtering the LC-MS/MS data for this diagnostic ion, the PrISM method can effectively highlight peptides derived from NRPS and PKS enzymes amidst the complex mixture of all cellular proteins. This strategy has successfully uncovered hidden biosynthetic pathways through proteomics. For instance, the PrISM workflow enabled the detection of peptides from the gramicidin S synthetase complex (GrsA/GrsB) in cultures of Bacillus brevis, establishing a direct link between these NRPS enzymes and the production of the antibiotic gramicidin S in that organism62. Importantly, the proteomic identification of GrsA/GrsB did not necessitate prior genomic knowledge of B. brevis, illustrating that in certain instances, proteomic analysis alone can uncover active NPs biosynthetic pathways even in organisms whose genomes have not yet been sequenced.
  3. Genome-informed proteomics
    When a genome sequence is available, the power of proteomics can be significantly enhanced by utilizing a strain-specific database for peptide identification. For example, in a study analyzing the proteome of Streptomyces lilacinus, researchers conducted two analyses: one by searching spectra against a general protein database and another against a custom database derived from the sequenced genome of that strain63. The genome-informed analysis yielded approximately ten times more identified peptides and detected peptides linked to six distinct BGCs within the organism. Notably, three of these identified clusters corresponded to previously known "orphan" metabolites, including graybactin, rakicidin D, and a specific siderophore, with their biosynthetic enzymes confirmed to be expressed. The remaining detected clusters appeared to be novel and uncharacterized. This example demonstrates that integrating genomic information into proteomic workflows can confirm which biosynthetic pathways are actively expressed in a strain, thus prioritizing those gene clusters for the isolation and characterization of their products.
    Additionally, proteomic methods can illuminate the molecular targets and modes of action of NPs, an area often referred to as chemical proteomics. For instance, activity-based protein profiling (ABPP) employs small molecule probes, typically derivatives or analogs of NPs, that react with the cellular targets of the compound, tagging those proteins for enrichment and identification through MS64. This technique can uncover the protein targets that a specific NP or related molecule binds to within the cell. Another method, thermal proteome profiling (TPP), involves heating protein samples in the presence or absence of a compound to determine which proteins show altered thermal stability, indicating a binding interaction65. These chemical proteomic approaches are invaluable for validating the biological targets of NPs and elucidating their mechanisms of action, thereby complementing the discovery of the NPs themselves.

5. ML and AI for omics integration and prediction

  1. Integration of multi-omics data with ML
    An exciting frontier in the discovery of NPs is the application of ML and AI to integrate omics data for functional predictions. In the realm of genomics, supervised ML algorithms, including random forests, support vector machines, and deep neural networks, have been developed to identify patterns in BGCs that are associated with specific types of NPs. For instance, these ML models can categorize BGCs based on the general class of molecules they produce or even forecast certain characteristics of the chemical structure derived from the gene sequence. A review by Dason MS and colleagues highlights various ML tools designed to decode the "language" of BGCs, utilizing DNA or amino acid sequence data to deduce the structures and bioactivities of the corresponding NPs66. These models typically rely on extensive databases of known gene clusters and compounds to learn the relationships between them, enabling predictions about the potential biological activity of new compounds. For example, they can assess whether a product from an uncharacterized BGC might possess antibiotic or anticancer properties. Significant advancements that have facilitated these capabilities include the utilization of large training datasets, such as thousands of known NP structures and gene clusters, innovative representation techniques like sequence embeddings and graph neural networks that capture the features of gene clusters and chemical structures, and multi-parametric models that combine various genomic and chemical descriptors to enhance prediction accuracy.
    In the field of metabolomics, AI techniques are significantly enhancing the interpretation of complex mass spectral data. A notable example is the tool CSI:FingerID, which employs ML to predict a molecule's structural fingerprint from its MS/MS spectrum. This capability aids in identifying unknown metabolites by suggesting potential substructures and candidate compounds67. Similarly, deep learning models have been utilized for spectral annotation; convolutional neural networks and, more recently, transformer-based architectures, such as the "DreaMS" model, learn fragmentation patterns from extensive spectral libraries. These models can propose structures for unknown spectra based on the patterns they have learned, often outperforming traditional heuristic methods, particularly for metabolites that do not have exact matches in existing databases68. Beyond structural annotation, ML also contributes to metabolomics by identifying global patterns. For example, unsupervised visualization algorithms such as t-SNE (t-distributed stochastic neighbor embedding) and UMAP (Uniform Manifold Approximation and Projection) can reduce the dimensionality of untargeted metabolomics datasets, facilitating the clustering of samples based on similarities in their metabolite profiles. Meanwhile, supervised classifiers can connect specific metabolite signatures to the phenotypic traits or bioactivities of the samples, further enriching the analysis69,70,71.
    ML is increasingly being utilized to integrate various omics datasets, creating a more comprehensive understanding of NPs biosynthesis and function. By merging data from genomics, transcriptomics, proteomics, and metabolomics, integrative models can reveal connections that may be overlooked when examining each data type separately. For instance, researchers can construct gene-metabolite correlation networks that link the expression levels of genes or proteins to the production levels of metabolites. ML models trained on this integrated data can then predict which gene clusters are likely responsible for specific metabolites or biological activities72,73. Advanced multivariate techniques, such as multi-omics factor analysis (available in tools like MOFA) and regularized multivariate methods (found in packages like mixOmics), facilitate the identification of latent factors that encompass different data types, such as a set of co-expressed genes alongside a group of co-occurring metabolites74,75,76. In practical terms, for NPs discovery, various approaches assist in prioritizing candidate gene clusters by demonstrating a connection to observed metabolites or phenotypes. A notable example of this integrative, AI-guided strategy is the discovery of a new family of ribosomally synthesized and RiPPs using the DecRiPPter algorithm. This algorithm utilizes a support vector machine (SVM)-based model to analyze genomic data for cryptic RiPP precursor peptides and then cross-references these predictions with pan-genome analysis to eliminate known compounds. When applied to 1,295 Streptomyces genomes, DecRiPPter successfully identified 42 putative novel RiPP families that had previously been missed by traditional genome mining methods. Among these predicted clusters, one was experimentally validated to produce a new class V lanthipeptide, which the researchers named "pristinin A3." By inducing the expression of this silent gene cluster, the compound pristinin A3 was isolated, and its structure was determined using NMR and MS, revealing two previously unknown lanthionine-forming enzymes encoded within the pathway77. This achievement highlights how AI-driven analyses can uncover hidden NPs: the ML model detected an otherwise unrecognized genetic signal, guiding experimental efforts to discover a novel molecule.
  2. Emerging AI applications
    In addition to the current applications, several emerging AI-driven strategies are set to further revolutionize NPs research. One promising area is computational retrosynthesis and pathway design, where deep learning models are being developed to suggest biosynthetic routes or synthetic chemistry steps for known NPs and their analogs, which could significantly enhance the design and production of novel compounds78,79. Another exciting frontier is the prediction of enzyme function using AI. By leveraging modern protein structure prediction tools based on deep learning algorithms, alongside ML techniques like contrastive learning on protein sequences, researchers are starting to deduce the likely activities of uncharacterized enzymes found in BGCs. This could provide insights into the types of chemical transformations these enzymes catalyze80,81,82. Fully integrated omics-to-chemistry pipelines are being developed, where AI platforms aim to automatically connect mass spectral features to genomic data83,84. In principle, such a system could analyze an LC-MS/MS dataset from a complex sample alongside genome sequences from the same environment, allowing it to predict which gene cluster is responsible for each unidentified metabolite by scoring gene cluster-metabolite matches. Although these approaches are still in their infancy, they suggest a future where AI could autonomously link genes to molecules, greatly speeding up the process from omics data to the discovery of new NPs.
  3. Data governance and ethical challenges in AI-assisted NPs discovery
    Modern multi-omics research produces a significant amount of diverse data, but the effectiveness of AI predictions relies heavily on the quality and consistency of these datasets. When training sets are poorly annotated or biased, they can result in misleading outcomes and unsuccessful validations85. To address this issue, researchers should implement the FAIR principles-Findable, Accessible, Interoperable, and Reusable-to promote data transparency, reproducibility, and extensive reusability86. This approach involves sharing both positive and negative results, thoroughly documenting data preprocessing pipelines, and providing analysis code to facilitate independent verification. Additionally, the adoption of emerging digital infrastructures, such as electronic laboratory notebooks (ELNs) and containerized workflows, is essential for improving data capture and ensuring standardized curation87,88.
    While AI significantly speeds up the discovery of NPs, it also brings forth crucial ethical concerns. Algorithms that are trained on incomplete or biased datasets can reinforce existing inequalities, underscoring the necessity for diverse and representative training data89. Additionally, employing explainable AI methods is vital for enhancing transparency and assisting scientists in understanding predictions, ensuring that AI serves as a supportive tool under human control rather than a decision-maker85. Ethical considerations also encompass data privacy, especially when using clinical or human-derived datasets, as well as the implications for the workforce as automation increasingly influences the research environment90. To tackle these issues, AI systems should be subject to regular audits, bias assessments, and stringent regulatory oversight, which will help maintain fairness, accountability, and trustworthiness in the field of NPs research.

Access restricted. Please log in or start a trial to view this content.

Conclusions

Loading...
$$\rightleftharpoonup{xx}$$ $$\longleftharp{xx}$$, $$\longrightharp{xx}$$,

The current state of NPs discovery leverages a combination of analytical chemistry and bioinformatics, creating a powerful synergy. As outlined in Figure 1, advanced omics technologies such as LC-MS metabolomics, high-throughput sequencing, RNA-Seq, and proteomics work together with computational analytics, including networking and ML, to form an effective pipeline for discovery. Recent developments in the field include single-cell omics, which allow for the ...

Access restricted. Please log in or start a trial to view this content.

Disclosures

Loading...
$$\rightleftharpoonup{xx}$$ $$\longleftharp{xx}$$, $$\longrightharp{xx}$$,

All authors declare that there are no financial or personal relationships that could be perceived as potential conflicts of interest.

Acknowledgements

Loading...
$$\rightleftharpoonup{xx}$$ $$\longleftharp{xx}$$, $$\longrightharp{xx}$$,

This work was supported by the National Natural Science Foundation of China (32500813), Postdoctoral Fellowship Program of CPSF (GZC20251503), Sichuan Science and Technology Program (2024ZYD0149), Special Research Foundation for the Postdoctoral Program of Sichuan Province (TB2024017), Key Research and Development Project of Deyang Science and Technology Bureau (2024SZY016, 2023SZZ009), National Health Commission Capacity Building and Continuing Education Center (GWJJMB202510025060), and Special Fund for Incubation Projects of Deyang People's Hospital (FRH202501, FHT202501). We acknowledge that Figure 1 was created using BioRender,

Access restricted. Please log in or start a trial to view this content.

References

Loading...
$$\rightleftharpoonup{xx}$$ $$\longleftharp{xx}$$, $$\longrightharp{xx}$$,
  1. Newman, D. J., Cragg, G. M. Natural products as sources of new drugs over the nearly four decades from 01/1981 to 09/2019. J Nat Prod. 83 (3), 770-803 (2020).
  2. Shang, X., et al. Natural products in antiparasitic drug discovery: Advances, opportunities and challenges. Nat Prod Rep. 42 (9), 1419-1458 (2025).
  3. Atanasov, A. G., et al. Natural products in drug discovery: Advances and opportunities. Nat Rev Drug Discov. 20 (3), 200-216 (2021).
  4. Xie, S., Zhan, F., Zhu, J., Xu, S., Xu, J. The latest advances with natural products in drug discovery and opportunities for the future: A 2025 update. Expert Opin Drug Discov. 20 (7), 827-843 (2025).
  5. Sahana, S., et al. Multi-omics approaches: Transforming the landscape of natural product isolation. Funct Integr Genomics. 25 (1), 132(2025).
  6. Zhang, H. W., et al. Application of omics- and multi-omics-based techniques for natural product target discovery. Biomed Pharmacother. 141, 111833(2021).
  7. Blin, K., et al. antiSMASH 5.0: Updates to the secondary metabolite genome mining pipeline. Nucleic Acids Res. 47 (W1), W81-W87 (2019).
  8. Song, C., et al. Comparative transcriptomics unveil the crucial genes involved in coumarin biosynthesis in Peucedanum praeruptorum Dunn. Front Plant Sci. 13, 899819(2022).
  9. Bayona, L. M., De Voogd, N. J., Choi, Y. H. Metabolomics on the study of marine organisms. Metabolomics. 18 (3), 17(2022).
  10. Du, P., Fan, R., Zhang, N., Wu, C., Zhang, Y. Advances in integrated multi-omics analysis for drug-target identification. Biomolecules. 14 (6), 692(2024).
  11. Elgahamy, A. A., El-Desoky, A. H., Otify, A. M., El Fishawy, A. M., El-Beih, A. A. Molecular networking derived from untargeted LC-MS/MS analysis to discover inhibitors of RANKL-induced osteoclastogenesis from Egyptian marine sponge-associated fungi. Sci Rep. 15 (1), 27137(2025).
  12. Bustamam, M. S. A., et al. Complementary analytical platforms of NMR spectroscopy and LCMS analysis in the metabolite profiling of Isochrysis galbana. Mar Drugs. 19 (3), 139(2021).
  13. Hu, G., Qiu, M. Machine learning-assisted structure annotation of natural products based on MS and NMR data. Nat Prod Rep. 40 (11), 1735-1753 (2023).
  14. Gaudencio, S. P., et al. Advanced methods for natural products discovery: Bioactivity screening, dereplication, metabolomics profiling, genomic sequencing, databases and informatics tools, and structure elucidation. Mar Drugs. 21 (5), 308(2023).
  15. Sedighikamal, H., Mashayekhan, S. Critical assessment of quenching and extraction/sample preparation methods for microorganisms in metabolomics. Metabolomics. 21 (2), 40(2025).
  16. Recommended Coverage and Read Depth for NGS Applications. , https://genohub.com/recommended-sequencing-coverage-by-application/?utm_source (2025).
  17. Brockbals, L., Ueland, M., Fu, S., Padula, M. P. Development and thorough evaluation of a multi-omics sample preparation workflow for comprehensive LC-MS/MS-based metabolomics, lipidomics and proteomics datasets. Talanta. 286, 127442(2025).
  18. Graw, S., et al. Multi-omics data integration considerations and study design for biological systems and disease. Mol Omics. 17 (2), 170-185 (2021).
  19. Han, E., Kwon, H., Jung, I. A review on multi-omics integration for aiding study design of large-scale TCGA cancer datasets. BMC Genomics. 26 (1), 769(2025).
  20. Lusk, R. W. Diverse and widespread contamination evident in the unmapped depths of high-throughput sequencing data. PloS one. 9 (10), e110808(2014).
  21. Queiroz, E. F., Guillarme, D., Wolfender, J. L. Advanced high-resolution chromatographic strategies for efficient isolation of natural products from complex biological matrices: From metabolite profiling to pure chemical entities. Phytochem Rev. 23 (5), 1415-1442 (2024).
  22. Huang, Z., Bi, T., Jiang, H., Liu, H. Review on NMR as a tool to analyse natural products extract directly: Molecular structure elucidation and biological activity analysis. Phytochem Anal. 35 (1), 5-16 (2024).
  23. Li, T., et al. Direct infusion-tandem mass spectrometry combining with data mining strategies enables rapid chemome characterization of medicinal plants: A case study of Polygala tenuifolia. J Pharm Biomed Anal. 204, 114281(2021).
  24. Zou, Y., Tang, W., Li, B. Exploring natural product biosynthesis in plants with mass spectrometry imaging. Trends Plant Sci. 30 (1), 69-84 (2025).
  25. Zhang, A., et al. Qualitative and quantitative determination of chemical constituents in Jinbei oral liquid, a modern Chinese medicine for coronavirus disease 2019, by ultra-performance liquid chromatography coupled with mass spectrometry. Front Chem. 11, 1079288(2023).
  26. Liu, Y., et al. Discovery of bioactive-chemical Q-markers of Acanthopanax sessiliflorus leaves: An integrated strategy of plant metabolomics, fingerprint and spectrum-efficacy relationship research. J Chromatogr B Analyt Technol Biomed Life Sci. 1233, 124009(2024).
  27. LC-MS analysis: Mini review frequently used open source softwares. Binanto, I., Warnars, H. L. H. S., Sianipar, N. F., Abbas, B. S. 2019 6th International Conference on Information Technology, Computer and Electrical Engineering (ICITACEE), , IEEE. 1-5 (2019).
  28. Blazenovic, I., Kind, T., Ji, J., Fiehn, O. Software tools and approaches for compound identification of LC-MS/MS data in metabolomics. Metabolites. 8 (2), 31(2018).
  29. Wang, M., et al. Sharing and community curation of mass spectrometry data with Global Natural Products Social Molecular Networking. Nat Biotechnol. 34 (8), 828-837 (2016).
  30. Leclair, M. M., et al. Discovery of Levesquamide B through Global Natural Products Social Molecular Networking. Molecules. 27 (22), 7794(2022).
  31. Yanagisawa, K., et al. A new pyranonaphthoquinone, actinoquinonal A, and its congeners from the combined-culture of Streptomyces sp. 23-50 and Tsukamurella pulmonis TP-B0596. J Antibiot (Tokyo). 78 (6), 350-358 (2025).
  32. Kum, E., Ince, E. Metabolomics approach to explore bioactive natural products derived from plant-root-associated streptomyces. Appl Biochem Biotechnol. 196 (10), 7293-7306 (2024).
  33. Poynton, E. F., et al. The Natural Products Atlas 3.0: Extending the database of microbially derived natural products. Nucleic Acids Res. 53 (D1), D691-D699 (2025).
  34. Zhang, S., et al. Global analysis of natural products biosynthetic diversity encoded in fungal genomes. J Fungi (Basel). 10 (9), 653(2024).
  35. Wang, Z., et al. The deep mining era: Genomic, metabolomic, and integrative approaches to microbial natural products from 2018 to 2024. Mar Drugs. 23 (7), 261(2025).
  36. Blin, K., et al. antiSMASH 7.0: New and improved predictions for detection, regulation, chemical structures and visualisation. Nucleic Acids Res. 51 (W1), W46-W50 (2023).
  37. Hannigan, G. D., et al. A deep learning genome-mining strategy for biosynthetic gene cluster prediction. Nucleic Acids Res. 47 (18), e110(2019).
  38. Sanchez-Navarro, R., et al. Long-read metagenome-assembled genomes improve identification of novel complete biosynthetic gene clusters in a complex microbial activated sludge ecosystem. mSystems. 7 (6), e0063222(2022).
  39. Waschulin, V., et al. Biosynthetic potential of uncultured Antarctic soil bacteria revealed through long-read metagenomic sequencing. ISME J. 16 (1), 101-111 (2022).
  40. Xu, M., et al. Functional genome mining for metabolites encoded by large gene clusters through heterologous expression of a whole-genome bacterial artificial chromosome library in Streptomyces spp. Appl Environ Microbiol. 82 (19), 5795-5805 (2016).
  41. Liu, Z., Zhao, Y., Huang, C., Luo, Y. Recent advances in silent gene cluster activation in Streptomyces. Front Bioeng Biotechnol. 9, 632230(2021).
  42. Blin, K., et al. antiSMASH 6.0: Improving cluster detection and comparison capabilities. Nucleic Acids Res. 49 (W1), W29-W35 (2021).
  43. Skinnider, M. A., et al. Comprehensive prediction of secondary metabolite structure and biological activity from microbial genome sequences. Nat Commun. 11 (1), 6058(2020).
  44. Navarro-Munoz, J. C., et al. A computational framework to explore large-scale biosynthetic diversity. Nat Chem Biol. 16 (1), 60-68 (2020).
  45. Zhu, S., et al. Computational advances in biosynthetic gene cluster discovery and prediction. Biotechnol Adv. 79, 108532(2025).
  46. Zhao, C., et al. Genome sequencing provides potential strategies for drug discovery and synthesis. Acupunc Herbal Med. 3 (4), 244-255 (2023).
  47. Bomfiglio, I. F., Mendes, I. S. M., Bonatto, D. A review of DNA restriction-free overlapping sequence cloning techniques for synthetic biology. Biotechnol J. 20 (7), e70084(2025).
  48. Culp, E. J., et al. Evolution-guided discovery of antibiotics that inhibit peptidoglycan remodelling. Nature. 578 (7796), 582-587 (2020).
  49. Chu, J., et al. Discovery of MRSA active antibiotics using primary sequence from the human microbiome. Nat Chem Biol. 12 (12), 1004-1006 (2016).
  50. Kloosterman, A. M., et al. Expansion of RiPP biosynthetic space through integration of pan-genomics and machine learning uncovers a novel class of lanthipeptides. PLoS Biol. 18 (12), e3001026(2020).
  51. Amos, G. C. A., et al. Comparative transcriptomics as a guide to natural product discovery and biosynthetic gene cluster functionality. Proc Natl Acad Sci U S A. 114 (52), E11121-E11130 (2017).
  52. Imdahl, F., Saliba, A. E. Advances and challenges in single-cell RNA-seq of microbial communities. Curr Opin Microbiol. 57, 102-110 (2020).
  53. Llorens-Rico, V., Simcock, J. A., Huys, G. R. B., Raes, J. Single-cell approaches in human microbiome research. Cell. 185 (15), 2725-2738 (2022).
  54. Hirsch, P., et al. ABC-HuMi: The Atlas of biosynthetic gene clusters in the human microbiome. Nucleic Acid Res. 52 (D1), D579-D585 (2024).
  55. Tyanova, S., Temu, T., Cox, J. The MaxQuant computational platform for mass spectrometry-based shotgun proteomics. Nat Protoc. 11 (12), 2301-2319 (2016).
  56. Lepretre, M., et al. From shotgun to targeted proteomics: rapid Scout- MRM assay development for monitoring potential immunomarkers in Dreissena polymorpha. Anal Bioanal Chem. 412 (26), 7333-7347 (2020).
  57. Rani, P., Alam, S. I., Singh, S., Kumar, S. Elucidation of peptide screen for targeted identification of Yersinia pestis by nano-liquid chromatography tandem mass spectrometry. Sci Rep. 15 (1), 1096(2025).
  58. Beller, N. C., Hummon, A. B. Advances in stable isotope labeling: Dynamic labeling for spatial and temporal proteomic analysis. Mol Omics. 18 (7), 579-590 (2022).
  59. Meng, J. Y., et al. Identification of differentially expressed proteins in sugarcane in response to infection by Xanthomonas albilineans using iTRAQ quantitative proteomics. Microorganisms. 8 (1), 76(2020).
  60. Vidova, V., Spacil, Z. A review on mass spectrometry-based quantitative proteomics: Targeted and data independent acquisition. Anal Chim Acta. 964, 7-23 (2017).
  61. Rauniyar, N. Parallel reaction monitoring: A targeted experiment performed using high resolution and high mass accuracy mass spectrometry. Int J Mol Sci. 16 (12), 28566-28581 (2015).
  62. Bumpus, S. B., Evans, B. S., Thomas, P. M., Ntai, I., Kelleher, N. L. A proteomics approach to discovering natural products and their biosynthetic pathways. Nat Biotechnol. 27 (10), 951-956 (2009).
  63. Albright, J. C., Goering, A. W., Doroghazi, J. R., Metcalf, W. W., Kelleher, N. L. Strain-specific proteogenomics accelerates the discovery of natural products via their biosynthetic pathways. J Ind Microbiol Biotechnol. 41 (2), 451-459 (2014).
  64. Chen, X., et al. Target identification with quantitative activity-based protein profiling (ABPP). Proteomics. 17 (3-4), (2017).
  65. Cui, Z., Li, C., Chen, P., Yang, H. An update of label-free protein target identification methods for natural active products. Theranostics. 12 (4), 1829-1854 (2022).
  66. Dason, M. S., Cora, D., Re, A. Sequence modeling tools to decode the biosynthetic diversity of the human microbiome. mSystems. 10 (7), e0033325(2025).
  67. Ludwig, M., Duhrkop, K., Bocker, S. Bayesian networks for mass spectrometric metabolite identification via molecular fingerprints. Bioinformatics. 34 (13), i333-i340 (2018).
  68. Bushuiev, R., et al. Self-supervised learning of molecular representations from millions of tandem mass spectra using DreaMS. Nat Biotechnol. , (2025).
  69. Tian, M., et al. Pure ion chromatograms combined with advanced machine learning methods improve accuracy of discriminant models in LC-MS-based untargeted metabolomics. Molecules. 26 (9), 2715(2021).
  70. Mildau, K., et al. Effective data visualization strategies in untargeted metabolomics. Nat Prod Rep. 42 (6), 982-1019 (2025).
  71. Yuan, X., Smith, N. S., Moghe, G. D. Analysis of plant metabolomics data using identification-free approaches. Applications in Plant Sciences. 13 (4), e70001(2025).
  72. Ji, J., Jung, S. PredCMB: Predicting changes in microbial metabolites based on the gene-metabolite network analysis of shotgun metagenome data. Bioinformatics. 41 (1), btaf020(2024).
  73. Wang, X., Kadarmideen, H. N. Metabolite genome-wide association study (m GWAS) and gene-metabolite interaction network analysis reveal potential biomarkers for feed efficiency in pigs. Metabolites. 10 (5), 201(2020).
  74. Argelaguet, R., et al. Multi-Omics Factor Analysis-a framework for unsupervised integration of multi-omics data sets. Mol Syst Biol. 14 (6), e8124(2018).
  75. Rohart, F., Gautier, B., Singh, A., Le Cao, K. A. mixOmics: An R package for 'omics feature selection and multiple data integration. PLoS Comput Biol. 13 (11), e1005752(2017).
  76. Arikan, M., Muth, T. Integrated multi-omics analyses of microbial communities: A review of the current state and future directions. Mol Omics. 19 (8), 607-623 (2023).
  77. Kloosterman, A. M., et al. Integration of machine learning and pan-genomics expands the biosynthetic landscape of RiPP natural products. bioRxiv. , (2020).
  78. Sathyanarayana, S. V., et al. DeepRetro: Retrosynthetic pathway discovery using iterative LLM reasoning. arXiv e-prints. , (2025).
  79. Zheng, S., et al. Deep learning driven biosynthetic pathways navigation for natural products with BioNavi-NP. Nat Commun. 13 (1), 3342(2022).
  80. Wang, W., Shuai, Y., Zeng, M., Fan, W., Li, M. DPFunc: Accurately predicting protein function via deep learning with domain-guided structure information. Nat Commun. 16 (1), 70(2025).
  81. Yu, T., et al. Enzyme function prediction using contrastive learning. Science. 379 (6639), 1358-1363 (2023).
  82. Casadevall, G., Duran, C., Osuna, S. AlphaFold2 and deep learning for elucidating enzyme conformational flexibility and its application for design. JACS Au. 3 (6), 1554-1562 (2023).
  83. Lee, Y. Y., et al. HypoRiPPAtlas as an Atlas of hypothetical natural products for mass spectrometry database search. Nat Commun. 14 (1), 4219(2023).
  84. Leao, T. F., et al. NPOmix: A machine learning classifier to connect mass spectrometry fragmentation data to biosynthetic gene clusters. PNAS Nexus. 1 (5), pgac257(2022).
  85. Hermann, E., Hermann, G., Tremblay, J. C. Ethical artificial intelligence in chemical research and development: A dual advantage for sustainability. Sci Eng Ethics. 27 (4), (2021).
  86. Mugahid, D., et al. A practical guide to FAIR data management in the age of multi-OMICS and AI. Front Immunol. 15, 1439434(2024).
  87. Lin, C. L., et al. Addressing standardization and semantics in an electronic lab notebook for multidisciplinary use: LabIMotion. J Cheminform. 17 (1), 75(2025).
  88. Alser, M., et al. Packaging and containerization of computational methods. Nat Protoc. 19 (9), 2529-2539 (2024).
  89. Blanco-Gonzalez, A., et al. The role of AI in drug discovery: Challenges, opportunities, and strategies. Pharmaceuticals (Basel). 16 (6), 891(2023).
  90. Mirakhori, F., Niazi, S. K. Harnessing the AI/ML in drug and biological products discovery and development: The regulatory perspective. Pharmaceuticals (Basel). 18 (1), 47(2025).

Access restricted. Please log in or start a trial to view this content.

Reprints and Permissions

Request permission to reuse the text or figures of this JoVE article

Request Permission

Tags

Multi Omics AnalyticsMetabolomics PlatformsGenomics ApproachesProteomics TechniquesBioinformatics ToolsMolecular NetworkingGenome MiningArtificial Intelligence

Related Articles