$$\rightleftharpoonup{xx}$$
$$\longleftharp{xx}$$,
$$\longrightharp{xx}$$,
To compile this review, we performed a comprehensive literature search across multiple databases and indexing platforms, including PubMed, Scopus, Web of Science, ScienceDirect, and Google Scholar. The search covered publications from January 2015 through July 2025. Keywords and Boolean combinations included: "natural product discovery," "omics," "metabolomics," "genomics," "transcriptomics," "proteomics," "machine learning," "artificial intelligence," "biosynthetic gene clusters," and "multi-omics integration." Reference lists of key articles and recent reviews were also screened to identify additional relevant publications. In addition to peer-reviewed studies, a limited number of preprint articles from platforms such as bioRxiv and medRxiv were consulted when they provided timely and relevant insights not yet available in published literature. All preprints are clearly labeled to maintain transparency. This section reviews the latest developments and future outlooks in the key areas of omics-based NPs discovery, specifically focusing on metabolomics, genomics, transcriptomics, and proteomics. As well as their integration via AI/ML. The primary tools and resources discussed across these fields are compiled for reference in Table 1.
1. Metabolomics in NPs discovery
- Analytical platforms for metabolomics
Metabolomics is crucial for discovering NPs as it allows for a detailed analysis of the small molecules generated by various organisms. The primary analytical tools used for characterizing the metabolomes of NPs include high-resolution MS techniques, such as LC-MS/MS and gas chromatography-MS (GC-MS), along with NMR spectroscopy21,22. Untargeted LC-MS/MS workflows, like ultra-high-performance LC-MS paired with MS/MS, enable the identification of a wide range of metabolites in complex mixtures, while GC-MS excels in profiling volatile compounds. Additionally, new methods such as direct infusion-tandem MS and MS imaging are enhancing the metabolomics toolkit for NPs23,24. Advanced mass spectrometric techniques, including multimodal MS, which combines positive and negative ion modes or employs multi-stage MSn fragmentation, and hyphenated methods like LC-MS coupled with NMR, offer deeper structural insights into the metabolites that are detected25,26.
- Data processing and computational tools
Specialized software tools play a crucial role in the processing and interpretation of metabolomics data. Programs like XCMS, MZmine, and OpenMS are commonly used to detect and align chromatographic features, or peaks, while meticulous data preprocessing steps such as peak alignment, normalization, and filtering are essential for achieving reproducible metabolic profiles27. The identification of compounds is supported by MS/MS spectral libraries, including NIST and mzCloud, as well as in silico fragmentation tools like SIRIUS and MetFrag28. The Global Natural Products Social (GNPS) platform has emerged as a key resource for molecular networking, which organizes related MS/MS spectra to emphasize chemical families and facilitate the dereplication of known compounds29. For instance, an analysis using GNPS on LC-MS/MS data from Streptomyces cultures successfully identified known antibiotics like spiramycin and actinomycin, alongside previously unreported analogues30,31. The molecular networking technique effectively grouped related unknown spectra, and the addition of ion-identity networking enhanced the connectivity of features. These network-based analyses, when combined with multivariate statistical tools such as principal component analysis and orthogonal partial least squares discriminant analysis, as well as correlation network analysis, can link patterns of metabolite production to specific biological or environmental conditions. In one particular study, GNPS molecular networking of fermentation extracts from a plant-associated Streptomyces demonstrated distinct metabolite profiles that varied depending on the growth mediator used. Known metabolites, including spiramycin, actinomycin, and a carcinogenic compound called lyngbyatoxin, were detected; however, many network nodes did not match any entries in spectral databases32. This finding suggests the presence of genuinely novel metabolites, highlighting how untargeted metabolomics, in conjunction with GNPS, can identify potentially new NPs for further exploration.
Combining metabolomic screening with functional bioassays facilitates the swift prioritization of strains or extracts that yield bioactive compounds. Additionally, emerging community resources like the Natural Products Atlas (NPAtlas), along with integrated metabolome-genome analysis pipelines, are instrumental in associating detected metabolites with their BGCs in the source organisms33,34. This integrative strategy enables researchers to correlate metabolite signals with genomic data, thus enhancing the efficiency of identifying NPs and tracing their origins.
2. Genomics in NPs discovery
Over the past decade, microbial NPs research has entered a transformative "deep-mining era" fueled by advances such as high-throughput sequencing, ultra-sensitive high-resolution MS, and integrated multi-omics approaches35. These tools have shifted discovery efforts from serendipitous isolation toward data-driven, targeted exploration. Genome mining pipelines, including antiSMASH 7.0 and DeepBGC, now enable systematic identification of cryptic BGCs, particularly in under-explored taxa36,37.
- High-throughput sequencing and hybrid assembly
High-throughput DNA sequencing technologies, such as short-read high-throughput DNA sequencing combined with long-read sequencing platforms, have significantly transformed the genomic study of organisms that produce NPs. By employing hybrid assembly strategies that integrate both long and short reads, researchers can achieve high-contiguity genome assemblies, which are essential for capturing entire BGCs within a genome38. Additionally, metagenomic sequencing, which includes the recovery of metagenome-assembled genomes, broadens the discovery of BGCs to uncultured microbes, enabling scientists to explore environmental DNA for its biosynthetic potential39. In instances where conventional genome assembly methods fall short in resolving large or repetitive gene clusters, library-based techniques, such as cosmid or bacterial artificial chromosome libraries, can be utilized to isolate and clone substantial genomic fragments. This approach facilitates the characterization of complete gene clusters through targeted sequencing or heterologous expression, thereby enhancing our understanding of the genetic basis for NPs biosynthesis40,41.
- Genome mining tools
Specialized bioinformatic tools are essential for analyzing genomic data to identify the distinct signatures of secondary metabolite pathways. Programs like antiSMASH, PRISM, and DeepBGC play a crucial role in this process by automatically detecting candidate BGCs. They do this by recognizing specific biosynthetic domains, such as those associated with nonribosomal peptide synthetases, polyketide synthases, ribosomally synthesized and post-translationally modified peptide (RiPP) pathways, and terpene cyclases, among others. To aid in dereplication, databases like MIBiG, which compiles known BGCs and their products, and NPAtlas, which catalogs known NPs, help researchers match new sequences or compounds to established examples37,42,43. Clustering algorithms, such as BiG-SCAPE, along with tools like CORASON, organize related BGCs into families, offering a network view of biosynthetic diversity and facilitating the identification of novel cluster families through comparisons with known ones44. Additional bioinformatic analyses, including BLAST searches and hidden Markov model scans, can predict the potential functions of enzymes encoded within a BGC, while comparative genomics approaches, such as synteny analysis and the co-evolution of genes, help refine these functional predictions45. Furthermore, genomic data enable the annotation of regulatory elements linked to NPs biosynthesis, including pathway-specific promoters and transcriptional regulators46. These insights provide a foundation for pathway engineering; for example, synthetic biology techniques like transformation-associated recombination cloning and Red/ET recombineering empower researchers to capture and express silent or cryptic BGCs in a heterologous host organism, thereby activating the production of their metabolites47.
Genome mining efforts have successfully led to the discovery of new NPs by concentrating on specific genetic signals. For instance, a systematic search for genes linked to glycopeptide antibiotic resistance revealed several unusual BGCs predicted to encode new antibiotics. This method resulted in the identification of two novel compounds, compstatin and cobramycin, both of which exhibit unique mechanisms of action by targeting bacterial autolysins48. In a similar vein, exploring the human microbiome for biosynthetic gene content yielded a lipopeptide antibiotic candidate known as "humicin," recognized for its potential effectiveness against Staphylococcus and Streptococcus species49. In these instances, the compounds were initially identified through computational methods, with their chemical structures predicted based on genomic data. Subsequently, these predicted molecules were chemically synthesized and demonstrated the expected antibiotic activities, thereby validating the genome-based discovery strategy. On a broader scale, large-scale sequencing of various bacteria, including many rare actinomycetes, is uncovering a wealth of "cryptic" BGCs-gene clusters whose products remain unknown. For example, investigations into Streptomyces genomes have revealed thousands of uncharacterized BGCs50. An emerging approach to linking these orphan clusters to actual metabolites involves integrating metabolomic analyses with genomic surveys. By correlating the presence or expression of a gene cluster with the detection of unique metabolite features through platforms like GNPS or mass spectral databases, researchers can start to assign tentative chemical products to the numerous new BGCs, thus facilitating the discovery of truly novel NPs from genomic data.
3. Transcriptomics in natural product discovery
Transcriptomic analyses offer a dynamic perspective on the activity of BGCs by measuring mRNA expression under various conditions. Specifically, RNA-seq experiments can uncover which BGCs are actively transcribed and how their expression levels fluctuate in response to environmental or experimental changes. By comparing transcript levels across different conditions, differential expression analysis, utilizing tools such as DESeq2 or edgeR, can pinpoint BGCs that are either up-regulated or down-regulated, thereby highlighting gene clusters that may be inducible or remain silent under standard conditions51. Although conventional RNA-seq conducted on bulk cultured cells is the typical approach for investigating BGC expression, more advanced techniques like single-cell RNA-seq are starting to be utilized in complex microbiomes. This shift allows for the observation of gene expression in environmental organisms that are difficult to culture52,53. A standard RNA-seq workflow generally includes mapping reads to a reference genome using aligners such as STAR or HISAT2, quantifying gene expression with tools like featureCounts or Kallisto, and normalizing the data to identify significant expression changes. Following this, co-expression network analysis, often performed with the WGCNA package, can cluster genes exhibiting similar expression patterns. When metabolite data is also available, these networks can be further expanded to incorporate metabolites, thus linking specific compounds with co-expressed genes54.
Transcriptomic data have proven to be valuable in identifying conditionally active biosynthetic pathways. A notable example of this approach is found in a detailed comparison of four Salinispora strains51. In this study, it was revealed that more than half of the BGCs in each strain were expressed to some degree, with many clusters that were inactive in one strain showing expression in another. By analyzing gene expression profiles across these strains, the researchers were able to pinpoint several regulatory factors that seemed to either silence or activate specific clusters. This integrative method led to the identification of a previously unknown family of NPs, known as salinipostins. Transcriptomic data indicated a cryptic gene cluster as a potential source for an unidentified compound; when this cluster was induced to express (through regulatory changes), it resulted in the production of salinipostin molecules, which were subsequently isolated and structurally characterized. This study exemplifies how transcriptomics can facilitate the discovery of NPs: by analyzing differential gene expression, researchers can highlight candidate BGCs, and the correlation between gene expression and metabolite profiles provides supporting evidence for gene-metabolite relationships that can be validated through targeted experiments.
4. Proteomics in NPs discovery
- Shotgun and targeted proteomics approaches
Proteomics, which involves the large-scale study of proteins, provides a vital perspective in NPs research by directly measuring the enzymes and proteins that play roles in biosynthetic pathways. Generally, proteomic approaches can be categorized into two main types: shotgun (untargeted) proteomics and targeted proteomics. In shotgun proteomics, proteins extracted from microbial or plant samples undergo enzymatic digestion, typically with trypsin, and the resulting peptides are analyzed using high-resolution LC-MS/MS, often employing advanced mass spectrometers such as quadrupole time-of-flight (QTOF) or high-resolution orbital trap mass spectrometers55. The MS data collected, known as MS/MS spectra, are then matched to peptide sequences by searching against a protein database derived from the organism's genome or a closely related species, utilizing software tools such as MaxQuant or Mascot56,57. This process allows for the quantification of changes in protein abundance under different conditions, which can be achieved either through label-free methods or by employing isotope labeling techniques, such as stable isotope labeling by amino acids in cell culture (SILAC) or isobaric tagging with TMT/iTRAQ reagents, to determine which biosynthetic enzymes are up-regulated or down-regulated58,59. In contrast, targeted proteomics methods, including selected reaction monitoring (SRM) or parallel reaction monitoring (PRM), concentrate on a specific set of peptide ions, often unique peptides from key biosynthetic enzymes or regulatory proteins, enabling precise and sensitive quantification of these peptides across multiple samples60,61.
- Innovative approaches in proteomics
One significant challenge in the proteomic analysis of secondary metabolism is the detection of large biosynthetic enzymes, such as nonribosomal peptide synthetases and polyketide synthases, which often exceed 100 kDa in size and yield few detectable peptides with standard digestion protocols, making them difficult to observe. To tackle this issue, Bumpus et al. developed a method known as PrISM, which enhances the detection of NRPS and PKS proteins by utilizing a specific cofactor found on these enzymes62. PrISM integrates conventional shotgun proteomics with a phosphopantetheinyl (Ppant) ejection assay. Notably, NRPS and PKS enzymes possess a covalently attached Ppant arm on their acyl carrier or peptidyl carrier domains; during MS/MS fragmentation, this prosthetic group frequently generates a unique fragment ion, referred to as the "Ppant ejection" ion. By computationally filtering the LC-MS/MS data for this diagnostic ion, the PrISM method can effectively highlight peptides derived from NRPS and PKS enzymes amidst the complex mixture of all cellular proteins. This strategy has successfully uncovered hidden biosynthetic pathways through proteomics. For instance, the PrISM workflow enabled the detection of peptides from the gramicidin S synthetase complex (GrsA/GrsB) in cultures of Bacillus brevis, establishing a direct link between these NRPS enzymes and the production of the antibiotic gramicidin S in that organism62. Importantly, the proteomic identification of GrsA/GrsB did not necessitate prior genomic knowledge of B. brevis, illustrating that in certain instances, proteomic analysis alone can uncover active NPs biosynthetic pathways even in organisms whose genomes have not yet been sequenced.
- Genome-informed proteomics
When a genome sequence is available, the power of proteomics can be significantly enhanced by utilizing a strain-specific database for peptide identification. For example, in a study analyzing the proteome of Streptomyces lilacinus, researchers conducted two analyses: one by searching spectra against a general protein database and another against a custom database derived from the sequenced genome of that strain63. The genome-informed analysis yielded approximately ten times more identified peptides and detected peptides linked to six distinct BGCs within the organism. Notably, three of these identified clusters corresponded to previously known "orphan" metabolites, including graybactin, rakicidin D, and a specific siderophore, with their biosynthetic enzymes confirmed to be expressed. The remaining detected clusters appeared to be novel and uncharacterized. This example demonstrates that integrating genomic information into proteomic workflows can confirm which biosynthetic pathways are actively expressed in a strain, thus prioritizing those gene clusters for the isolation and characterization of their products.
Additionally, proteomic methods can illuminate the molecular targets and modes of action of NPs, an area often referred to as chemical proteomics. For instance, activity-based protein profiling (ABPP) employs small molecule probes, typically derivatives or analogs of NPs, that react with the cellular targets of the compound, tagging those proteins for enrichment and identification through MS64. This technique can uncover the protein targets that a specific NP or related molecule binds to within the cell. Another method, thermal proteome profiling (TPP), involves heating protein samples in the presence or absence of a compound to determine which proteins show altered thermal stability, indicating a binding interaction65. These chemical proteomic approaches are invaluable for validating the biological targets of NPs and elucidating their mechanisms of action, thereby complementing the discovery of the NPs themselves.
5. ML and AI for omics integration and prediction
- Integration of multi-omics data with ML
An exciting frontier in the discovery of NPs is the application of ML and AI to integrate omics data for functional predictions. In the realm of genomics, supervised ML algorithms, including random forests, support vector machines, and deep neural networks, have been developed to identify patterns in BGCs that are associated with specific types of NPs. For instance, these ML models can categorize BGCs based on the general class of molecules they produce or even forecast certain characteristics of the chemical structure derived from the gene sequence. A review by Dason MS and colleagues highlights various ML tools designed to decode the "language" of BGCs, utilizing DNA or amino acid sequence data to deduce the structures and bioactivities of the corresponding NPs66. These models typically rely on extensive databases of known gene clusters and compounds to learn the relationships between them, enabling predictions about the potential biological activity of new compounds. For example, they can assess whether a product from an uncharacterized BGC might possess antibiotic or anticancer properties. Significant advancements that have facilitated these capabilities include the utilization of large training datasets, such as thousands of known NP structures and gene clusters, innovative representation techniques like sequence embeddings and graph neural networks that capture the features of gene clusters and chemical structures, and multi-parametric models that combine various genomic and chemical descriptors to enhance prediction accuracy.
In the field of metabolomics, AI techniques are significantly enhancing the interpretation of complex mass spectral data. A notable example is the tool CSI:FingerID, which employs ML to predict a molecule's structural fingerprint from its MS/MS spectrum. This capability aids in identifying unknown metabolites by suggesting potential substructures and candidate compounds67. Similarly, deep learning models have been utilized for spectral annotation; convolutional neural networks and, more recently, transformer-based architectures, such as the "DreaMS" model, learn fragmentation patterns from extensive spectral libraries. These models can propose structures for unknown spectra based on the patterns they have learned, often outperforming traditional heuristic methods, particularly for metabolites that do not have exact matches in existing databases68. Beyond structural annotation, ML also contributes to metabolomics by identifying global patterns. For example, unsupervised visualization algorithms such as t-SNE (t-distributed stochastic neighbor embedding) and UMAP (Uniform Manifold Approximation and Projection) can reduce the dimensionality of untargeted metabolomics datasets, facilitating the clustering of samples based on similarities in their metabolite profiles. Meanwhile, supervised classifiers can connect specific metabolite signatures to the phenotypic traits or bioactivities of the samples, further enriching the analysis69,70,71.
ML is increasingly being utilized to integrate various omics datasets, creating a more comprehensive understanding of NPs biosynthesis and function. By merging data from genomics, transcriptomics, proteomics, and metabolomics, integrative models can reveal connections that may be overlooked when examining each data type separately. For instance, researchers can construct gene-metabolite correlation networks that link the expression levels of genes or proteins to the production levels of metabolites. ML models trained on this integrated data can then predict which gene clusters are likely responsible for specific metabolites or biological activities72,73. Advanced multivariate techniques, such as multi-omics factor analysis (available in tools like MOFA) and regularized multivariate methods (found in packages like mixOmics), facilitate the identification of latent factors that encompass different data types, such as a set of co-expressed genes alongside a group of co-occurring metabolites74,75,76. In practical terms, for NPs discovery, various approaches assist in prioritizing candidate gene clusters by demonstrating a connection to observed metabolites or phenotypes. A notable example of this integrative, AI-guided strategy is the discovery of a new family of ribosomally synthesized and RiPPs using the DecRiPPter algorithm. This algorithm utilizes a support vector machine (SVM)-based model to analyze genomic data for cryptic RiPP precursor peptides and then cross-references these predictions with pan-genome analysis to eliminate known compounds. When applied to 1,295 Streptomyces genomes, DecRiPPter successfully identified 42 putative novel RiPP families that had previously been missed by traditional genome mining methods. Among these predicted clusters, one was experimentally validated to produce a new class V lanthipeptide, which the researchers named "pristinin A3." By inducing the expression of this silent gene cluster, the compound pristinin A3 was isolated, and its structure was determined using NMR and MS, revealing two previously unknown lanthionine-forming enzymes encoded within the pathway77. This achievement highlights how AI-driven analyses can uncover hidden NPs: the ML model detected an otherwise unrecognized genetic signal, guiding experimental efforts to discover a novel molecule.
- Emerging AI applications
In addition to the current applications, several emerging AI-driven strategies are set to further revolutionize NPs research. One promising area is computational retrosynthesis and pathway design, where deep learning models are being developed to suggest biosynthetic routes or synthetic chemistry steps for known NPs and their analogs, which could significantly enhance the design and production of novel compounds78,79. Another exciting frontier is the prediction of enzyme function using AI. By leveraging modern protein structure prediction tools based on deep learning algorithms, alongside ML techniques like contrastive learning on protein sequences, researchers are starting to deduce the likely activities of uncharacterized enzymes found in BGCs. This could provide insights into the types of chemical transformations these enzymes catalyze80,81,82. Fully integrated omics-to-chemistry pipelines are being developed, where AI platforms aim to automatically connect mass spectral features to genomic data83,84. In principle, such a system could analyze an LC-MS/MS dataset from a complex sample alongside genome sequences from the same environment, allowing it to predict which gene cluster is responsible for each unidentified metabolite by scoring gene cluster-metabolite matches. Although these approaches are still in their infancy, they suggest a future where AI could autonomously link genes to molecules, greatly speeding up the process from omics data to the discovery of new NPs.
- Data governance and ethical challenges in AI-assisted NPs discovery
Modern multi-omics research produces a significant amount of diverse data, but the effectiveness of AI predictions relies heavily on the quality and consistency of these datasets. When training sets are poorly annotated or biased, they can result in misleading outcomes and unsuccessful validations85. To address this issue, researchers should implement the FAIR principles-Findable, Accessible, Interoperable, and Reusable-to promote data transparency, reproducibility, and extensive reusability86. This approach involves sharing both positive and negative results, thoroughly documenting data preprocessing pipelines, and providing analysis code to facilitate independent verification. Additionally, the adoption of emerging digital infrastructures, such as electronic laboratory notebooks (ELNs) and containerized workflows, is essential for improving data capture and ensuring standardized curation87,88.
While AI significantly speeds up the discovery of NPs, it also brings forth crucial ethical concerns. Algorithms that are trained on incomplete or biased datasets can reinforce existing inequalities, underscoring the necessity for diverse and representative training data89. Additionally, employing explainable AI methods is vital for enhancing transparency and assisting scientists in understanding predictions, ensuring that AI serves as a supportive tool under human control rather than a decision-maker85. Ethical considerations also encompass data privacy, especially when using clinical or human-derived datasets, as well as the implications for the workforce as automation increasingly influences the research environment90. To tackle these issues, AI systems should be subject to regular audits, bias assessments, and stringent regulatory oversight, which will help maintain fairness, accountability, and trustworthiness in the field of NPs research.