$$\rightleftharpoonup{xx}$$
$$\longleftharp{xx}$$,
$$\longrightharp{xx}$$,
Metabolomics efficiently identifies and quantifies metabolites, targeting the metabolic cycles that become deranged during disease. The quality of the results depends on the meticulous execution of each step in the metabolomics approach. Every stage, from sample selection and collection to pathway identification, is critical in accurately identifying the primary factors contributing to the disease. Before performing metabolomics, a thorough review of the literature is essential, and careful attention must be paid at every stage.
In designing a metabolomics study, the selection of a sample is crucial for extracting significant information and ensuring ideal sample preparation, whether for targeted or untargeted metabolic fingerprinting of the clinical condition in consideration. Proper initial sample processing is crucial, and protocols should be consistently followed for all samples throughout the study49. Extra caution is needed when working with clinical samples. Commonly used samples for NMR-based metabolomics include biofluids, cell extracts, and tissue extracts50,51. Cerebrospinal fluid (CSF) and urine often require minimal pre-treatment and are frequently studied in contexts such as kidney-related ailments, neurodegenerative diseases, and drug toxicity. Whole blood, plasma, and serum provide comprehensive information for a broad spectrum of critical diseases and are preferred because of ease of collection. A wide range of other biofluids, such as seminal fluid, saliva, bile, dialysis fluid, amniotic fluid, exhaled breath condensate, lung aspirates, synovial fluid, and mini bronchoalveolar lavage fluid (mBALF), have been extensively used for their specificity and accuracy using NMR52,53,54. Each biofluid presents its advantages and challenges concerning collection, availability, and informational content, depending on the study or experimental design55. We have focused on serum samples of ARDS patients for this study as blood samples typically offer information at a whole-body level and are not significantly influenced by extreme daily variations. On the other hand, urine samples are subject to various fluctuations and can be impacted by diet, lifestyle, medications, and environmental factors27. Proper sample storage is also important to avoid enzymatic activity or microbial degradation. Samples are typically cryopreserved at -80 °C with minimal freeze-thaw cycles to prevent metabolic breakdown and maintain the stability of the sample.The sample can be stored at -80 °C for extended periods, allowing NMR experimentation to be conducted conveniently. However, planning and performing the experiment as soon as possible is advisable for optimal results.
In this study, we employed an 800 MHz NMR spectrometer for data acquisition, which provides high-resolution spectral data. For this experiment, we adhered to the same acquisition parameters mentioned above. However, it is important to note that the same acquisition parameters outlined in this protocol can be applied using NMR spectrometers operating at other magnetic field strengths. The number of scans can be increased when operating at a lower magnetic field to improve resolution. While higher field strengths may offer enhanced resolution, the core aspects of the acquisition process, such as pulse sequences, relaxation delays, and temperature control, remain consistent across different instruments, ensuring the reproducibility and reliability of the results.
Following a standardized procedure is essential to maintain reproducibility and impartiality, as deviations can hinder subsequent data analysis and evaluation. Each step of the NMR experimentation must be performed meticulously, as the quality of the spectra depends on them. Special attention should be given to shimming, as it defines peak shape, and acquisition parameters should be optimized according to the sample and study design. For samples with low metabolic concentration in the sample, the number of scans can be increased. Data acquisition in NMR involves standard optimization of both 1D and 2D experiments tailored to the choice of sample and the specific information required. Various NMR experiments have been optimized to obtain high-quality spectra from biofluids, minimizing the impact of unwanted macromolecules. The CPMG pulse sequence56 is preferred for samples like serum, plasma, or cerebrospinal fluid. This technique suppresses the broad signals from macromolecules such as lipids and proteins, which have long transverse relaxation times and can obscure the resonances of low molecular weight metabolites.
Data pre-processing encompasses the intermediate methods used to ensure uniform and homogeneous data analysis and its interpretation. This process includes avoiding undesirable signals or spectral regions prone to inconsistencies and correcting peak shifts. In biofluids, excluding unwanted regions often means removing signals, such as those from water, that dominate the 4.6 ppm to 5 ppm range. Peak shifts, caused by variations in pH, temperature, salt concentration, dilution factors, and ionic composition, can be minimized by using buffers at a homeostatic pH. Additionally, the use of standard reference compounds, peak alignment, and binning can help reduce peak shifts. Peak alignment is crucial to correlate patterns across different spectra acquired under identical conditions.
Spectra processing must be done properly to ensure accurate binning later in the analysis. All spectra should be calibrated and aligned uniformly, and base and phase corrections can be performed manually or automatically. Proper phase and base correction, as well as calibration, are critical. Performing baseline correction is essential to avoid distortions that can affect the identification and quantification of metabolites57. Phase correction is necessary to ensure no distorted or inverted peaks can affect the interpretation of results. Carefully decide on the criteria for selecting significant metabolites and filter the results accordingly. Since binning provides results in the form of ppm values, accurately identify the metabolite's name using various sources mentioned in the protocol. If multiple metabolites are present at the same ppm values, identify them through peak shape.
Before performing the statistical analysis, normalization is done to reduce variations in the dataset, whether induced or not induced, due to differences in magnitude and fluctuations in metabolite concentrations. This process involves various methods of centering, scaling, and transformation, chosen based on the data's distribution and variability. The goal is to minimize unwanted systematic bias while preserving the biological factors of interest. Multivariate methods, including PCA, PLS-DA, and OPLS-DA, are used to establish differential metabolites with VIP values greater than 1. Model predictability and cross-validation are determined using R2 (goodness of fit) and Q2 (goodness of prediction) values, while permutation test statistics validate the model. A Pearson correlation-based heat map assesses model robustness by enumerating upregulated and downregulated metabolites in sub-phenotypes and the resultant endotypes.
Univariate methods, such as the student t-test and ANOVA, like Bonferroni correction or minimization of FDR (false discovery rate) and Benjamini-Hochberg, are used to reduce the probability of false positives. These methods provide a rough measure of potentially important features independent of their correlation within or between molecules and other confounding factors. Significant metabolites are identified based on a p-value less than 0.05, which are then used to substantiate metabolic endotypes in outcome subgroups45.
Individual box whisker plots and AUROC are employed to evaluate the specificity and sensitivity of the endotypes in outcome subgroups. The clinical score is combined with the model in AUROC to validate the model's accuracy. The AUC value > 0.8 is considered significant46. Various other tests are available online to test the accuracy of the discrimination and efficacy of metabolites identified. Some of them are random forest classification, heat map, correlation, etc.45.
Various approaches are utilized to quantify significant metabolites, including peak integration, which can be challenging for beginners due to their advanced nature. Identify characteristic peaks for metabolites by referring to individual chemical shifts for precise identification. Determine the area of these characteristic peaks relative to the area of the reference peak, TSP. To calculate the absolute concentration of a metabolite (Cm), use the following equation:
Cm = Im.nr.Cr/Ir.nm
Where Cr is the concentration of the reference compound (TSP) in the solution, Im and Ir are the integrated peak areas (can be obtained using NMR processing and analysis tool) of the metabolite and reference (TSP), respectively, and nm and nr are the number of protons representing the metabolite and reference peaks, respectively24.
Alternatively, automated NMR metabolite quantification software is used to obtain metabolite concentrations for single or multiple spectra based on the concentration of the reference standard, TSP. Quantification using software simplifies the process and reduces the likelihood of errors compared to other methods. While this research paper highlights the use of software for metabolite quantification, there can be issues if the metabolite of interest is not present in the software library. Additionally, the software is paid, so only researchers with an authorized license can use it.
Metabolomics has proven instrumental in unraveling the complexities of various diseases by identifying dysregulated metabolic cycles contributing to disease severity. However, translating these findings into clinical practice has been challenging. The medical community is now emphasizing a personalized medicine approach to enhance disease outcomes. While identifying dysregulated metabolites was previously sufficient, current efforts focus on precise quantification to define disease-specific ranges. This information is invaluable for clinicians in customizing treatment strategies. The advancements in quantification techniques have facilitated research efforts, with numerous research groups worldwide successfully employing these approaches using specialized software.
This progress underscores the potential of metabolomics to significantly impact personalized medicine and improve clinical outcomes. Moreover, the integration of metabolomics data with other omics data, such as genomics and proteomics, promises to provide a more comprehensive understanding of disease mechanisms and further refine therapeutic interventions. As technology continues to evolve, the precision and utility of metabolomics in clinical settings are expected to increase, paving the way for more effective and personalized healthcare solutions.