Understanding the interactions between genes, environment and management practices in agriculture could allow more accurate prediction and management of product yield and quality. Plant metabolites are influenced by factors such as the genome, environment (climate, rainfall etc.), and in an agriculture setting, the way crops are managed (i.e., application of fertilizer, fungicide etc.). Unlike the genome, the metabolome is influenced by all of these factors and hence metabolomics data provides a biochemical fingerprint of these interactions at a particular time. There are usually one of two goals for a metabolomics-based study: firstly, to achieve a deeper understanding of the organism's biochemistry and help explain the mechanism of response to perturbation (abiotic or biotic stress) in relation to the physiology; and secondly, to associate biomarkers with the perturbation under study. In both cases, the outcome of having this knowledge is a more precise management strategy to achieve the goal of improved yield size and quality.
The plant metabolome is predicted to contain thousands1 of small molecules with varied physicochemical properties. Currently, no metabolomics platforms (predominantly mass spectrometry and nuclear magnetic resonance spectroscopy) can capture the entire metabolome in a single analysis. Developing such techniques (sample preparation, metabolite extraction and analysis), which provide as great a coverage of the metabolome as possible within a single analytical run, is a key aim for metabolomics researchers. Previous untargeted metabolomics analyses of wheat grain have combined data from multiple chromatographic separations and acquisition polarities and/or instrumentation for greater metabolome coverage. However, this has required samples to be prepared and acquired separately for each modality. For example, Beleggia et al.2 prepared a derivatized sample for the GC-MS analysis of polar analytes in addition to the GC-MS analysis of the nonpolar analytes. Das et al.3 used both GC- and LC-MS methods to improve coverage in their analyses; however, this approach would generally require separate sample preparations as described above as well as two independent analytical platforms. Previous analyses of wheat grain using GC-MS2,3,4 and LC-MS3,5 platforms have yielded 50 to 412 (55 identified) features for GC-MS, 409 for combined GC-MS and LC-MS and several thousand for an LC-MS lipidomics analysis5. By combining at least two modes into a single analysis, extended metabolome coverage can be maintained, increasing the richness of biological interpretation while also offering savings in both time and cost.
To permit the efficient separation of a wide range of lipid species by reversed-phase chromatography, modern lipidomics methodologies commonly use a high proportion of isopropanol in the elution solvent6, providing amenability to lipid classes that might otherwise be unresolved by the chromatography. For an efficient lipid separation, the starting mobile phase is also much higher in organic composition7 than the typical reversed phase chromatographic methods, which consider other classes of molecules. The high organic composition at the start of the gradient makes these methods less suitable to many other classes of molecules. Most notably, reversed phase liquid chromatography employs a binary solvent gradient, starting with a mostly aqueous composition and increasing in organic content as the elution strength of the chromatography is increased. To this end, we sought to combine the two approaches to achieve separation of both lipid and non-lipid classes of metabolites within a single analysis.
Here, we present a chromatographic method that uses a third mobile phase and enables a combined traditional reversed phase and lipidomics-appropriate chromatography method using a single sample preparation and one analytical column. We have adopted many of the quality control measures and data filtering steps that have previously been implemented in predominantly clinical metabolomics studies. These approaches are useful in determining robust features with high technical reproducibility and biological relevance and excludes those which do not meet these criteria. For example, we describe repeat analysis of the pooled QC sample8, QC correction9, data filtering9,10 and imputation of missing features11.