Large-scale data analysis based on the genomes, transcriptomes, proteomes and metabolomes of plants provides unprecedented insight into the behavior of complex systems, such as the terroir characteristics of wine which reflect the interactions between grapevine plants and their environment. Because the terroir of a wine can be distinct even when identical grapevine clones are grown in different vineyards, genomics analysis is of little use because the clonal genomes are identical. Instead it is necessary to look at correlations between gene expression and the metabolic properties of the berries, which determine the quality traits of wine. The analysis of gene expression at the level of the transcriptome benefits from the similar chemical properties of all transcripts, which facilitates quantitative analysis by exploiting universal characteristics such as hybridization to immobilized probes on microarrays. In contrast, universal analytical methods in proteomics and metabolomics are more challenging because of the huge physical and chemical diversity of individual proteins and metabolites. In the case of metabolomics this diversity is even more extreme because individual metabolites differ vastly in size, polarity, abundance and volatility, so no single extraction process or analytical method offers a holistic approach.
Among the analytical platforms suitable for non-volatile metabolites, those based on high performance liquid chromatography coupled to mass spectrometry (HPLC-MS) are much more sensitive than alternatives such as HPLC with ultraviolet or diode array detectors (HPLC-UV, HPLC-DAD) or nuclear magnetic resonance (NMR) spectroscopy, but quantitative analysis by HPLC-MS can be influenced by phenomena such as the matrix effect and ion suppression/enhancement1-3. The investigation of such effects during the analysis of Corvina grape berries by HPLC-MS using an electrospray ionization source (HPLC-ESI-MS), showed that sugars and other molecules with the lowest retention times were strongly underreported, probably also reflecting the large number of molecules in this zone, and that the abundance of other molecules could be underestimated, overestimated or unaffected by the matrix effect, but the data normalization for the matrix effect seemed to have limited impact on the overall results4,5. The method described herein is optimized for the analysis of medium-polarity metabolites that accumulate at high levels in grape berries during ripening, and which are significantly impacted by the terroir. They include anthocyanins, flavonols, flavan-3-ols, procyanidins, other flavonoids, resveratrol, stilbenes, hydroxycinnamic acids and hydroxybenzoic acids, which together determine the color, taste and health-related properties of wines. Other metabolites, such as sugars and aliphatic organic acids, are ignored because quantitation by HPLC-MS is unreliable due to the matrix effect and ion suppression phenomena5. Within the polarity range selected by this method, the approach is untargeted in that it aims to detect as many different metabolites as possible6.
Transcriptomics methods that allow thousands of grapevine transcripts to be monitored simultaneously are facilitated by the availability of the complete grapevine genome sequence7,8. Early transcriptomics methods based on high-throughput cDNA sequencing have evolved with the advent of next-generation sequencing into a collection of procedures collectively described as RNA-Seq technology. This approach is rapidly becoming the method of choice for transcriptomics studies. However, a large body of literature based on microarray, which allow thousands of transcripts to be quantified in parallel by hybridization, has accumulated for grapevine. Indeed, before RNA-Seq became a mainstream technology, many dedicated commercial microarray platforms had been developed allowing grapevine transcriptome to be inspected in great detail. Among the vast variety of platforms, only two allowed genome-wide transcriptome analysis9. The most evolved array allowed the hybridization of up to 12 independent samples on a single device, thus reducing the costs of each experiment. The 12 sub-arrays each comprised 135,000 60-mer probes representing 29,549 grapevine transcripts. This device has been used in a large number of studies10-24. These two platforms have now been discontinued but a new custom microarray has recently been designed and represents a more recent development as it contains an even greater number of probes representing additional newly discovered grapevine genes25.
The large-sale datasets produced by transcriptomics and metabolomics analysis require suitable statistical methods for data analysis, including multivariate techniques to determine correlations between different forms of data. The most widely used multivariate techniques are those based on projection, and these can be unsupervised, such as principal component analysis (PCA), or supervised, such as bidirectional orthogonal projection to latent structures discriminant analysis (O2PLS-DA)26. The protocol presented in this article utilizes PCA for exploratory data analysis and O2PLS-DA to identify differences between groups of samples.