$$\rightleftharpoonup{xx}$$
$$\longleftharp{xx}$$,
$$\longrightharp{xx}$$,
GC×GC-TOF MS patterns of high-quality extra-virgin olive oil volatilome exhibit about 500 2D peaks above a signal-to-noise ratio (SNR) threshold of 100. Such a threshold was defined by previous investigations on food volatiles14,27 as the minimum relative signal over threshold to obtain reliable spectra for cross-comparative analysis. Components are distributed over the chromatographic space according to their relative retention in the two chromatographic dimensions, and specifically based on their volatility/polarity in the 1D and volatility in the 2D. Here, column combination is polar × semi-polar (i.e., Carbowax 20M × OV1701).
The 2D pattern shows a high degree of order. Relative retention patterns for homologous series and classes are shown in Figure 1A with annotations (graphics for groups and bubbles for peaks) for linear saturated hydrocarbons (black), unsaturated hydrocarbons (yellow), linear saturated aldehydes (blue), mono-unsaturated aldehydes (red), polyunsaturated aldehydes (salmon), primary alcohols (green), and short-chain fatty acids (cyano).
Detected 2D peaks can then be identified by comparing the average MS spectrum extracted from the entire 2D peak (blob spectrum) or from the largest spectrum (apex spectrum). Figure 2 illustrates the output of the apex spectrum search for blob 5 and returns a high similarity match (first 10 hits) for (E)-2-hexenal. Databases explored are those pre-selected by the analyst in step 8 of the method.
The identification is validated by active retention indexing. The experimental IT value was calculated for the 2D peaks, so that at this stage the library search prioritizes results with coherent values of tabulated IT. Tolerance windows can be customized based on analyst experience, reliability of reference database values according to stationary phase, and analytical conditions applied. New tools for smart calibration of linear retention indices without experimental calibration with n-alkanes, have been recently developed and discussed in a study by Reichenbach et al19.
The collection of identified 2D peaks (i.e., targeted peaks) can be adopted to build a template of targeted peaks to promptly establish reliable correspondences between the same compound across all sample chromatograms. The collection of targeted template peaks is visualized in Figure 1B. Red circles correspond to the 196 targeted compounds, including two Internal Standards (IS) linked to template peaks with connection lines. IS are used for response normalization and connection lines help to visualize which of the included IS will be adopted to normalize each 2D peak/blob response.
In Figure 1B, filled circles indicate positive matches between template peak and the actual pattern while empty circles are for template peaks for which the correspondence was not verified. False negative matches can be limited by appropriate selection of threshold parameters, reference spectra and constraint functions13,14,18,19. For complex patterns with multiple co-elutions, ion peak detection functions that are based on spectral deconvolution are advisable and could be a valid option19. Template peak metadata are shown in the enlarged panel of Figure 1B for (E)-2-hexenal.
The specificity of template matching relies on the possibility to apply constraint functions that limit positive correspondence to those candidate peaks that, falling within the search window of the algorithm, have MS spectral similarity above a certain threshold. In this case, in step 11, similarity thresholds23 were set at 700 according to previous experiments aimed at defining optimal parameters limiting false negative matches14. Highlighted areas of the template peak properties in Figure 1B show the information about the reference MS spectrum string and the qCLIC constraint function (i.e., (Match("
By applying the template to all chromatograms of a set, one could encounter challenging situations as in the case of partial misalignment of patterns. This can be due to oven temperature inconsistencies, carrier gas flow/pressure instabilities, or because of a manual intervention on the system as in the case of column substitution or modulator loop-capillary replacement14,28. Figure 3 shows a situation of a partial misalignment between the targeted template and the actual chromatogram. For minimal misalignments, interactive template transforms (Figure 3, control panel) can reposition template peaks for a better fit. Once repositioned, the template can be matched to establish correspondences. In the example, the template (Figure 3, step 12) peaks correctly match with the actual 2D pattern. In case of severe misalignments, not discussed here, the repetition of match-transform-update actions can iteratively adapt the template peaks position to the actual peak pattern12,13,14.
Here, the targeted peaks (i.e., known analytes) provide about 40% of the chromatographic result (196 targeted peaks of about 500 detectable peaks on average). The other 60% of compounds, together with the information they bring, are not taken into consideration in targeted analysis. To make the investigation truly comprehensive, consistent cross-alignment of untargeted 2D peaks should also be established. The first application where template matching was extended to all detectable analytes dealt with the complex volatilome of roasted coffee7. This process is automated with a software (e.g., Investigator), shown here in steps 14–15.
In this process, pre-targeted images belonging to the sample set under study (20 samples) are used to define reliable peaks by cross-matching of all image patterns29. Subsequently, a composite chromatogram is built from which one can identify UT reliable peaks and peak regions (i.e., 2D peaks footprint) in the so-called feature template17.
For analyses acquired at 70 eV, the process determined 144 reliable peaks with relaxed reliability29, 76 of which belong to the targeted peaks list. Based on these 144 reliable peaks, the process aligns all chromatograms consistently with the average retention times of the reliable peaks and then combines them to create a composite chromatogram. Figure 4 shows a list of all samples labeled according to the production region of the oil (left) and the list of reliable peaks/blob volumes in each sample (right).
The untargeted feature template is composed of 2D peaks from analytes detected in the composite chromatogram, shown in Figure 5A, that are matched by the reliable-peaks template (n = 168 – red circles for targeted peaks and green circles for untargeted peaks). The mass spectra of the composite peaks, as well as their retention times, are recorded in the feature template as shown for (Z)-3-hexenol acetate in the enlarged area. Peak-regions are shown in Figure 5B as red colored graphics; they are instead defined by the outlines of all 2D peaks detected in the composite chromatogram (n = 3578).
When unsupervised pattern recognition by Principal Component Analysis is applied to targeted peaks distribution within the 20 analyzed samples, Sicilian and Tuscany oils cluster separately suggesting that pedo-climatic conditions and terroir impact the relative prevalence of volatiles. The results are shown in Figure 6A and the PCA results from the reliable peaks distribution are shown in Figure 6B. The two approaches cross-validate that oils from different geographical areas have different, while coherent, chemical signatures whether targeted or untargeted compounds, or both, are mapped.
Finally, the software enables prompt and effective re-alignment of patterns across parallel detection channels. In this application, the re-alignment is proposed for tandem ionization signals. The ion source of the MS multiplexes between two ionization energies (i.e., 70 and 12 eV) at an acquisition frequency of 50 Hz per channel30. The two resulting chromatographic patterns are closely aligned while spectral data (i.e., spectral signatures and responses) bring complementary information with different dynamic ranges of response26,27. The aligned patterns allow extracting features (2D peaks and peak-regions) with univocal IDs (i.e., chemical names for targeted peaks and unique numbering # for untargeted peaks and peak-regions).
Template matching allows effective cross-alignment. In this situation, there is not much misalignment, but MS constraints must be relaxed to allow matches for UT peaks. On the other hand, featured UT peak-regions, that have no MS constraints, are promptly matched without any false negative matches. Figure 5C shows an enlarged area of a 12 eV chromatogram where the feature template built from 70 eV data is matched. Reliable UT peaks are positively matched because of the lowered qCLIC constraints (e.g., DMF threshold at 600). To note, at 12 eV, there are fewer detected peaks due to the limited fragmentation induced by low ionization energy.

Figure 1: Bidimensional contour plot and targeted template. (A) Contour plot of the volatile fraction of an extra-virgin olive oil from Tuscany. Ordered patterns of homolog series and classes are highlighted with different colors and lines: linear saturated hydrocarbons (black line and 2D contours) unsaturated hydrocarbons (yellow), linear saturated aldehydes (blue) mono-unsaturated aldehydes (red), polyunsaturated aldehydes (salmon), primary alcohols (green) and short-chain fatty acids (cyano). (B) Overimposed targeted template of known analytes (red colored circles) with connection lines linking Internal Standards (ISs). Panels show 2D peak/blob properties metadata (Decanal) or Template peak properties. Please click here to view a larger version of this figure.

Figure 2: Apex MS search. Output of the apex MS search for blob 5. List of the database entries with the highest similarity match and related metadata available from the library. Please click here to view a larger version of this figure.

Figure 3: Template realignment. Workflow illustrating the steps that allow re-alignment of the template by transformation. Please click here to view a larger version of this figure.

Figure 4: GC Investigator interface. Investigator panel with all selected images labeled according to the production Region of the oil (left) and the list of reliable peaks/blob volumes in each sample (right). Please click here to view a larger version of this figure.

Figure 5: Targeted and UT template. (A) Reliable peaks as resulting from the automated processing in step 11; red circles correspond to known analytes while green circles are unknowns. In the superimposed panel, template object properties are shown for the (Z)-3-hexenal. (B) Enlarged area that shows the UT peaks (red and green circles) and peak-regions (red graphics) of the UT template matched on a sample oil acquired at 70 eV ionization energy. (C) UT template matched on a sample oil acquired at 12 eV ionization energy. Please click here to view a larger version of this figure.

Figure 6: PCA loading plots. They show the natural conformation of samples (oils from Tuscany and Sicily) as they result by (A) targeted peaks distribution or (B) UT peaks distribution. Please click here to view a larger version of this figure.
Supplemental Files. Please click here to download these files.