$$\rightleftharpoonup{xx}$$
$$\longleftharp{xx}$$,
$$\longrightharp{xx}$$,
The field of biological research has witnessed significant advancements in recent years, particularly in the area of omics technologies. These technologies provide valuable insights into the complex nature of biological systems. However, each omics technology offers a unique perspective on biological components, necessitating the integration of multi-omics data to obtain a comprehensive understanding.
Multi-omics encompasses various classes of biomolecules that can be quantitatively defined thanks to the advent of new and powerful high-throughput sequencing techniques. Among the different types of omics technologies are genomics, epigenomics, transcriptomics, proteomics, metabolomics, metagenomics, lipidomics, and glycomics. Genomics involves the study of an organism's genomes, while epigenomics focuses on the supporting structure of the genome, including protein and RNA binders, alternative DNA structures, and chemical modifications on DNA. Transcriptomics encompasses the study of all RNA molecules, including mRNA, rRNA, tRNA, and other non-coding RNA. Proteomics involves the study of proteins, including modifications made to specific groups of proteins. Metabolomics focuses on the ensemble of small molecules (metabolites) within a biological matrix. Metagenomics involves the study of microbial communities in well-defined habitats with specific physico-chemical properties. Lipidomics encompasses the study of the entire complement of cellular lipids, while glycomics focuses on the study of the glycome, including carbohydrates and sugars1.
The integration of multi-omics data has gained increasing attention in the scientific community due to its potential to unravel complex biological phenomena. By combining data from multiple omics technologies, researchers can overcome the limitations of individual datasets and gain a more holistic view of biological systems. This integrated approach enables the identification of novel biomarkers, the discovery of disease mechanisms, and the elucidation of intricate biological pathways.
The number of citations of the terms "Multiomics" and "Multi-omics" in PubMed has significantly increased over the years, from 307 in 2018 to 1414 in 2021 to 3933 in 2023. The integration of different types of omics variables is becoming increasingly common as it allows for a deeper investigation into the mechanisms underlying diseases and dysfunctions of organisms. Single omics approaches provide a limited, partial view of the hidden biology, as they focus on a single perspective. However, by integrating multi-omics data, we can shed light on the interplay of different biomolecules, understanding the relationships within multiple layers, and bridging the gap between genotype and phenotype. Overall, multi-omics approaches can help answer important questions such as classifying different subgroups of diseases, predicting fundamental disease-associated biomarkers, and gaining a better understanding of biological pathways and mechanisms. In the following sections, the different omics datasets can also be referred to as data "views" or data "blocks".
Techniques for multi-omics integration can be classified into three main subgroups, as described by Reel et al. (2021)2 and Ritchie et al. (2015)3 (Figure 1).
Low-level, early integration or concatenation: This approach involves concatenating variables from each single dataset into a single matrix. However, early integration does not consider the unique distribution of each omics data type and may assign more weight to certain omics data types with larger dimensions. It also poses challenges such as an increased risk of the curse of dimensionality, added noise, highly correlated variables, and computational scalability issues. Despite these limitations, early integration allows for the identification of coordinated changes across multiple omic layers and enhances biological interpretation.
Mid-level, middle integration or transformation-based: In this approach, mathematical integration models are applied to the multiple layers of omics data. Middle integration focuses on the fusion of subsets or representations extracted from the sources. Two sub-approaches within middle integration are the middle-up approach and the middle-down approach. The middle-up approach involves concatenated scores obtained from dimensionality reduction on each block, making it suitable for handling heterogeneous data. However, it may lack interpretability. The middle-down approach involves local variable selection and subsequent analysis on concatenated variable subsets, allowing for easier interpretation of the models. Middle integration offers advantages such as improved signal-to-noise ratio, reduced dimensionality, and improved statistical power.
High-level, late integration or model-based: This approach involves performing analyses at each single omic level and combining the results in an ad-hoc fashion. It includes the fusion of results from single block models to identify biomarkers from each source and provide a joint interpretation of the results. Late integration does not increase the dimensionality of the input space and works with the unique distribution of each omics data. It is particularly appropriate when one omic layer is more predictive than others. However, late integration may overlook cross-omics relationships and face challenges related to the lack of understanding of the connections between initial data blocks and the potential loss of biological information through individual modeling.