$$\rightleftharpoonup{xx}$$
$$\longleftharp{xx}$$,
$$\longrightharp{xx}$$,
Animal models are widely used to study human diseases, because of their assumed similarity to humans in terms of genetics, anatomy, and physiology. Moreover, animal models often serve as gatekeepers to clinical therapies and can have a huge impact on the success of translational research. Careful selection of the optimal animal model can reduce the number of misleading animal studies. Recently, the relevance of animal models for translational research has been controversially discussed, particularly because analyzing the same datasets obtained from human inflammatory diseases and related mouse models led to contradictory conclusions 1,2. This discussion revealed a fundamental problem during analyzing omics data: standardized approaches for systematic data analysis are needed in order to reduce biased gene selection and to increase the robustness of interspecies comparisons 3.
Traditionally, the analysis of transcriptomics data (and other omics data) is done at the single-gene level and includes an initial step of gene selection based on stringent cut-off parameters (e.g., fold change >2.0, p value <0.05). However, the setting of initial cut-off parameters often is subjective, arbitrary and not biologically justified, and can even lead to opposite conclusions1,2. Furthermore, initial gene selection generally restricts the analysis to a few highly up- and downregulated genes and is thus not sensitive enough to include the majority of genes that were differentially expressed to a lesser extent.
With the rise of the genomics era in the early 2000s and the increasing knowledge of biological pathways and contexts, alternative statistical approaches were developed that allowed to circumvent the limitations of single-gene level analyses. Gene set enrichment analysis (GSEA)4, which is one of the widely accepted methods for the analysis of transcriptomics data, makes use of a-priori defined groups of genes (e.g., signaling pathways, proximal location on a chromosome etc.). GSEA first maps all detected unfiltered genes to the intended gene sets (e.g., pathways), irrespective of their individual change in expression. This approach thus also includes moderately regulated genes that would otherwise be lost with single-gene level analyses. The additive change in expression within gene sets is subsequently performed using running sum statistics.
Despite its wide use in medical research, GSEA and related set enrichment approaches are not self-evidently taken into account for the analysis of complex omics data. Here, we describe a protocol for comparing omics data from human samples with those from mouse models in order to identify the ideal model for translational studies. We demonstrate the applicability of the protocol based on a collection of mouse models that are used for mimicking human inflammatory disorders. However, this analysis pipeline is not restricted to human-mouse comparisons and is amendable to further research questions.