May 16th, 2022
LEfSe (LDA Effect Size) is a tool for high-dimensional biomarker mining to identify genomic features (such as genes, pathways, and taxonomies) that significantly characterize two or more groups in microbiome data.
The application of linear discriminant analysis effect size can solve the problem of finding good biomarkers with statistical differences between biologic groups. Linear discriminant analysis effect size provides a convenient method for identifying genomic biomarkers to characterize statistical differences between biologic groups. Be sure to take care with each step of the procedure, because the outcome of each preceding step may effect the subsequent step.
After generating a linear discriminant analysis effect size input file, run the commands as indicated to exclude the possibility of dependencies conflict and create a conda environment for linear discriminant analysis effect size. Use the commands as indicated to activate the created environment and to install linear discriminant analysis effect size with channel bioBakery. To format data for linear discriminant analysis effect size, run the command to format the original file to the internal linear discriminant analysis effect size format.
To calculate the linear discriminant analysis effect size, run the command to perform a linear discriminant analysis and to generate the resulting data file. After the analysis, use the commands as indicated to plot the effect size of the biomarkers in a PDF file and to draw the species tree to display the biomarkers in a cladogram. To plot the differences of a single biomarker among different groups, use the command as indicated.
All of the features can also be drawn using the command if desired. For online LEfSe analysis using the Galaxy server, navigate to the server. To upload the relevant files, click up and Choose local file to select the files.
Then, select tabular format and click Start. To format the data for linear discriminant analysis effect size, click LEfSe and Format data for LEfSe, select the specific rows for class, and click Execute. To calculate the linear discriminant analysis effect size, click LEfSe and LDA Effect Size.
Select the parameter values according to the analysis requirements, and click Execute. To plot the linear discriminant analysis effect size results, click LEfSe and Plot LEfSe Results and click Execute. To plot the cladogram, select the appropriate parameter values and click Plot Cladogram and Execute.
To plot one feature, select the appropriate parameter values and click Plot One Feature and Execute. To plot differential features, select the appropriate parameter values and click Plot Differential Features and Execute button. Here, the linear discrimination analysis scores of microbial communities with significant differences in each group, determined by analyzing the 16 S-ribosomal RNA gene sequences of three samples, is shown.
In this figure, the biomarkers with significant differences in species trees between different classification levels can be observed. The circles radiating from inside to the outside represent the classification levels from phylum to genus, with the diameter of each species circle representing the level of abundance of each classification. The species with no significant differences appear in yellow, and the significantly different species biomarkers are colored to match the corresponding groups.
The corresponding species names of the biomarkers shown in the plot are listed here. Here, a representative of abundance bar plot for one biomarker that shows differences among the different groups according to the linear discriminant analysis effect size results is shown. The solid line represents the average relative abundance, the dotted line represents the median relative abundance, and each column represents the relative abundance of each sample in different groups.
Principal component analysis can also be performed, as the dimensionality education is directly related to the data dimension, and the projected coordinate system is orthogonal. As the demand for high-dimensional data analysis increases, this method will aid in the exploration of biomarker features of interest.
View the full transcript and gain access to thousands of scientific videos
LEfSe (LDA Effect Size) is a powerful tool designed for high-dimensional biomarker mining. It identifies genomic features that significantly characterize different groups in microbiome data.
LEfSe enables biopharma teams to systematically identify statistically robust microbial biomarkers that distinguish biological groups in high-dimensional microbiome datasets. This capability supports early discovery, target validation, and translational research by providing quantitative evidence for group-specific features. Integrating LEfSe into discovery pipelines enhances predictive confidence and informs risk-adjusted portfolio decisions.
LEfSe fits within the discovery-to-preclinical continuum by enabling robust biomarker identification, statistical validation, and visualization of group differences in microbiome data.