A subscription to JoVE is required to view this content. Sign in or start your free trial.

Research Article

Gut Microbiota and Metabolomic Changes In Type 2 Diabetes Mellitus: Insights From 16S rDNA Sequencing and Bioinformatics

83 views

DOI:

10.3791/70219

June 26th, 2026

In This Article

Summary

This study identifies reduced gut microbiota diversity and altered metabolic pathways in type 2 diabetes (T2DM), highlighting increased proteobacteria and decreased beneficial taxa, suggesting that microbial dysbiosis is associated with T2DM and may contribute to its pathophysiology.

Abstract

The global rise in type 2 diabetes mellitus (T2DM) underscores the need to better understand its underlying biological mechanisms, particularly those involving host–microbiome interactions. This study aimed to characterize gut microbial diversity, taxonomic composition, and predicted metabolic pathways in newly diagnosed T2DM patients compared with the non-diabetic matched (NM) group. Fresh stool samples were analyzed using 16S rDNA sequencing. Alpha diversity (Chao, ACE, Shannon, and Simpson indices) and beta diversity were calculated to assess microbial community structure. Taxonomic differences were evaluated using Wilcoxon rank-sum tests and linear discriminant analysis effect size (LEfSe). Functional pathway prediction was performed using phylogenetic investigation of communities by reconstruction of unobserved states (PICRUSt2) based on KEGG and MetaCyc annotations. Mendelian randomization (MR) analysis, including inverse-variance weighting, MR-Egger, and weighted median methods, was applied to assess genetically predicted associations between microbial taxa and T2DM. Results showed reduced microbial richness, as reflected by lower Chao and ACE indices, and altered diversity structure, as reflected by lower Shannon and higher Simpson indices, in T2DM patients, accompanied by significant compositional shifts. Increased relative abundance of Proteobacteria and decreased abundance of beneficial taxa such as Lachnospiraceae and Blautia were observed. Functional prediction indicated reduced abundance of pathways related to the non-oxidative pentose phosphate pathway, isobutanol biosynthesis, and L-isoleucine biosynthesis. MR analysis provided complementary evidence supporting associations between specific microbial taxa and T2DM susceptibility. In conclusion, T2DM is associated with reduced microbial richness, altered diversity structure, and distinct taxonomic and functional changes. These findings highlight the relevance of gut microbiota in T2DM and support the potential utility of microbiome-based biomarkers and therapeutic strategies. Further studies are required to validate these findings and clarify underlying mechanisms.

Introduction

Type 2 diabetes mellitus (T2DM) has become one of the fastest-growing chronic diseases worldwide, driven by rapid urbanization, Westernized dietary patterns, and increasingly sedentary lifestyles1,2. It is a major contributor to global morbidity and mortality and imposes a substantial public health burden3. Although genetic susceptibility plays a role, the rapid rise in T2DM incidence highlights the importance of environmental and metabolic determinants4. Identifying modifiable biological mechanisms is therefore essential for improving early detection and prevention strategies. Recent advances have identified the gut microbiota as a key regulator of metabolic homeostasis5. The intestinal microbial ecosystem influences glucose metabolism, immune function, intestinal barrier integrity, and chronic low-grade inflammation, all of which are central to the pathogenesis of T2DM6. Dysbiosis has been associated with obesity, insulin resistance, endotoxemia, and altered energy balance, suggesting a mechanistic link between microbial alterations and disease progression7. In addition, microbial metabolites, including short-chain fatty acids, bile acids, and branched-chain amino acids, have been implicated in modulating insulin sensitivity and host metabolism8.

While accumulating evidence supports the role of the gut microbiome in T2DM, most existing studies rely on observational designs, which limit causal inference. Alternative approaches, such as shotgun metagenomics and targeted metabolomics, can provide higher-resolution functional insights but are often constrained by cost, computational complexity, and limited feasibility in exploratory or pilot-scale studies. In contrast, 16S rDNA sequencing is a cost-effective, widely adopted method for profiling microbial community composition, although its taxonomic resolution is limited to the species level and it cannot directly quantify functional metabolites. Therefore, 16S rDNA sequencing was selected for this study as a cost-effective, scalable approach suitable for exploratory clinical cohorts, while enabling integration with downstream functional prediction and complementary analytical frameworks. Therefore, integrating 16S-based microbial profiling with bioinformatics functional prediction and complementary approaches such as Mendelian randomization (MR) may provide a balanced framework for exploring both compositional and potentially causal relationships.

Notably, there remains a relative lack of data focusing specifically on newly diagnosed, treatment-naïve T2DM patients. In such populations, microbial alterations are less likely to be confounded by pharmacological interventions, particularly metformin, which is known to significantly reshape gut microbiota composition and function9. This represents an important gap in the current literature, as many existing studies include heterogeneous patient populations with varying treatment exposures and limited integration of microbiome profiling with genetically informed analytical approaches.

In this context, the present study aimed to characterize gut microbiota diversity, taxonomic composition, and predicted functional pathways in newly diagnosed T2DM patients compared with the NM group, using 16S rDNA sequencing and bioinformatics analysis. In addition, Mendelian randomization was applied using publicly available genome-wide association data to explore potential genetically predicted associations between specific microbial taxa and T2DM risk. This combined analytical strategy allows complementary insights from microbial community profiling and population-level genetic data. Specifically, 16S rDNA sequencing provides direct characterization of microbial composition in the study population, whereas Mendelian randomization leverages genetic instruments from independent cohorts to assess whether observed associations are consistent with genetically predicted relationships. This integration is intended to address a key limitation of observational microbiome studies by providing complementary evidence that supports the robustness of microbiota–T2DM associations. Therefore, integrating these approaches is intended to provide cross-validation rather than redundancy, thereby strengthening the interpretability of microbiota–T2DM associations.

The present study focuses on three aspects. First, it evaluates gut microbiota alterations in newly diagnosed, treatment-naïve T2DM patients to reduce confounding from pharmacological interventions. Second, it integrates taxonomic profiling with predicted functional pathways to provide a systems-level view of microbiota-associated metabolic changes. Third, it incorporates Mendelian randomization analysis to assess whether observed associations are consistent with genetically predicted relationships, thereby providing methodological triangulation. Collectively, these aspects provide a complementary and reproducible framework for microbiome analysis rather than a purely descriptive study, distinguishing this work from prior studies that typically rely on single-method observational designs and do not integrate microbial profiling, functional prediction, and genetically informed inference within a unified analytical framework.

Access restricted. Please log in or start a trial to view this content.

Protocol

All procedures involving human participants were reviewed and approved by the Ethics Committee of Chongqing Chenjiaqiao Hospital (approval No. 20240131). Written informed consent was obtained from all participants before sample collection. All stool samples and participant information were handled using de-identified study codes, as listed in the Table of Materials.

Participant screening and enrollment

Newly diagnosed patients with T2DM who visited Chongqing Chenjiaqiao Hospital between June 2023 and December 2023 were screened for eligibility. Participants were enrolled in the DM group if they were between 20 and 65 years of age, had a new diagnosis of T2DM, agreed to provide a fresh stool sample, and completed the required medical history assessment and clinical examinations. Healthy volunteers recruited during the same period were included as the non-diabetic matched (NM) group and were matched to the DM group as closely as possible with respect to age, sex, and general demographic characteristics. Participants were excluded if they had hematological disorders, central nervous system diseases, active rheumatic disease, autoimmune disease, acute or chronic gastrointestinal infection, chronic diarrhea, constipation, active or healing gastrointestinal ulcer, inflammatory bowel disease, irritable bowel syndrome, intestinal tuberculosis, gastrointestinal tumors, other malignancies, severe cardiac insufficiency, severe hepatic insufficiency, severe renal insufficiency, other severe metabolic diseases, malnutrition, immunodeficiency, congenital metabolic disorders, psychiatric illness, sedative-hypnotic use, or drug abuse. Participants were also excluded if they had received antibiotics, acid-suppressing agents, gastrointestinal motility drugs, probiotics, glucocorticoids, or immunosuppressants within 1 month before stool collection. In addition, individuals who had experienced diarrhea, undergone gastrointestinal surgery or gastrointestinal endoscopy, or reported sudden changes in living environment or dietary habits within 1 month before stool collection were excluded. A unique study code was assigned to each eligible participant before sample collection. From the eligible cohort, 12 newly diagnosed T2DM and 12 NM participants were selected for downstream 16S rDNA sequencing and microbiome analyses based on sample availability and sequencing quality requirements.

Stool sample collection and storage

Each participant was provided with a sterile stool collection container labeled only with the assigned study code to ensure de-identification. Participants were instructed to empty their bladders before defecation to minimize urine contamination and to collect fresh stool directly into the sterile container without contact with toilet water, urine, disinfectants, or other potential contaminants. Immediately after collection, approximately 1–2 g of stool was transferred into a sterile cryogenic tube using a sterile disposable sampling spoon. The tube was tightly sealed and immediately placed on dry ice or in a pre-cooled transfer container. All samples were transported to the laboratory within 2 h of collection. Upon arrival at the laboratory, the study code was verified, and each tube was inspected for leakage or visible contamination. Collection and storage times were recorded for all samples. Stool samples were subsequently stored at −80 °C until genomic DNA extraction. Samples were considered acceptable for downstream analysis only if they remained completely frozen during storage and transport and showed no evidence of tube leakage or external contamination.

Genomic DNA extraction

Stool samples were removed from −80 °C storage and placed on ice prior to processing. Each sample was thawed only until homogenization was possible to minimize degradation associated with repeated freeze–thaw cycles. Approximately 200 mg of stool was transferred into a sterile microcentrifuge tube, and the lysis buffer provided in the stool DNA extraction kit was added according to the manufacturer’s instructions listed in the Table of Materials. Samples were vortexed vigorously for 30 s to achieve complete homogenization and then incubated at room temperature for 5–10 min to facilitate cell lysis. Following lysis, samples were centrifuged at 12,000 × g. for 10 min at 4 °C. The supernatant was carefully transferred to a new sterile microcentrifuge tube without disturbing the pellet. Genomic DNA was subsequently eluted using the elution buffer supplied with the kit or nuclease-free water according to the manufacturer’s protocol. Extracted DNA was stored at −20 °C for short-term use or at −80 °C for long-term preservation. DNA samples considered suitable for downstream analysis appeared clear and free of visible particulate material.

DNA quality assessment

DNA quality was assessed using agarose gel electrophoresis and fluorescence-based quantification. A 1% agarose gel was prepared in electrophoresis buffer, and the extracted DNA samples, mixed with loading buffer, were loaded into the gel wells. Electrophoresis was performed until the DNA bands were adequately separated, after which the gel was examined under ultraviolet or blue-light illumination. DNA integrity was considered acceptable when an intact genomic DNA band was observed without marked degradation or excessive smearing. DNA concentration was subsequently quantified using a fluorescence-based DNA quantification system. Samples were then diluted to the concentration required for downstream PCR amplification. High-quality DNA samples typically exhibited a distinct band on agarose gel electrophoresis and were sufficiently concentrated for subsequent amplification procedures.

PCR amplification of the 16S rDNA V3–V4 region

The V3–V4 hypervariable region of bacterial 16S rDNA was amplified using barcoded primers. The primer sequences were as follows: forward primer 5′-ACTCCTACGGGAGGCAG-3′ and reverse primer 5′-GGACTACHVGGGTWTCTAAT-3′. PCR reactions were prepared on ice in a final volume of 20 µL containing 4 µL of PCR buffer, 2 µL of nucleotide mixture (2.5 mmol/L), 0.8 µL each of forward and reverse primers (5 µmol/L), 0.4 µL of high-fidelity DNA polymerase, 10 ng of template DNA, and nuclease-free water to volume. Reaction mixtures were gently mixed by pipetting and briefly centrifuged to collect the liquid at the bottom of the tube. PCR amplification was performed using the following cycling conditions: initial denaturation at 95 °C for 5 min, followed by 25 cycles of denaturation at 95 °C for 30 s, annealing at 55 °C for 30 s, and extension at 72 °C for 30 s, with a final extension step at 72 °C for 10 min. To minimize amplification bias, each sample was amplified in triplicate reactions, and the resulting PCR products from the same sample were subsequently pooled into a single tube. Successful amplification was confirmed by the presence of a clear amplicon band of the expected size on agarose gel electrophoresis.

PCR product verification and purification

PCR products were verified using 2% agarose gel electrophoresis. The combined PCR products from each sample were loaded onto the gel and electrophoresed until the target amplicon bands were clearly separated. Bands corresponding to the expected amplicon size were excised using a clean sterile blade and purified with a gel extraction kit according to the manufacturer’s instructions listed in the Table of Materials. Purified PCR products were subsequently eluted in elution buffer or nuclease-free water. The concentration of each purified PCR product was quantified using a fluorescence-based DNA quantification system, and the concentration values were recorded for downstream library preparation. Successful amplification and purification were confirmed by a single, clear band at the expected amplicon size on agarose gel electrophoresis.

Library preparation and pooling

Sequencing libraries were prepared using a DNA library preparation kit according to the manufacturer’s instructions listed in the Table of Materials. Sequencing adapter sequences were ligated to the purified amplicons following the standard library preparation workflow. Adapter-ligated products were subsequently purified using either gel extraction or a bead-based purification method, depending on the selected library preparation protocol. Purified library products were evaluated using 2% agarose gel electrophoresis to confirm appropriate fragment distribution. Library concentrations were quantified using a fluorescence-based DNA quantification system. Based on DNA concentration and fragment size, each library was normalized to the same molar concentration to ensure balanced sequencing depth across samples. Equal molar amounts of individual libraries were then pooled at a 1:1 ratio to generate the final sequencing library pool. Prior to sequencing, the pooled library was denatured according to the sequencing platform protocol before sample loading. Libraries considered suitable for sequencing exhibited a clear fragment distribution within the expected size range and sufficient concentration for downstream sequencing analysis.

Paired-end sequencing

The final pooled library was loaded onto a paired-end sequencing platform according to the manufacturer’s standard operating procedures. Paired-end sequencing was performed using either a 2 × 250 bp or 2 × 300 bp run configuration to ensure adequate sequencing depth for comprehensive characterization of microbial diversity in each sample. Sequencing quality was monitored throughout the run using platform-generated quality metrics. Following completion of sequencing, raw paired-end FASTQ files were exported for each sample for downstream bioinformatics analysis. Sequencing datasets were considered acceptable when they demonstrated sufficient read depth per sample, appropriate Q30 quality scores, and successful barcode assignment.

Raw sequence processing

Raw paired-end FASTQ files were imported into a microbiome bioinformatics analysis pipeline for downstream processing. Sequences were demultiplexed based on the barcodes assigned to each sample. Quality filtering was subsequently performed to remove reads containing low-quality scores, ambiguous bases, or sequencing artifacts. Low-quality bases located at the ends of reads were trimmed based on quality score distribution to improve overall sequence reliability. After quality control, paired-end reads were merged across overlapping regions to reconstruct full-length sequences. Only high-quality merged sequences were retained for subsequent microbiome analyses.

Amplicon sequence variant (ASV) generation and taxonomic annotation

Sequence denoising was performed using a validated denoising algorithm, including DADA2 or Deblur, to correct sequencing errors and improve sequence accuracy. Chimeric sequences were identified and removed during the denoising process. Representative ASV sequences and corresponding ASV abundance tables were subsequently generated for each sample. Taxonomic assignment of ASV sequences was performed using a Naive Bayes classifier, and taxonomic annotation was conducted against the SILVA 138 reference database. Taxonomic abundance tables were exported at the phylum, family, and genus levels for downstream analysis. The final ASV dataset included representative sequences, abundance information, and corresponding taxonomic annotations for each sample.

Alpha and beta diversity analysis

Alpha diversity indices, including Chao, ACE, Shannon, Simpson, and Coverage indices, were calculated to evaluate within-sample microbial diversity and richness. Alpha diversity indices were expressed as mean ± standard deviation. Data normality was assessed using the Shapiro–Wilk test. Normally distributed variables were compared between the DM and NM groups using Student’s t-tests, whereas non-normally distributed variables were analyzed using Wilcoxon rank-sum tests. Beta diversity distances were subsequently calculated using a suitable distance matrix to evaluate differences in microbial community composition between samples. Principal coordinates analysis (PCoA) was performed to visualize differences in microbial community structure between groups. In addition, analysis of similarities (ANOSIM) was conducted to determine whether intergroup differences exceeded intragroup variation, and the corresponding ANOSIM R and p values were reported. Alpha diversity analysis reflected within-sample microbial diversity, whereas beta diversity analysis evaluated differences in microbial community composition between groups.

Differential taxonomic analysis

Microbial community composition was summarized at the phylum, family, and genus levels. Stacked bar plots were generated to visualize the relative abundance of dominant taxa across individual samples and study groups. Pan/Core curves and Venn diagrams were further constructed to compare shared and group-specific ASVs between the DM and NM groups. Differences in taxon abundance between groups were evaluated using Wilcoxon rank-sum tests. LEfSe analysis was subsequently performed to identify microbial taxa that contributed most strongly to intergroup discrimination. Differential taxonomic analysis enabled the identification of microbial taxa enriched in either the DM or NM group.

Functional prediction

PICRUSt2 or an equivalent validated functional prediction pipeline was used to infer microbial functional pathways from 16S rDNA sequencing profiles. ASV abundance data were normalized according to the requirements of the selected analytical pipeline, and predicted functions were mapped to KEGG or MetaCyc pathway annotations. Predicted pathway abundances were subsequently compared between the DM and NM groups. Significantly different pathways were visualized using heatmaps or other appropriate graphical approaches. Predicted functions were interpreted as inferred microbial metabolic potential rather than directly measured metabolite abundance. Functional prediction analysis enabled the identification of candidate microbial pathways that differed between groups, including pathways associated with carbohydrate and amino acid metabolism.

Mendelian randomization analysis

Genome-wide association study (GWAS) summary statistics for gut microbiota were obtained from publicly available datasets10. Genome-wide significant association data for microbial taxa were retrieved from the NHGRI-EBI GWAS Catalog (https://www.ebi.ac.uk/gwas/) using accession numbers ranging from GCST90032172 to GCST90032644. Additional metagenomic data were accessed from the FINRISK 2002 cohort through the European Genome-Phenome Archive (Research ID: EGAS00001005020). GWAS summary statistics for type 2 diabetes mellitus were also obtained from the NHGRI-EBI GWAS Catalog (accession number: ebi-a-GCST006867). Genetic variants associated with microbial taxa were selected as instrumental variables according to predefined statistical thresholds, and variants in linkage disequilibrium were excluded. Exposure and outcome datasets were harmonized to ensure consistent allele orientation. Mendelian randomization analysis was subsequently performed using the inverse variance weighted (IVW) method as the primary analytical approach. Sensitivity analyses were conducted using the MR-Egger and weighted median methods. Instrument strength was evaluated using F-statistics, whereas heterogeneity among instrumental variables was assessed using Cochran’s Q test. Horizontal pleiotropy was examined using the MR-Egger intercept test. Bonferroni correction was applied to account for multiple comparisons. Mendelian randomization findings were interpreted as genetically predicted associations rather than definitive evidence of causality.

Data output and endpoint

Final analytical outputs included ASV abundance tables, taxonomic composition plots, alpha diversity metrics, beta diversity analyses, differential taxonomic results, predicted functional pathway profiles, and Mendelian randomization estimates. All sample identifiers in the sequencing datasets were verified to ensure consistency with the corresponding de-identified study codes. Raw sequencing files, processed ASV tables, statistical outputs, and figure source files were archived for downstream analysis and data management. The protocol was considered complete once high-quality sequencing data, taxonomic profiles, diversity metrics, predicted functional pathway analyses, and Mendelian randomization results had been successfully generated and verified.

Access restricted. Please log in or start a trial to view this content.

Results

Baseline demographic and clinical characteristics of the initially screened clinical cohort are presented in Table 1. A representative subgroup consisting of 12 newly diagnosed T2DM and 12 NM participants was subsequently selected for 16S rDNA sequencing and downstream microbiome analyses.

Sequencing depth evaluation

The Shannon index was calculated, and the rarefaction curve was plotted to assess the sequencing depth. A total of 1,...

Access restricted. Please log in or start a trial to view this content.

Discussion

In recent years, the human microbiome, particularly the gut microbiota, has emerged as a complex and influential factor in host metabolism, immune regulation, and disease development. In a healthy host, the gut microbiota contributes to immune homeostasis and intestinal barrier integrity11. Dysbiosis has been associated with a range of pathological conditions, including metabolic disorders, autoimmune diseases, and malignancies12,13,<...

Access restricted. Please log in or start a trial to view this content.

Disclosures

The authors declare that the research was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.

AUTHORS’ CONTRIBUTIONS:

Linrui Xie and Kemiao Chen performed data analysis and drafted the manuscript. Xican Pan participated in sample collection and experimental procedures. Xuemei Zhong and Xiaoyu Li conceived and designed the study, supervised the project, and critically revised the manuscript. All authors reviewed and approved the final manuscript.

Acknowledgements

The authors thank all participants and staff involved in sample collection and data processing for their contributions to this study.

Access restricted. Please log in or start a trial to view this content.

Materials

List of materials used in this article
NameCompanyCatalog NumberComments
16S rRNA Primers (V3–V4)Sangon BiotechCustomUniversal primers for bacterial 16S region
80 °C FreezerHaier BiomedicalDW-86L626Long-term sample storage
Agarose Gel System (Sub-Cell GT)Bio-Rad170-4401PCR product validation
AMPure XP Magnetic BeadsBeckman CoulterA63880DNA purification and size selection
Bead-beating Tubes (Lysing Matrix E)MP Biomedicals116913050Mechanical lysis of microbial cells
Bioinformatics WorkstationDellPrecision 7920Data analysis using QIIME2 and R
DADA2 pipelineBioconductorVersion 1.28Sequence denoising and ASV generation; https://benjjneb.github.io/dada2
DeblurQIIME2Version 2023.5Sequence denoising and error correction; https://docs.qiime2.org
European Genome-Phenome ArchiveEMBL-EBIEGAS00001005020Public metagenomic dataset repository; https://ega-archive.org
GWAS CatalogNHGRI-EBIAccessions GCST90032172–GCST90032644Public genome-wide association summary statistics database; https://www.ebi.ac.uk/gwas
Illumina MiSeq Sequencing SystemIlluminaSY-410-1003High-throughput sequencing platform
KEGG databaseKyoto Encyclopedia of Genes and GenomesRelease 107Functional pathway annotation database; https://www.genome.jp/kegg
MetaCyc databaseSRI InternationalVersion 27.1Metabolic pathway database; https://metacyc.org
MicrocentrifugeEppendorf5424RCooling centrifuge for sample preparation
Nextera XT Library Prep KitIllumina20018705Library prep for sequencing
PCR Master Mix (2x PrimeSTAR Max Premix)Takara BioRR350AFor high-fidelity PCR
PCR ThermocyclerBio-RadT100PCR amplification instrument
PICRUSt2PICRUSt2 Development TeamVersion 2.5.2Functional pathway prediction from 16S rDNA sequencing data; https://github.com/picrust/picrust2
QIAamp Fast DNA Stool Mini KitQIAGEN51504DNA extraction from stool samples
QIAquick Gel Extraction KitQIAGEN28704DNA purification from agarose gel
Qubit FluorometerThermo Fisher ScientificQ33226DNA quantification
R softwareR Foundation for Statistical ComputingVersion 4.3.1Statistical analysis and Mendelian randomization analysis; https://www.r-project.org
SILVA 138 databaseSILVA ribosomal RNA database projectRelease 138Reference database for taxonomic annotation; https://www.arb-silva.de
Stool Collection TubeSarstedt76.9923.001Sterile stool specimen collection tube
Stool DNA KitOmegaD4015-02DNA extraction kit
TwoSampleMR packageMR-BaseVersion 0.5.7Mendelian randomization analysis in R; https://mrcieu.github.io/TwoSampleMR

Reprints and Permissions

Tags

16S rDNA SequencingMicrobial DiversityBioinformatics AnalysisTaxonomic CompositionFunctional Pathway PredictionMendelian RandomizationMicrobiome Biomarkers