All procedures involving human participants were reviewed and approved by the Ethics Committee of Chongqing Chenjiaqiao Hospital (approval No. 20240131). Written informed consent was obtained from all participants before sample collection. All stool samples and participant information were handled using de-identified study codes, as listed in the Table of Materials.
Participant screening and enrollment
Newly diagnosed patients with T2DM who visited Chongqing Chenjiaqiao Hospital between June 2023 and December 2023 were screened for eligibility. Participants were enrolled in the DM group if they were between 20 and 65 years of age, had a new diagnosis of T2DM, agreed to provide a fresh stool sample, and completed the required medical history assessment and clinical examinations. Healthy volunteers recruited during the same period were included as the non-diabetic matched (NM) group and were matched to the DM group as closely as possible with respect to age, sex, and general demographic characteristics. Participants were excluded if they had hematological disorders, central nervous system diseases, active rheumatic disease, autoimmune disease, acute or chronic gastrointestinal infection, chronic diarrhea, constipation, active or healing gastrointestinal ulcer, inflammatory bowel disease, irritable bowel syndrome, intestinal tuberculosis, gastrointestinal tumors, other malignancies, severe cardiac insufficiency, severe hepatic insufficiency, severe renal insufficiency, other severe metabolic diseases, malnutrition, immunodeficiency, congenital metabolic disorders, psychiatric illness, sedative-hypnotic use, or drug abuse. Participants were also excluded if they had received antibiotics, acid-suppressing agents, gastrointestinal motility drugs, probiotics, glucocorticoids, or immunosuppressants within 1 month before stool collection. In addition, individuals who had experienced diarrhea, undergone gastrointestinal surgery or gastrointestinal endoscopy, or reported sudden changes in living environment or dietary habits within 1 month before stool collection were excluded. A unique study code was assigned to each eligible participant before sample collection. From the eligible cohort, 12 newly diagnosed T2DM and 12 NM participants were selected for downstream 16S rDNA sequencing and microbiome analyses based on sample availability and sequencing quality requirements.
Stool sample collection and storage
Each participant was provided with a sterile stool collection container labeled only with the assigned study code to ensure de-identification. Participants were instructed to empty their bladders before defecation to minimize urine contamination and to collect fresh stool directly into the sterile container without contact with toilet water, urine, disinfectants, or other potential contaminants. Immediately after collection, approximately 1–2 g of stool was transferred into a sterile cryogenic tube using a sterile disposable sampling spoon. The tube was tightly sealed and immediately placed on dry ice or in a pre-cooled transfer container. All samples were transported to the laboratory within 2 h of collection. Upon arrival at the laboratory, the study code was verified, and each tube was inspected for leakage or visible contamination. Collection and storage times were recorded for all samples. Stool samples were subsequently stored at −80 °C until genomic DNA extraction. Samples were considered acceptable for downstream analysis only if they remained completely frozen during storage and transport and showed no evidence of tube leakage or external contamination.
Genomic DNA extraction
Stool samples were removed from −80 °C storage and placed on ice prior to processing. Each sample was thawed only until homogenization was possible to minimize degradation associated with repeated freeze–thaw cycles. Approximately 200 mg of stool was transferred into a sterile microcentrifuge tube, and the lysis buffer provided in the stool DNA extraction kit was added according to the manufacturer’s instructions listed in the Table of Materials. Samples were vortexed vigorously for 30 s to achieve complete homogenization and then incubated at room temperature for 5–10 min to facilitate cell lysis. Following lysis, samples were centrifuged at 12,000 × g. for 10 min at 4 °C. The supernatant was carefully transferred to a new sterile microcentrifuge tube without disturbing the pellet. Genomic DNA was subsequently eluted using the elution buffer supplied with the kit or nuclease-free water according to the manufacturer’s protocol. Extracted DNA was stored at −20 °C for short-term use or at −80 °C for long-term preservation. DNA samples considered suitable for downstream analysis appeared clear and free of visible particulate material.
DNA quality assessment
DNA quality was assessed using agarose gel electrophoresis and fluorescence-based quantification. A 1% agarose gel was prepared in electrophoresis buffer, and the extracted DNA samples, mixed with loading buffer, were loaded into the gel wells. Electrophoresis was performed until the DNA bands were adequately separated, after which the gel was examined under ultraviolet or blue-light illumination. DNA integrity was considered acceptable when an intact genomic DNA band was observed without marked degradation or excessive smearing. DNA concentration was subsequently quantified using a fluorescence-based DNA quantification system. Samples were then diluted to the concentration required for downstream PCR amplification. High-quality DNA samples typically exhibited a distinct band on agarose gel electrophoresis and were sufficiently concentrated for subsequent amplification procedures.
PCR amplification of the 16S rDNA V3–V4 region
The V3–V4 hypervariable region of bacterial 16S rDNA was amplified using barcoded primers. The primer sequences were as follows: forward primer 5′-ACTCCTACGGGAGGCAG-3′ and reverse primer 5′-GGACTACHVGGGTWTCTAAT-3′. PCR reactions were prepared on ice in a final volume of 20 µL containing 4 µL of PCR buffer, 2 µL of nucleotide mixture (2.5 mmol/L), 0.8 µL each of forward and reverse primers (5 µmol/L), 0.4 µL of high-fidelity DNA polymerase, 10 ng of template DNA, and nuclease-free water to volume. Reaction mixtures were gently mixed by pipetting and briefly centrifuged to collect the liquid at the bottom of the tube. PCR amplification was performed using the following cycling conditions: initial denaturation at 95 °C for 5 min, followed by 25 cycles of denaturation at 95 °C for 30 s, annealing at 55 °C for 30 s, and extension at 72 °C for 30 s, with a final extension step at 72 °C for 10 min. To minimize amplification bias, each sample was amplified in triplicate reactions, and the resulting PCR products from the same sample were subsequently pooled into a single tube. Successful amplification was confirmed by the presence of a clear amplicon band of the expected size on agarose gel electrophoresis.
PCR product verification and purification
PCR products were verified using 2% agarose gel electrophoresis. The combined PCR products from each sample were loaded onto the gel and electrophoresed until the target amplicon bands were clearly separated. Bands corresponding to the expected amplicon size were excised using a clean sterile blade and purified with a gel extraction kit according to the manufacturer’s instructions listed in the Table of Materials. Purified PCR products were subsequently eluted in elution buffer or nuclease-free water. The concentration of each purified PCR product was quantified using a fluorescence-based DNA quantification system, and the concentration values were recorded for downstream library preparation. Successful amplification and purification were confirmed by a single, clear band at the expected amplicon size on agarose gel electrophoresis.
Library preparation and pooling
Sequencing libraries were prepared using a DNA library preparation kit according to the manufacturer’s instructions listed in the Table of Materials. Sequencing adapter sequences were ligated to the purified amplicons following the standard library preparation workflow. Adapter-ligated products were subsequently purified using either gel extraction or a bead-based purification method, depending on the selected library preparation protocol. Purified library products were evaluated using 2% agarose gel electrophoresis to confirm appropriate fragment distribution. Library concentrations were quantified using a fluorescence-based DNA quantification system. Based on DNA concentration and fragment size, each library was normalized to the same molar concentration to ensure balanced sequencing depth across samples. Equal molar amounts of individual libraries were then pooled at a 1:1 ratio to generate the final sequencing library pool. Prior to sequencing, the pooled library was denatured according to the sequencing platform protocol before sample loading. Libraries considered suitable for sequencing exhibited a clear fragment distribution within the expected size range and sufficient concentration for downstream sequencing analysis.
Paired-end sequencing
The final pooled library was loaded onto a paired-end sequencing platform according to the manufacturer’s standard operating procedures. Paired-end sequencing was performed using either a 2 × 250 bp or 2 × 300 bp run configuration to ensure adequate sequencing depth for comprehensive characterization of microbial diversity in each sample. Sequencing quality was monitored throughout the run using platform-generated quality metrics. Following completion of sequencing, raw paired-end FASTQ files were exported for each sample for downstream bioinformatics analysis. Sequencing datasets were considered acceptable when they demonstrated sufficient read depth per sample, appropriate Q30 quality scores, and successful barcode assignment.
Raw sequence processing
Raw paired-end FASTQ files were imported into a microbiome bioinformatics analysis pipeline for downstream processing. Sequences were demultiplexed based on the barcodes assigned to each sample. Quality filtering was subsequently performed to remove reads containing low-quality scores, ambiguous bases, or sequencing artifacts. Low-quality bases located at the ends of reads were trimmed based on quality score distribution to improve overall sequence reliability. After quality control, paired-end reads were merged across overlapping regions to reconstruct full-length sequences. Only high-quality merged sequences were retained for subsequent microbiome analyses.
Amplicon sequence variant (ASV) generation and taxonomic annotation
Sequence denoising was performed using a validated denoising algorithm, including DADA2 or Deblur, to correct sequencing errors and improve sequence accuracy. Chimeric sequences were identified and removed during the denoising process. Representative ASV sequences and corresponding ASV abundance tables were subsequently generated for each sample. Taxonomic assignment of ASV sequences was performed using a Naive Bayes classifier, and taxonomic annotation was conducted against the SILVA 138 reference database. Taxonomic abundance tables were exported at the phylum, family, and genus levels for downstream analysis. The final ASV dataset included representative sequences, abundance information, and corresponding taxonomic annotations for each sample.
Alpha and beta diversity analysis
Alpha diversity indices, including Chao, ACE, Shannon, Simpson, and Coverage indices, were calculated to evaluate within-sample microbial diversity and richness. Alpha diversity indices were expressed as mean ± standard deviation. Data normality was assessed using the Shapiro–Wilk test. Normally distributed variables were compared between the DM and NM groups using Student’s t-tests, whereas non-normally distributed variables were analyzed using Wilcoxon rank-sum tests. Beta diversity distances were subsequently calculated using a suitable distance matrix to evaluate differences in microbial community composition between samples. Principal coordinates analysis (PCoA) was performed to visualize differences in microbial community structure between groups. In addition, analysis of similarities (ANOSIM) was conducted to determine whether intergroup differences exceeded intragroup variation, and the corresponding ANOSIM R and p values were reported. Alpha diversity analysis reflected within-sample microbial diversity, whereas beta diversity analysis evaluated differences in microbial community composition between groups.
Differential taxonomic analysis
Microbial community composition was summarized at the phylum, family, and genus levels. Stacked bar plots were generated to visualize the relative abundance of dominant taxa across individual samples and study groups. Pan/Core curves and Venn diagrams were further constructed to compare shared and group-specific ASVs between the DM and NM groups. Differences in taxon abundance between groups were evaluated using Wilcoxon rank-sum tests. LEfSe analysis was subsequently performed to identify microbial taxa that contributed most strongly to intergroup discrimination. Differential taxonomic analysis enabled the identification of microbial taxa enriched in either the DM or NM group.
Functional prediction
PICRUSt2 or an equivalent validated functional prediction pipeline was used to infer microbial functional pathways from 16S rDNA sequencing profiles. ASV abundance data were normalized according to the requirements of the selected analytical pipeline, and predicted functions were mapped to KEGG or MetaCyc pathway annotations. Predicted pathway abundances were subsequently compared between the DM and NM groups. Significantly different pathways were visualized using heatmaps or other appropriate graphical approaches. Predicted functions were interpreted as inferred microbial metabolic potential rather than directly measured metabolite abundance. Functional prediction analysis enabled the identification of candidate microbial pathways that differed between groups, including pathways associated with carbohydrate and amino acid metabolism.
Mendelian randomization analysis
Genome-wide association study (GWAS) summary statistics for gut microbiota were obtained from publicly available datasets10. Genome-wide significant association data for microbial taxa were retrieved from the NHGRI-EBI GWAS Catalog (https://www.ebi.ac.uk/gwas/) using accession numbers ranging from GCST90032172 to GCST90032644. Additional metagenomic data were accessed from the FINRISK 2002 cohort through the European Genome-Phenome Archive (Research ID: EGAS00001005020). GWAS summary statistics for type 2 diabetes mellitus were also obtained from the NHGRI-EBI GWAS Catalog (accession number: ebi-a-GCST006867). Genetic variants associated with microbial taxa were selected as instrumental variables according to predefined statistical thresholds, and variants in linkage disequilibrium were excluded. Exposure and outcome datasets were harmonized to ensure consistent allele orientation. Mendelian randomization analysis was subsequently performed using the inverse variance weighted (IVW) method as the primary analytical approach. Sensitivity analyses were conducted using the MR-Egger and weighted median methods. Instrument strength was evaluated using F-statistics, whereas heterogeneity among instrumental variables was assessed using Cochran’s Q test. Horizontal pleiotropy was examined using the MR-Egger intercept test. Bonferroni correction was applied to account for multiple comparisons. Mendelian randomization findings were interpreted as genetically predicted associations rather than definitive evidence of causality.
Data output and endpoint
Final analytical outputs included ASV abundance tables, taxonomic composition plots, alpha diversity metrics, beta diversity analyses, differential taxonomic results, predicted functional pathway profiles, and Mendelian randomization estimates. All sample identifiers in the sequencing datasets were verified to ensure consistency with the corresponding de-identified study codes. Raw sequencing files, processed ASV tables, statistical outputs, and figure source files were archived for downstream analysis and data management. The protocol was considered complete once high-quality sequencing data, taxonomic profiles, diversity metrics, predicted functional pathway analyses, and Mendelian randomization results had been successfully generated and verified.