Evidence acquisition
The study was designed and reported in accordance with the PRISMA guidelines18. The protocol was registered in the PROSPERO database (registration number: CRD420261307956). The completed PRISMA 2020 checklist has been uploaded as a separate reporting checklist (Supplementary File 1). All the materials used in this study are listed in the Table of Materials.
Literature search strategy
Seven databases—PubMed, Embase, the Cochrane Library, Web of Science, CNKI, Wanfang Data, and VIP—were searched from their respective inception dates through January 20, 2026. Database-specific controlled vocabulary (MeSH or Emtree) was combined with free-text terms for diabetic kidney disease, traditional Chinese medicine, and randomized controlled trials. The reference lists of pertinent reports were also checked to identify additional eligible records. Title and abstract screening began on February 14, 2026; full-text assessment began on February 25, 2026; and data extraction began on March 5, 2026. Two reviewers worked independently at each stage, resolving disagreements by discussion and, when needed, adjudication by a third reviewer.
Inclusion and exclusion criteria
Inclusion criteria
Eligible studies were randomized controlled trials involving participants diagnosed with diabetic kidney disease, without restrictions on age, sex, or disease duration. Interventions consisted of traditional Chinese herbal decoctions administered alone or in combination with conventional Western medical therapy. Eligible comparators included conventional treatment, Western medical therapy, or placebo. The primary outcomes were glycated hemoglobin (HbA1c) and 24-hour urinary protein excretion (24 hUP), while the secondary outcomes were serum creatinine (Cr) and blood urea nitrogen (BUN).
Exclusion criteria
Case reports, retrospective studies, reviews, animal experiments, and duplicate publications were excluded. Studies evaluating interventions other than decoction-based Chinese herbal medicine, including single-herb preparations and proprietary Chinese medicine formulations, were also excluded. Additional reasons for exclusion included unavailable or non-extractable outcome data, unclear identification of primary or secondary outcomes, failure of the study population to meet the diagnostic criteria for diabetic kidney disease, and insufficient distinction between the intervention and control groups.
Literature screening and data extraction
Records retrieved from the databases were imported into EndNote 21 (Clarivate Analytics) and de-duplicated. Two reviewers then assessed eligibility independently against the prespecified criteria. Screening proceeded in two passes: records that were clearly irrelevant were removed after title and abstract review, and reports retained at that stage underwent full-text evaluation. Reviewer disagreements were settled through discussion; a third reviewer adjudicated unresolved cases. Study selection and reasons for exclusion were documented in a PRISMA flow diagram. Using a prespecified form, two reviewers independently recorded the first author, year of publication, participant and arm sample sizes, participant characteristics, intervention and comparator details, treatment duration, and, for each reported outcome, its definition, unit, assessment time, mean, and standard deviation. Differences between extracted entries were reconciled by consensus. When essential numerical information was missing or ambiguous, the corresponding author was contacted. An outcome was omitted from the relevant synthesis if usable data could not be obtained.
Standardization of decoction composition and preparation
For each intervention, two reviewers extracted the base formula, individual herbal components, reported dose, formula modifications, preparation method, administration frequency, and treatment duration. Herbal doses were converted to grams when direct conversion was possible; unclear doses or preparation procedures were recorded as not reported and were not imputed. Modified prescriptions were assigned to the same network node only when the original study explicitly identified the base formula and retained its principal components. The extracted preparations were considered members of the same formula family rather than pharmacologically identical products. Differences in dosage, processing, and modified compatibility were treated as potential sources of clinical heterogeneity.
Risk of bias
Two reviewers independently evaluated each included trial with Cochrane's revised risk-of-bias tool for randomized trials (RoB 2)19. The assessment covered bias arising from randomization, deviations from intended interventions, missing outcome data, outcome measurement, and selection of the reported result. RoB 2 decision rules were used to assign each domain—and the overall result—to low risk, some concerns, or high risk. Differences in judgment were reconciled through discussion.
GRADE assessment
For each outcome, certainty in the network estimates was rated under the grading of recommendations assessment, development, and evaluation (GRADE) framework20 and its extension for network meta-analysis32. Evidence from randomized trials started at high certainty and was downgraded when warranted for risk of bias, inconsistency, indirectness, imprecision, or publication bias. Final ratings were reported as high, moderate, low, or very low.
Statistical analysis
Analyses used post-treatment means and standard deviations; endpoint and change-from-baseline values were not pooled together. Effects were summarized as mean differences (MDs) with 95% credible intervals (CrIs). Units were harmonized to percentage points for HbA1c, g/24 h for 24-hour urinary protein, µmol/L for serum creatinine, and mmol/L for blood urea nitrogen, with other reported units converted by standard factors. When an SD was absent, it was derived from a standard error or 95% confidence interval if sufficient information was available; values were not borrowed from other studies. An outcome contribution was excluded when no usable dispersion measure could be recovered. An arm-level Bayesian network meta-analysis was fitted separately for each outcome. Continuous measurements were modeled with a normal likelihood and identity link within a random-effects consistency model. Basic treatment effects received diffuse normal priors with mean 0 and standard deviation 15 × om.scale, whereas the between-study heterogeneity standard deviation received a Uniform (0, om.scale) prior. All arms of multi-arm trials entered the model jointly; shared comparators were represented once so that correlations between contrasts were retained.
Four Markov chain Monte Carlo (MCMC) chains were run for each model, with 5,000 adaptation iterations and 20,000 subsequent draws per chain; the thinning interval was 1. Chain mixing and convergence were examined using trace plots, density plots, and the Brooks–Gelman–Rubin potential scale reduction factor (PSRF), for which a value below 1.05 was considered acceptable. Model adequacy was judged by comparing posterior residual deviance with the number of unconstrained data points and by contrasting deviance information criterion (DIC) values from consistency and inconsistency models. Bayesian models were fitted with the gemtc (version 1.1-1) package in R version 4.4.133, and Stata 15 was used for supplementary analyses. Ranking distributions were summarized as surface under the cumulative ranking curve (SUCRA) values. Exploratory univariable meta-regression analyses were performed separately for each outcome to assess country, treatment duration, total sample size, and publication year as potential sources of heterogeneity. Regression coefficients (β) and two-sided p values were reported. A p value of <0.05 was considered statistically significant. Because these analyses were exploratory, the findings were interpreted cautiously.
Subgroup and sensitivity analyses
Because substantial heterogeneity was observed in several direct comparisons, post hoc exploratory subgroup analyses were conducted using random-effects pairwise meta-analysis. Studies were stratified according to treatment duration (≤12 weeks vs >12 weeks), DKD stage (early-stage/stage III vs advanced-stage/stage IV or later), concomitant Western medical treatment (ACE inhibitor or angiotensin receptor blocker-based treatment vs other conventional treatment), and total sample size (<100 vs ≥100 participants). Subgroup differences were evaluated using an interaction test rather than by comparing statistical significance within individual subgroups. Analyses were performed only when at least two studies were available in each subgroup. Studies with mixed or insufficiently reported DKD stages or treatment regimens were classified as unclear and excluded from the corresponding interaction test.
Between-study heterogeneity was assessed using the I2 statistic. Potential sources of heterogeneity were explored using univariable meta-regression for country, treatment duration, total sample size, and publication year. Subgroup analyses according to treatment duration, diabetic kidney disease stage, baseline conventional treatment, and patient characteristics were considered when the relevant study-level information was sufficiently reported, and the resulting network remained connected. Sensitivity analysis excluding trials with a high overall risk of bias was prespecified. When an analysis could not be performed because of incomplete reporting, sparse strata, or network disconnection, the reason was reported, and the remaining uncertainty was considered in the interpretation.