This retrospective study was reviewed and approved by the Medical Ethics Committee of Baoquanling Hospital of Beidahuang Group on May 9, 2025 (Approval No. BQH-BDHG-EC-2025-056), and informed consent was waived. The study was conducted in accordance with the Declaration of Helsinki and the hospital's data privacy management policy, and all personally identifiable information was de-identified before analysis. The research tools used in this study are listed in the Table of Materials.
1. Study design
This was a single-center retrospective cross-sectional study conducted in the Department of Endocrinology and Metabolism of Baoquanling Hospital of Beidahuang Group, including 126 eligible patients with type 2 diabetes who attended the hospital between January 1, 2022 and December 31, 2024. The date of the outpatient visit or hospital admission was defined as the index date, and laboratory test results obtained on that date were used as the baseline. Exposure and outcome measurements were obtained at the same time point or within 7 days before or after the index date. All data were derived from historical medical records without any intervention.
The report was prepared in accordance with the Strengthening the Reporting of Observational Studies in Epidemiology (STROBE)11 statement for cross-sectional studies, and the research question, variables, statistical methods, and sensitivity analysis strategies were prespecified, with results presented according to the prespecified plan.
2. Study population
The study population was screened consecutively from the electronic medical record system according to the prespecified inclusion and exclusion criteria. The inclusion criteria were age 18–80 years, a documented diagnosis of type 2 diabetes in the medical record, available results for fasting plasma glucose, fasting insulin, hs-CRP, urine albumin, urine creatinine, and serum creatinine within 7 days before or after the index date, eGFR ≥ 60 mL/min/1.73 m2 (CKD-EPI 2021), a freshly collected urine sample processed according to the standard testing procedure, and complete clinical data and key covariates or a missing proportion that met the prespecified handling criteria.
The exclusion criteria included evidence of acute inflammation or infection, hs-CRP > 10 mg/L, urinary tract infection, hematuria or pyuria, pregnancy or lactation, a definite history of non-diabetic kidney disease or imaging evidence of structural kidney disease, UACR ≥ 300 mg/g, systemic glucocorticoid therapy within the previous 3 months, acute cardiovascular or cerebrovascular events or major surgery within the previous 3 months, and malignancy receiving chemotherapy or immunotherapy.
After records that did not meet the criteria were excluded, the final sample was obtained. Under a two-sided α = 0.05 and 80% power, the final sample size of 126 in this study corresponded to a detectable absolute correlation coefficient of approximately 0.25. This statement was only a post hoc description of the detectable effect range and did not constitute an a priori sample size estimation. Analyses involving HOMA-IR calculations were restricted to participants not receiving exogenous insulin therapy, and the sample size of this subgroup was reported as observed.
3. Data sources and collection procedures
According to a unified data dictionary, the investigators extracted demographic information (age, sex), diabetes duration, height, weight, smoking status, alcohol use, systolic blood pressure (SBP), and diastolic blood pressure (DBP) (the second reading after two measurements on the index date), comorbidity records, and medication information from the electronic medical records.
Medication variables included renin-angiotensin system inhibitors (RASi), sodium-glucose cotransporter 2 inhibitors (SGLT2i), glucagon-like peptide-1 receptor agonists (GLP-1RA), and statins; continuous use for ≥ 3 months before the index date was recorded as "use". Diabetes duration was defined as the number of years from diagnosis to the index date. Body mass index (BMI) was calculated as weight (kg)/height2 (m2).
Laboratory testing was performed on a unified platform and was subject to internal quality control and external quality assessment. hs-CRP was measured by high-sensitivity immunoturbidimetry, with a detection limit of ≤ 0.1 mg/L; FINS was measured by chemiluminescence with within-batch calibration using calibrators; FPG was measured by the hexokinase method; glycated hemoglobin (HbA1c) was measured by high-performance liquid chromatography; urine albumin was measured by immunoturbidimetry, urine creatinine by the enzymatic method, and the ratio was expressed as mg/g; serum creatinine was measured by the enzymatic method with IDMS-traceable calibration.
All samples were tested within 2 h after collection or after short-term storage at 4 °C. If multiple results were available for the same visit, results obtained on the same day as the index date were prioritized.
4. Variable definitions and measurement
Exposure variables included hs-CRP and HOMA-IR. hs-CRP was expressed in mg/L, entered into the primary analyses as a continuous variable, and categorized into sample tertiles for trend analyses. FINS was expressed in µU/mL, FPG in mmol/L, and the calculation formula for HOMA-IR was as follows6:
HOMA - IR = (FINS FPG)/22.5
Analyses involving HOMA-IR were restricted to participants not using exogenous insulin who had same-day FPG and FINS results. In this subgroup, HOMA-IR and hs-CRP were entered into the multivariable model together. UACR was calculated using urine albumin and urine creatinine measured in the same sample and was expressed as mg/g12. Urine creatinine was harmonized to grams for ratio calculation when required. In this dataset, no UACR value was zero. For urine albumin values below the lower detection limit (<2.0 mg/L), the laboratory reported the result as <2.0 mg/L, and these below‑detection values were replaced with half of the lower detection limit (1.0 mg/L) prior to data analysis. To reduce the influence of right skewness, a natural log transformation was applied, and ln(UACR) was used as the primary outcome.
If multiple urine tests were available for the same participant within 7 days before or after the index date, only the sample closest to the index date, labeled as morning urine or first-void urine, was retained. If only a random urine sample was available, the sample type was recorded, and sensitivity analyses were restricted to the morning urine subset. Microalbuminuria was defined as UACR ≥ 30 mg/g and was used as a binary surrogate outcome in logistic regression.
eGFR was calculated using the CKD-EPI 2021 creatinine equation13. When serum creatinine was reported in µmol/L, it was converted to mg/dL (µmol/L ÷ 88.4) before calculation. eGFR was expressed as mL/min/1.73 m2 and used as a continuous secondary outcome in sensitivity analyses. Consistency of the results was examined in the eGFR 60–89 mL/min/1.73 m2 subgroup. The primary model was prespecified to adjust for age, sex, diabetes duration, SBP, and HbA1c, and the sensitivity models additionally included BMI, smoking status, alcohol use, use of RASi, use of SGLT2i, use of GLP-1RA, and statin use.
5. Data management and preprocessing
After de-identification, data were exported as the analytical dataset, and a brief data dictionary summarizing variable names, definitions, units, and coding rules for the analytical variables was provided in Supplementary Table 1. Duplicate records were merged by index date, and key variables were checked for logical consistency.
When both EMR and LIS timestamps were available, the LIS specimen collection time was used as the primary timestamp for temporal alignment; if unavailable, the LIS report/verification time was used, whereas the EMR visit date was used only to define the index date. When multiple eligible results were available within the prespecified window, the result closest to the index date was retained; if two results were equally close, the same-day result was prioritized, with urine samples further selected according to the prespecified sample-type rule.
Missing values were handled according to the prespecified hierarchical strategy: if the missing proportion of any single covariate was ≤ 10%, the primary analyses used a complete-case approach; if it exceeded 10%, multiple imputation by chained equations with 20 imputations was performed, with the imputation model including the exposures, outcome, and all covariates, and the imputed results were compared with the complete-case results.
Participants with hs-CRP > 10 mg/L or evidence of acute inflammation were excluded from the primary analyses. Continuous variables were assessed for distribution using the Shapiro-Wilk test and Q-Q plots, and UACR was natural log-transformed. If hs-CRP or HOMA-IR showed marked skewness, rank or log transformations were performed in sensitivity analyses. Categorical variables were coded as binary or ordinal variables according to prespecified rules. Influential observations were identified by absolute studentized residuals > 3 or Cook's distance > 4/n, and the primary models were repeated after excluding these observations in sensitivity analyses.
Prespecified subgroup analyses included sex subgroups and subgroups based on HbA1c < 7% and ≥ 7%. Prespecified sensitivity analyses included restriction to morning urine samples, addition of medication variables to the primary models, exclusion of influential observations, use of robust standard errors instead of conventional standard errors, and restriction to participants with eGFR 60–89 mL/min/1.73 m2.
6. Statistical analysis
All statistical analyses were performed in R version 4.3.2. Continuous variables were assessed for distribution using the Shapiro-Wilk test and Q-Q plots. Normally distributed data were expressed as mean ± standard deviation, non-normally distributed data as median (interquartile range), and categorical variables as frequency and percentage.
Baseline characteristics were described according to whether the microalbuminuria threshold was reached, and between-group comparisons were performed using the independent-samples t test, Mann-Whitney U test, chi-square test, or Fisher's exact test, according to variable type and distribution. Spearman rank correlation analysis was performed between ln-UACR and hs-CRP and between ln-UACR and HOMA-IR, and correlation coefficients and their 95% confidence intervals were calculated, with confidence intervals estimated by Fisher's z transformation. Correlation plots were presented as scatter plots overlaid with locally weighted regression fitted lines to show the trend.
Multivariable linear regression models were constructed with ln-UACR as the dependent variable. In the subgroup not using exogenous insulin, hs-CRP and HOMA-IR were entered into the model together, with adjustment for age, sex, diabetes duration, SBP, and HbA1c; in the overall sample, a separate model including only hs-CRP was fitted to examine the overall correlation. Standardized regression coefficients, 95% confidence intervals, and p values were reported, and the change in the coefficient of determination before and after inclusion of the exposure variables was reported. Collinearity was assessed using variance inflation factors, with a threshold of 5.
After hs-CRP and HOMA-IR were categorized into sample tertiles, defined by the empirical 33.3rd and 66.7th percentiles of the corresponding analytic sample (hs-CRP: the overall sample for overall analyses and the non-insulin subgroup for subgroup/joint analyses; HOMA-IR: the non-insulin subgroup only), with T1 ≤ the lower cutoff, T2 > the lower cutoff to ≤ the upper cutoff, and T3 > the upper cutoff, the median value of each tertile was entered as a continuous variable to test for linear trend, and the adjusted marginal means of ln-UACR across tertiles and the p value for trend were reported.
In the exploratory analysis, high hs-CRP and high HOMA-IR were defined by the upper tertile cutoffs, derived from the non-insulin subgroup for the 2 × 2 joint analysis, and a 2 × 2 grouping was constructed to compare the adjusted marginal means of ln-UACR across groups. An interaction term was added to test for statistical interaction, and the adjusted mean difference between the "high × high" and "low × low" groups was reported; this analysis was not interpreted causally.
For surrogate outcome analysis, multivariable logistic regression was constructed with reaching the microalbuminuria threshold as the dependent variable, and the odds ratio and 95% confidence interval corresponding to each 1 standard deviation increase in hs-CRP or HOMA-IR were reported.
Multivariable linear regression was constructed with eGFR as the dependent variable to examine the direction of the association between the exposures and glomerular filtration rate.
Sensitivity analyses included restriction to morning urine samples, additional adjustment for BMI, smoking status, alcohol use, and the four medication categories in the primary models, exclusion of influential observations, use of Huber-White robust standard errors, and repetition of the primary models in the restricted population with UACR < 300 mg/g and eGFR ≥ 60 mL/min/1.73 m2. Model diagnostics used residual-versus-fitted plots, Q-Q plots, and the Shapiro-Wilk test to assess residual normality and linearity, and the Breusch-Pagan test to assess homoscedasticity. If heteroscedasticity was present, robust standard errors were reported; if clear nonlinearity was found, tertile indicator variables were used in place of continuous variables in sensitivity analyses.
Correlation and regression analyses were performed using the stats package, collinearity was assessed using the car package, robust standard errors were obtained using the sandwich and lmtest packages, multiple imputation was performed using the mice package, and figures were generated using the ggplot2 package. All tests were two-sided, the significance threshold was set at α = 0.05, and 95% confidence intervals were reported.