$$\rightleftharpoonup{xx}$$
$$\longleftharp{xx}$$,
$$\longrightharp{xx}$$,
Patient enrollment and CIE determination
A total of 161 patients who underwent neurointerventional procedures were included in this study, among whom 12 (7.5%) developed CIE, while 149 (92.5%) did not. The detailed patient screening and grouping process is shown in Figure 2. The diagnosis of CIE was established based on the temporal association between symptom onset and neurointerventional procedures, combined with exclusionary imaging findings. Representative neuroimaging findings of CIE are presented in Figure 3A–D.
The baseline demographic, clinical, laboratory, and procedural characteristics of the two groups are summarized in Table 1. No significant differences were observed between the CIE and non-CIE groups in terms of age, sex distribution, or body weight (all p > 0.05). Similarly, the prevalence of common comorbidities, including hypertension, diabetes mellitus, and coronary heart disease, did not differ significantly between the two groups (all p > 0.05). Lifestyle factors, such as alcohol consumption and smoking status, were also comparable.
In contrast, several laboratory and procedure-related variables showed significant differences between the two groups. Patients in the CIE group had significantly higher serum creatinine levels compared with those in the non-CIE group (median: 85.2 vs 65.3 µmol/L, p = 0.004), while their eGFR was significantly lower (median: 81 vs 102.7 mL/min/1.73 m2, p < 0.001). Regarding procedural characteristics, embolization procedures were significantly more frequent in the CIE group than in the non-CIE group (41.7% vs 2.0%, p < 0.001). In addition, posterior circulation lesions were significantly more common among patients who developed CIE (58.3% vs 6.0%, p < 0.001). No significant difference was observed in the type of contrast media used between the two groups (p = 0.982). Furthermore, patients in the CIE group received a significantly higher volume of contrast agent during the procedure (median: 200.0 vs 165.0 mL, p = 0.005) and had longer procedure durations (median: 78.5 vs 50.0 min, p = 0.008). Notably, the CGR, a composite index proposed in this study, was significantly higher in the CIE group compared with the non-CIE group (median: 2.35 vs 1.63, p < 0.001), indicating a strong association between CGR and the occurrence of CIE.
Perioperative variable acquisition and CGR construction
Perioperative demographic, laboratory, and procedural variables were systematically collected and analyzed as described in the protocol. Among these variables, CGR, defined as the ratio of total contrast volume to eGFR, was constructed as a composite indicator reflecting the balance between contrast-agent exposure and renal clearance capacity. CGR was associated with CIE occurrence in this cohort and was included as a candidate predictor for feature-selection and model-development analyses. To compare CGR with its individual components, we performed a comparative ROC analysis using CGR, total contrast volume, and eGFR. CGR showed the highest discriminative ability for CIE, with an AUC of 0.959 (95% CI: 0.924–0.985), compared with total contrast volume (AUC = 0.745, 95% CI: 0.559–0.886) and eGFR analyzed in the inverse risk direction (AUC = 0.865, 95% CI: 0.764–0.951). The AUC difference between CGR and total contrast volume was 0.214, and the AUC difference between CGR and eGFR was 0.093. Detailed results are provided in Supplementary Table 4.
Feature selection and ML workflow
To identify the most relevant predictors while minimizing multicollinearity, LASSO regression with 10-fold cross-validation was applied to the candidate variables. The coefficient profile and cross-validation curves of the LASSO model are presented in Figure 4A and Figure 4B, respectively. The LASSO coefficient ranking of the final model is presented in Figure 5. At the optimal λ value, a subset of variables was retained, including CGR, procedure type, lesion location, triglyceride levels, serum creatinine, eGFR, and several clinical covariates. The quantitative ranking of retained predictors based on LASSO coefficients is provided in Supplementary Table 5.
Comparisons of model performance
Five ML approaches—Naive Bayes, SVM, KNN, LightGBM, and MLP—were developed to predict the risk of CIE. Their performance in the training and test sets is summarized in Table 2, the accuracy comparison of the five models is shown in Figure 6, and the corresponding confusion matrices are provided in Supplementary Figure 1, Supplementary Figure 2, Supplementary Figure 3, Supplementary Figure 4, and Supplementary Figure 5. Because the dataset was markedly imbalanced, with 12 CIE cases and 149 non-CIE cases, accuracy and threshold-based metrics should be interpreted cautiously and in conjunction with sensitivity, specificity, AUC, and confidence intervals. In the training set, all models exhibited good discriminative ability, with AUC values ranging from 0.961–0.999. In the internal test set, the Naive Bayes model showed the highest numerical (AUC = 0.952, 95% CI: 0.843–1.000), followed by LightGBM (AUC = 0.944, 95% CI: 0.822–1.000) and SVM (AUC = 0.935, 95% CI: 0.796–1.000).
At the optimal classification threshold determined by the Youden index, the Naive Bayes model yielded a sensitivity of 100% and a specificity of 90.3%. Because formal pairwise statistical comparisons among model AUCs were not performed, the observed differences among models should be interpreted as descriptive and exploratory. In addition, because the dataset was markedly imbalanced and the internal test cohort was small with only a very limited number of CIE events, threshold-based metrics such as sensitivity and specificity may be unstable and should be interpreted cautiously. The receiver operating characteristic (ROC) curves of the Naive Bayes model in the training and test sets are presented in Figure 7.
Although the MLP model performed well in the training set (AUC = 0.961), its performance declined substantially in the test set (AUC = 0.742), suggesting potential overfitting. Similarly, the KNN model demonstrated limited discriminative ability in the test set (AUC = 0.718) with a wide confidence interval, indicating instability. The Youden index–based strategy was used to determine the optimal classification threshold for threshold-dependent metrics, including sensitivity, specificity, and accuracy. The ROC-AUC values themselves are threshold-independent. Therefore, the reported threshold-dependent performance metrics reflect model behavior at the optimal operating point. The optimal Youden thresholds and corresponding threshold-based metrics are provided in Supplementary Table 6. Overall, the Naive Bayes model showed relatively stable descriptive performance between the training and internal test sets, with minimal numerical decline in AUC.
Decision curve analysis
DCA was performed to evaluate the clinical utility of the models across a range of threshold probabilities. The DCA curves are shown in Figure 8. All models demonstrated positive net benefit over specific ranges of threshold probabilities compared with the treat-all and treat-none strategies. Among the evaluated models, SVM demonstrated the widest range of positive net benefit, whereas the other models also showed potential clinical usefulness within specific threshold ranges.
DATA AVAILABILITY:
The de-identified raw data used to support the main analyses have been uploaded to GitHub and are available at: https://doi.org/10.5281/zenodo.20043331. All direct patient identifiers were removed before data sharing. Because this was a retrospective clinical study based on hospital medical records, access to any additional patient-level information remains restricted by institutional ethical and privacy requirements. Additional de-identified data may be made available from the corresponding author upon reasonable request and with approval from the institutional ethics committee. The statistical analysis scripts, ML model training code, and figure-generation code have been provided as supplementary files to support reproducibility.

Figure 1: Workflow of CIE risk modeling based on perioperative data. (A) Mechanisms of CIE and derivation of the contrast volume-to-eGFR ratio (CGR). (B) Data integration and LASSO-based feature selection. (C) ML model development and comparison, with Naive Bayes showing the highest numerical AUC in the internal test set (AUC ≈ 0.95). Abbreviations: CIE = contrast-induced encephalopathy; CGR = contrast volume-to-eGFR ratio; eGFR = estimated glomerular filtration rate; AUC = area under the curve. The figure was derived from the authors' own clinical data, and the schematic illustration was originally created by the authors using modified Microsoft PowerPoint icons. Please click here to view a larger version of this figure.

Figure 2: Patient selection and cohort allocation. Flowchart illustrating patient screening, eligibility assessment, exclusion criteria, and cohort assignment. A total of 225 patients were initially assessed. After applying inclusion criteria, including age between 30 and 80 years, availability of complete clinical data, and postoperative head CT within 3 days, 190 patients were eligible. Patients were excluded based on acute cerebral infarction, chronic kidney disease stage ≥ G4 (eGFR < 30 mL/min/1.73 m2), severe cardiovascular disease, absence of endovascular treatment, severe neurological deficit (modified Rankin Scale score ≥ 4), or incomplete data. Finally, 161 patients were included and randomly divided into the training cohort (n = 128) and internal test cohort (n = 33) at an 8:2 ratio. Abbreviations: CT = computed tomography; eGFR = estimated glomerular filtration rate; CKD = chronic kidney disease; mRS = modified Rankin Scale. Please click here to view a larger version of this figure.

Figure 3: Serial non-contrast CT images in a patient with CIE. (A) Baseline CT before procedure. (B) CT at symptom onset showing cortical hyper density in the right parietal lobe (arrow). (C) CT on postoperative day one showing partial resolution. (D) CT on postoperative day two showing further resolution. Please click here to view a larger version of this figure.

Figure 4: LASSO-based feature selection. (A) Ten-fold cross-validation curve used to determine the optimal λ based on the minimum MSE. (B) Coefficient profiles of candidate variables plotted against log(λ), showing the changes in variable coefficients with increasing regularization strength. Abbreviations: LASSO = least absolute shrinkage and selection operator; λ = regularization parameter; MSE = mean squared error. Please click here to view a larger version of this figure.

Figure 5: Coefficients of selected variables. Bar plot showing coefficients of variables retained after LASSO selection. Positive coefficients indicate positive association; negative coefficients indicate inverse association. Please click here to view a larger version of this figure.

Figure 6: Training and test accuracy of the five ML models. Accuracy of five ML models (Naive Bayes, SVM, KNN, LightGBM, MLP) in training and test sets. Abbreviations: SVM = support vector machine; KNN = k-nearest neighbor; LightGBM = light gradient boosting machine; MLP = multilayer perceptron. Please click here to view a larger version of this figure.

Figure 7: ROC curves of the Naive Bayes model. Receiver operating characteristic curves for the training and internal test cohorts with corresponding area under the curve values. Abbreviations: ROC = receiver operating characteristic; AUC = area under the curve. Please click here to view a larger version of this figure.

Figure 8: Decision curve analysis. DCA showing net benefit of different models across threshold probabilities compared with treat-all and treat-none strategies. Shaded regions indicate threshold probability ranges where the model achieved positive net benefit compared with both treat-all and treat-none strategies. Abbreviations: DCA = decision curve analysis. Please click here to view a larger version of this figure.
Supplementary Figure 1: Confusion matrix of the Naive Bayes model in the internal test cohort. The matrix displays the numbers of correctly and incorrectly classified samples at the selected classification threshold.Please click here to download this file.
Supplementary Figure 2: Confusion matrix of the SVM model in the internal test cohort. The matrix displays the numbers of correctly and incorrectly classified samples at the selected classification threshold.Please click here to download this file.
Supplementary Figure 3: Confusion matrix of the KNN model in the internal test cohort. The matrix displays the numbers of correctly and incorrectly classified samples at the selected classification threshold.Please click here to download this file.
Supplementary Figure 4: Confusion matrix of the LightGBM model in the internal test cohort. The matrix displays the numbers of correctly and incorrectly classified samples at the selected classification threshold.Please click here to download this file.
Supplementary Figure 5: Confusion matrix of the MLP model in the internal test cohort. The matrix displays the numbers of correctly and incorrectly classified samples at the selected classification threshold.Please click here to download this file.
| Characteristic | CIE ( n = 12 ) | Non-CIE ( n = 149 ) | p value |
| Demographics | | | |
| Age (years), median (IQR) | 67 (62.5–69.0 ) | 63 ( 55.0–70.0 ) | 0.194 |
| Male, n (%) | 9 (75.0) | 105 (70.5) | 1.000 |
| Weight (kg), median (IQR) | 65 (55.0–70.0) | 70 (60.0–77.5) | 0.160 |
| Comorbidities, n (%) | | | |
| Hypertension | 7 (58.3) | 102 (68.5) | 0.526 |
| Diabetes | 4 (33.0) | 57 (38.3) | 1.000 |
| Coronary heart disease | 2 (16.7) | 23 (15.4) | 1.000 |
| Habits, n (%) | | | |
| Alcohol Use | 7 (58.3) | 78 (52.3) | 0.770 |
| Smoking Status | 7 (58.3) | 59 (39.6) | 0.233 |
| Lab Values, median (IQR) | | | |
| Serum Creatinine (μmol/L) | 85.2 (65.0–89.0) | 65.3 (54.9–75.7) | 0.004 |
| eGFR (mL/min/1.73m²) | 81.0 (60.6–90.3) | 102.7 (89.0–116.3) | <0.001 |
| Total Cholesterol (mmol/L) | 3.3 (3.0–4.4) | 3.4 (2.9–4.1) | 0.775 |
| Triglycerides (mmol/L) | 1.8 (1.2–2.1) | 1.2 (0.8–1.6) | 0.017 |
| Procedural Characteristics | | | |
| Procedure Type, n ( % ) | | | <0.001 |
| - Stent Angioplasty | 7 (58.3) | 146(98.0) | |
| - Embolization | 5 (41.7) | 3 (2.0) | |
| Lesion Location, n (%) | | | <0.001 |
| - Posterior Circulation | 7 (58.3) | 9 (6.0) | |
| - Anterior Circulation | 5 (41.7) | 140 (94.0) | |
| Contrast Media Type, n (%) | | | 0.982 |
| - 2nd Generation | 6 (50.0) | 75 (50.3) | |
| - 3rd Generation | 6(50.0) | 74 (49.7) | |
| Total Contrast Volume (mL), median (IQR) | 200 (188.0–200.0) | 165 (155.0–180.0) | 0.005 |
| Procedure Duration (min), median (IQR) | 78.5 (52.5–124.5) | 50 (47.0–57.0) | 0.008 |
| Novel Index | | | |
| CGR, median (IQR) | 2.35 (2.17–2.70) | 1.63 (1.40–1.85) | <0.001 |
Table 1: Baseline clinical characteristics of patients. Demographic, clinical, laboratory, and procedure-related characteristics of patients were compared between the CIE and non-CIE groups. Continuous variables are presented as mean ± standard deviation or median (interquartile range), and categorical variables are presented as number and percentage. Statistical comparisons between groups were performed using appropriate parametric or nonparametric tests. Abbreviations: CIE = contrast-induced encephalopathy; eGFR = estimated glomerular filtration rate; CGR = contrast volume-to-eGFR ratio; IQR = interquartile range; SD = standard deviation.
| Model name | Accuracy | AUC | 95% CI | Sensitivity | Specificity | Dataset |
| Naive Bayes | 0.875 | 0.961 | 0.924–0.998 | 1.0 | 0.864 | train |
| Naive Bayes | 0.909 | 0.952 | 0.843–1.000 | 1.0 | 0.903 | test |
| SVM | 0.992 | 0.999 | 0.997–1.000 | 1.0 | 0.992 | train |
| SVM | 0.879 | 0.935 | 0.796–1.000 | 1.0 | 0.871 | test |
| KNN | 0.891 | 0.969 | 0.941–0.997 | 1.0 | 0.881 | train |
| KNN | 0.939 | 0.718 | 0.194–1.000 | 0.5 | 0.968 | test |
| LightGBM | 0.969 | 0.988 | 0.973–1.000 | 1.0 | 0.966 | train |
| LightGBM | 0.848 | 0.944 | 0.822–1.000 | 1.0 | 0.839 | test |
| MLP | 0.914 | 0.961 | 0.908–1.000 | 0.9 | 0.915 | train |
| MLP | 0.970 | 0.742 | 0.228–1.000 | 0.5 | 1.000 | test |
Table 2: Comparison of the performance of the models. The predictive performance of different ML models was evaluated using accuracy, area under the curve (AUC), 95% confidence interval (CI), sensitivity, and specificity in the training and internal test datasets. Abbreviations: CI = confidence interval.
Supplementary Table 1: Summary of missing data assessment before model development. Summary of missing values for all candidate predictor variables before model development. No missing values were observed among the included variables; therefore, no imputation procedure was performed.Please click here to download this file.
Supplementary Table 2: Variable dictionary of the analytical dataset. Definitions, coding methods, variable categories, and descriptions of clinical, laboratory, and procedure-related variables included in the ML analysis.Please click here to download this file.
Supplementary Table 3: Hyperparameter search space and selected parameters for ML models. Candidate hyperparameter ranges and optimized parameters obtained during model tuning for each ML algorithm are summarized.Please click here to download this file.
Supplementary Table 4: Comparative ROC analysis of CGR and its individual components. Receiver operating characteristic analysis was performed to compare the predictive performance of CGR and its individual components for CIE prediction.Please click here to download this file.
Supplementary Table 5: Quantitative feature ranking based on LASSO coefficients. Selected variables were ranked according to their coefficient values after LASSO regression, showing the relative contribution and direction of association of each variable in the predictive model.Please click here to download this file.
Supplementary Table 6: Youden thresholds determined in the training cohort and corresponding classification performance metrics. The optimal threshold for each model was determined exclusively in the training cohort using the Youden index and was subsequently fixed and applied to the internal test cohort. Accuracy, sensitivity, and specificity in the internal test cohort were calculated using these fixed thresholds.Please click here to download this file.