Characteristics of included studies
Through a systematic electronic database search, 3,923 studies were initially identified for screening. After excluding 2,032 duplicate records and 1,719 irrelevant studies, 172 articles underwent full-text review. Based on the predefined inclusion and exclusion criteria, 17 studies13,14,17,18,19,20,21,22,23,24,25,26,27,28,29,30,31 were ultimately included in the meta-analysis (Figure 1).
The basic characteristics of the included studies are summarized in Table 1. These 17 studies comprised a total of 1,708 patients with HCC, of whom 962 had high Ki-67 expression and 746 had low expression. Various cutoff values for Ki-67 expression were applied across studies: 8 studies used 10%, whereas the remaining 9 used thresholds >10%, ranging from 14% to 50%; MRI was the most common source of radiomic features (n = 9), followed by ultrasonography (n = 5) and CT (n = 3). Ten studies used radiomic features alone for Ki-67 prediction, whereas the other 7 combined radiomic and clinical features to construct predictive models. In terms of modeling techniques, logistic regression was employed in 14 studies, and machine learning models, including support vector machine (SVM, n = 2) and Xception (n = 1), were used in 3 studies. Additionally, 5 studies were prospective in design, and 12 were retrospective.
Diagnostic value of radiomic features for predicting Ki-67 expression
A total of 17 studies evaluating the diagnostic accuracy of radiomic features in predicting Ki-67 expression levels in HCC were included. The random-effects meta-analysis showed a pooled sensitivity of 0.87 (95% CI: 0.81–0.91) and a pooled specificity of 0.79 (95% CI: 0.71–0.85) for predicting high Ki-67 expression. Significant heterogeneity was observed among studies (sensitivity: I2 = 80.81%; specificity: I2 = 86.86%), as shown in Figure 2. The Spearman correlation coefficient was −0.062 (P = 0.814), indicating no statistically significant threshold effect. The SROC analysis yielded an AUC of 0.90 (95% CI: 0.87–0.93), indicating high pooled diagnostic performance (Figure 3).
Ki-67 cutoff values and diagnostic performance
In the subgroup using a Ki-67 cutoff value of 10%, the radiomics-based diagnostic model demonstrated strong performance, with a pooled sensitivity of 0.87 (95% CI: 0.78–0.93), a specificity of 0.76 (95% CI: 0.60–0.87) and an AUC of 0.89 (95% CI: 0.86–0.92), indicating high diagnostic accuracy at this threshold (Figure 4 and Supplemental Figure S1). In the subgroup with Ki-67 cutoff values >10%, the model showed a pooled sensitivity of 0.87 (95% CI: 0.79–0.92), a specificity of 0.80 (95% CI: 0.71–0.86), and an AUC of 0.90 (95% CI: 0.87–0.92) (Figure 5 and Supplemental Figure S2).
Imaging modalities and diagnostic value
In the subgroup using MRI-derived radiomic features, the pooled sensitivity was 0.86 (95% CI: 0.77–0.92), the specificity was 0.77 (95% CI: 0.65–0.85), and the AUC was 0.88 (95% CI: 0.85–0.91), as shown in Figure 6 and Supplemental Figure S3. The subgroup based on ultrasound radiomics showed a sensitivity of 0.88 (95% CI: 0.74–0.95), a specificity of 0.81 (95% CI: 0.61–0.92), and an AUC of 0.92 (95% CI: 0.89–0.94) (Figure 7 and Supplemental Figure S4). Additionally, three studies used CT-based radiomic features20,25,30; however, due to the limited sample size, meta-analysis was not performed for this subgroup. These studies reported sensitivities ranging from 0.778 to 0.963, specificities from 0.75 to 0.877, and AUC values between 0.836 and 0.903 (Table 2).
Predictive models and diagnostic performance
In the subgroup employing logistic regression models, the pooled sensitivity was 0.86 (95% CI: 0.79–0.90), the specificity was 0.78 (95% CI: 0.68–0.85), and the AUC was 0.89 (95% CI: 0.86–0.92) (Figure 8 and Supplemental Figure S5). These estimates were similar to the overall results. Moreover, three studies used machine learning models: two employed SVM14,28, reporting AUC values of 0.94 and 0.986, sensitivities of 0.95 and 0.973, and specificities of 0.91 and 0.8397, respectively. One study used the Xception model24, with an AUC of 0.8, sensitivity of 0.76 and specificity of 0.78 (Table 2).
Risk of bias
The QUADAS-2 assessment indicated that most domains were judged to have a low risk of bias, with some high or unclear ratings in patient selection and flow and timing (Supplemental Figure S6 and Supplemental Figure S7). Deeks’ funnel plot asymmetry test showed no evidence of statistically significant publication bias (P = 0.42) (Supplemental Figure S8).
Overall, the pooled estimates indicated high diagnostic performance for radiomic prediction of Ki-67 expression, but the substantial between-study heterogeneity limits generalizability.
Data Availability
The study-level data extracted from the 17 included studies and used in this meta-analysis are provided in Supplemental File 2.

Figure 1: Flowchart of study selection. The flowchart shows the literature search, screening, eligibility assessment, and study inclusion process in accordance with PRISMA 2020. Abbreviation: PRISMA = Preferred Reporting Items for Systematic Reviews and Meta-Analyses. Please click here to view a larger version of this figure.

Figure 2: Sensitivity and specificity of radiomics for predicting Ki-67 expression in hepatocellular carcinoma. Paired forest plots show study-specific and pooled sensitivity and specificity estimates. Squares represent individual-study estimates, horizontal lines represent 95% CIs, and diamonds represent pooled estimates. Abbreviation: CI = confidence interval. Please click here to view a larger version of this figure.

Figure 3: Overall diagnostic performance of radiomics for predicting Ki-67 expression in hepatocellular carcinoma. The SROC curve shows the summary operating point (sensitivity, 0.87; specificity, 0.79), 95% confidence contour, and 95% prediction contour; the AUC was 0.90. Abbreviations: AUC = area under the curve; SROC = summary receiver operating characteristic. Please click here to view a larger version of this figure.

Figure 4: Diagnostic performance of the subgroup with a Ki-67 cutoff of 10%. Please click here to view a larger version of this figure.
The SROC curve shows the diagnostic performance of studies using a Ki-67 cutoff of 10%; the AUC was 0.89. Abbreviations: AUC = area under the curve; SROC = summary receiver operating characteristic.

Figure 5: Diagnostic performance of the subgroup with Ki-67 cutoffs > 10%. The SROC curve shows the diagnostic performance of studies using Ki-67 cutoffs > 10%; the AUC was 0.90. Abbreviations: AUC = area under the curve; SROC = summary receiver operating characteristic. Please click here to view a larger version of this figure.

Figure 6: Diagnostic performance of MRI-based radiomic models. The SROC curve shows the pooled diagnostic performance of MRI-derived radiomic features; the AUC was 0.88. Abbreviations: AUC = area under the curve; MRI = magnetic resonance imaging; SROC = summary receiver operating characteristic. Please click here to view a larger version of this figure.

Figure 7: Diagnostic performance of ultrasound-based radiomic models. The SROC curve shows the pooled diagnostic performance of ultrasound-derived radiomic features; the AUC was 0.92. Abbreviations: AUC = area under the curve; SROC = summary receiver operating characteristic. Please click here to view a larger version of this figure.

Figure 8: Diagnostic performance of logistic regression models. The SROC curve shows the pooled diagnostic performance of radiomic models constructed using logistic regression; the AUC was 0.89. Abbreviations: AUC = area under the curve; SROC = summary receiver operating characteristic. Please click here to view a larger version of this figure.
| Study | Study design | Sample size | Age | Male (%) | Imaging modality | Feature set | Prediction model | Segmentation | Validation | Ki-67 cutoff | Ki-67 high | Ki-67 low |
| Hu, 2017 | Retrospective | 57 | 54.23 ± 11.13 | 78.95 | MRI | Radiomics | Logistic regression model | Manual | No validation | 10% | 41 | 16 |
| Yao, 2018 | Retrospective | 44 | 55.5 ± 10.4 | 42.37 | Ultrasound | Radiomics | SVM | Manual | Internal validation | 25% | 23 | 21 |
| Chen, 2020 | Retrospective | 180 | 51.22 ± 10.38 | 82.8 | MRI | Radiomics | Logistic regression model | Manual | No validation | 50% | 34 | 146 |
| Ye, 2019 | Prospective | 89 | 50.72 ± 11.40 | 76.4 | MRI | Radiomics and clinical factors | Logistic regression model | Manual | Internal validation | 15% | 49 | 40 |
| Ye, 2020 | Prospective | 103 | 50.90 ± 11.93 | 77.67 | MRI | Radiomics | Logistic regression model | Manual | No validation | 10% | 73 | 30 |
| Wu, 2020 | Retrospective | 74 | 58.61 | 81.08 | CT | Radiomics | Logistic regression model | Manual | No validation | 10% | 54 | 20 |
| Shi, 2020 | Prospective | 52 | 55.7 ± 12.8 | 75. | MRI | Radiomics | Logistic regression model | Manual | No validation | 10% | 35 | 17 |
| Fan, 2021 | Retrospective | 103 | 61.0 (50.3–68.0) | 76.7 | MRI | Radiomics and clinical factors | Logistic regression model | Manual | Internal validation | 14% | 80 | 23 |
| Jing, 2021 | Retrospective | 81 | 53.52 | 76.54 | MRI | Radiomics | Logistic regression model | Manual | No validation | 10% | 67 | 14 |
| Hu, 2022 | Retrospective | 87 | 59.38 ± 11.13 | 88.51 | MRI | Radiomics | Xception | Manual | Internal validation | 20% | 40 | 47 |
| Wu, 2022 | Retrospective | 120 | 58.12 | 90. | CT | Radiomics and clinical factors | Logistic regression model | Manual | Internal validation | 20% | 63 | 57 |
| Dong, 2022 | Prospective | 60 | 59.35 ± 10.07 | 77.2 | Ultrasound | Radiomics | Logistic regression model | Manual | Internal validation | 10% | 37 | 23 |
| Liu, 2022 | Retrospective | 73 | >55 years: 71.% | 87.7 | MRI | Radiomics and clinical factors | Logistic regression model | Manual | Internal validation | 25% | 35 | 38 |
| Zhang, 2023 | Retrospective | 168 | 57.0 (49.0–64.0) | 81.5 | Ultrasound | Radiomics and clinical factors | SVM | Manual | Internal validation | 10% | 131 | 37 |
| Huang, 2022 | Prospective | 120 | 55.2 ± 11.2 | 92.5 | Ultrasound | Radiomics | Logistic regression model | Manual | No validation | 10% | 36 | 84 |
| Zhao, 2023 | Retrospective | 120 | 56.55 ± 9.53 | 87.5 | CT | Radiomics and clinical factors | Logistic regression model | Manual | Internal validation | 14% | 71 | 49 |
| Zhang, 2024 | Retrospective | 177 | 55.2 ± 11.4 | 86.4 | Ultrasound | Radiomics and clinical factors | Logistic regression model | Manual | Internal validation | 20% | 93 | 84 |
Table 1: Basic characteristics of the included studies. Characteristics of the 17 studies, including study design, sample size, imaging modality, feature set, prediction model, segmentation, validation, Ki-67 cutoff, and expression-group counts. Abbreviations: CT = computed tomography; MRI = magnetic resonance imaging; SVM = support vector machine.
| Subgroup | Reference | Sensitivity | Specificity | AUC |
| CT | 20 | 0.963 | 0.75 | 0.836 |
| CT | 25 | 0.778 | 0.877 | 0.884 (95% CI, 0.813–0.936) |
| CT | 30 | 0.86 | 0.79 | 0.903 (95% CI, 0.849–0.956) |
| Machine-learning model | 14 | 0.95 | 0.91 | 0.94 |
| Machine-learning model | 24 | 0.76 | 0.78 | 0.8 |
| Machine-learning model | 28 | 0.973 | 0.8397 | 0.986 (95% CI, 0.955–0.998) |
Table 2: Diagnostic performance of CT-based radiomic and advanced machine-learning models. Sensitivity, specificity, and AUC values reported by the individual CT-based and advanced machine-learning studies. Abbreviations: AUC = area under the curve; CI = confidence interval; CT = computed tomography.
Supplemental Figure S1: Sensitivity and specificity in the subgroup with a Ki-67 cutoff of 10%. Paired forest plots show study-specific and pooled estimates with 95% CIs. Abbreviation: CI = confidence interval. Please click here to download this file.
Supplemental Figure S2: Sensitivity and specificity in the subgroup with Ki-67 cutoffs > 10%. Paired forest plots show study-specific and pooled estimates with 95% CIs. Abbreviation: CI = confidence interval. Please click here to download this file.
Supplemental Figure S3: Sensitivity and specificity of MRI-based radiomic models. Paired forest plots show study-specific and pooled estimates with 95% CIs. Abbreviations: CI = confidence interval; MRI = magnetic resonance imaging. Please click here to download this file.
Supplemental Figure S4: Sensitivity and specificity of ultrasound-based radiomic models. Paired forest plots show study-specific and pooled estimates with 95% CIs. Abbreviation: CI = confidence interval. Please click here to download this file.
Supplemental Figure S5: Sensitivity and specificity of logistic regression models. Paired forest plots show study-specific and pooled estimates with 95% CIs. Abbreviation: CI = confidence interval. Please click here to download this file.
Supplemental Figure S6: Methodological quality graph. The graph shows the proportions of studies rated as having low, unclear, or high risk of bias and concerns regarding applicability in each QUADAS-2 domain. Abbreviation: QUADAS-2 = Quality Assessment of Diagnostic Accuracy Studies 2. Please click here to download this file.
Supplemental Figure S7: Methodological quality summary. The study-level summary shows low, unclear, or high ratings for risk of bias and concerns regarding applicability in each QUADAS-2 domain. Abbreviation: QUADAS-2 = Quality Assessment of Diagnostic Accuracy Studies 2. Please click here to download this file.
Supplemental Figure S8: Publication-bias assessment. Deeks’ funnel plot asymmetry test shows the relationship between the inverse square root of ESS and DOR; P = 0.42. Abbreviations: DOR = diagnostic odds ratio; ESS = effective sample size. Please click here to download this file.
Supplemental File 1: PRISMA 2020 checklist. The completed checklist documents reporting compliance for the systematic review and diagnostic meta-analysis. Please click here to download this file.
Supplemental File 2: Extracted data used for the diagnostic meta-analysis. The workbook contains the extracted study-level data, including patient characteristics, model details, and diagnostic performance values used for the pooled analyses. Please click here to download this file.