Clinical features of NAFLD
In this study, the clinical characteristics of individuals with and without NAFLD were systematically analyzed. A range of methods was employed to clearly illustrate the distribution and differences of these features within the sample population (Table 1). A total of 2,677 individuals were included in the analysis, of whom 718 had NAFLD. The analysis covered a range of clinical indicators, including BMI, TG, GGT, UA, and others (as shown in Figure 1). By performing PLS-DA on the clinical and metabolic variables, a significant separation was observed between the NAFLD and non-NAFLD groups (Figure 2A), indicating that these clinical features are crucial for differentiating NAFLD from non-NAFLD. Moreover, the heatmap (Figure 2B) demonstrated substantial deviations in multiple clinical and biochemical parameters, including WBC, NE, MO, RBC, UA, GGT, ALT, AST, TG, BMI, and GLU, within the NAFLD group. The increased frequency of these abnormalities in the NAFLD group further suggests potential involvement in the pathogenesis and progression of NAFLD. Further comparative analyses of BMI, TG, and GGT profiles between NAFLD and non-NAFLD cohorts demonstrated significantly elevated BMI levels in the NAFLD group compared with controls (P < 0.001) (Figure 2C). Moreover, the NAFLD group manifested pronounced increases in both TG and GGT concentrations (P = 3.44e-59 and P = 5.15e-52, respectively) (Figures 2D and E), substantiating the clinical utility of these biomarkers for NAFLD identification. Finally, the diagnostic accuracy of BMI, TG, and GGT in distinguishing NAFLD from non-NAFLD was assessed using ROC curve analysis. The AUC values were 0.79 (95% CI: 0.77–0.81) for BMI, 0.73 (95% CI: 0.71–0.75) for TG, and 0.71 (95% CI: 0.69–0.73) for GGT (Figures 2F–H, Supplementary Table 1). These results suggest that BMI exhibits the highest diagnostic performance among the three markers, while all three demonstrate moderate discriminatory ability.
Establishment of the NAFLD diagnostic model
Multiple machine learning algorithms were integrated with LASSO regression for feature selection to enhance model performance. Initially, variable selection was conducted using LASSO regression, where the relationship between the regularization parameter (Log Lambda) and the variable coefficients was illustrated (Figure 3A, Supplementary Table 2). As Lambda increased, the coefficients of less important variables were gradually shrunk to zero, resulting in the selection of 14 key variables: EOS, RBC, LDH, LYM, MON, BUN, Scr, UA, GGT, GLU, hypertension, diabetes mellitus, TG, and BMI. These variables are well documented as closely associated with the onset and progression of NAFLD and were therefore incorporated into the subsequent machine learning models. The cross-validation error curve for LASSO regression was also presented. The x-axis denotes the regularization parameter, log Lambda, while the y-axis indicates binomial deviance (Figure 3B). By selecting the optimal Lambda value (marked by the vertical line), the best-performing model was obtained with 14 variables and minimal deviance, thereby demonstrating strong predictive capability for NAFLD.
Variable importance was analyzed for the selected features. The results indicated that diabetes mellitus, TG, and BMI contributed the most to the model, with the highest importance scores (Figure 3C). Further, to evaluate the performance of different machine learning models, we employed a variety of algorithms. The performance of the models on the training, testing, and validation sets was illustrated in Figure 3D. By comparing the AUC values and F1 scores of the models, the GBM model exhibited superior performance compared to others across the training, testing, and validation sets, particularly in terms of AUC and F1 scores. Beyond discrimination metrics, the probabilistic reliability of the model was quantitatively evaluated to ensure its practicality for clinical screening. The 95% confidence intervals (CIs) for AUC, sensitivity, and specificity were calculated to rigorously account for the class imbalance present in the external validation cohort. In the external validation cohort, which exhibited a highly skewed distribution (159 NAFLD cases versus 41 non-NAFLD cases), the GBM model maintained strong discriminatory power, achieving an AUC of 0.851 (95% CI: 0.82–0.88). The corresponding sensitivity was 0.712, and the specificity was 0.835, indicating that the model still provides reliable risk stratification despite the imbalance in the validation sample. Furthermore, the ROC curves of the GBM model in both the training and testing sets were visualized (Figures 3E and F), demonstrating excellent discriminatory power and validating the robustness of the model. Notably, the confusion matrices for the GBM model in the training and testing sets were also presented (Figures 3G and H). In the training set, the model achieved a sensitivity of 0.712, specificity of 0.885, and accuracy of 0.745. In the testing set, the corresponding values were 0.712 for sensitivity, 0.835 for specificity, and 0.745 for accuracy. Although there was a minor reduction in specificity in the testing set, the overall performance of the model remained satisfactory.
Key features of the best diagnostic model
In this study, the GBM model was selected as the best diagnostic model for NAFLD based on its excellent performance in the training, testing, and validation sets. Next, an evaluation of the model's key features and the relative contribution of each variable to the prediction results was also conducted. An assessment of variable importance was conducted for the optimized GBM model (Figure 4A). The results indicated that BMI, TG, GGT, diabetes mellitus, and GLU were the most important variables, significantly contributing to the model's predictions. BMI, TG, and GGT had the highest importance scores, indicating a crucial role in predicting NAFLD.
To further understand the contribution of these variables to the model's output, the SHAP value distribution for each variable was displayed (Figure 4B). SHAP values quantify the size and direction of each variable's contribution to the model's predictions. It was observed that higher values of BMI, TG, and GGT pushed the model's predictions toward the NAFLD direction, while higher values of diabetes mellitus and GLU negatively impacted the prediction of NAFLD. Further single-sample SHAP analysis showed the contribution of SHAP values for a specific sample (Figure 4C), confirming that BMI, TG, and GGT played a positive role in predicting NAFLD, while diabetes mellitus and GLU had an inhibitory effect. This finding suggests that BMI, TG, and GGT are strong predictors of NAFLD, with important clinical significance.
To further validate the effect of these key variables on the prediction results, Partial dependence plots for BMI, TG, GGT, DM, and Glu were presented (Figure 4D–H). These plots show the relationship between each variable and the model's predicted response. The plots showed that as BMI, TG, and GGT levels increased, the probability of predicting NAFLD significantly increased, especially the increase in BMI and TG, which had the most pronounced positive impact on the prediction. Finally, SHAP decision path analysis revealed the SHAP value contributions for different samples. For positive prediction samples, the SHAP values showed that BMI, TG, and UA were the main variables contributing to the prediction toward NAFLD (Figure 4I). For negative prediction samples, BMI and GGT were the main negative contributing variables (Figure 4J). The analysis further confirmed the critical role of BMI and GGT in the diagnosis of NAFLD.
DATA AVAILABILITY:
The dataset for non-alcoholic fatty liver disease in the NHANES database is accessible via the following link: https://wwwn.cdc.gov/nchs/nhanes/continuousnhanes/default.aspx?BeginYear=2017.
All analytical codes and supplementary materials generated and utilized in this study are fully available and can be accessed at the following link: https://doi.org/10.6084/m9.figshare.32105248.

Figure 1: Schematic illustration of the current study. The inclusion flow, exclusion criteria, and the division of datasets into training, testing, and validation sets in this study. Please click here to view a larger version of this figure.

Figure 2: Comparison of clinical characteristics and diagnostic performance between NAFLD and non-NAFLD groups. (A) PLS-DA scatter plot: The Partial Least Squares Discriminant Analysis (PLS-DA) scatter plot based on clinical and metabolic variables shows significant separation between the NAFLD group (blue) and the non-NAFLD group (red). (B) Heatmap displaying the distribution differences of various clinical and biochemical indicators (including WBC, NEUT, MON, RBC, UA, GGT, ALT, AST, TG, BMI, and Glu) between the NAFLD (orange) and non-NAFLD (green) groups. Vertical axis: Group information (GROUP) and corresponding clinical indicators. Horizontal axis: Samples are arranged according to hierarchical clustering results, with similarities between samples represented by a dendrogram. Colors represent: Green for levels above or equal to the cutpoint, orange for levels below. (C) Bar chart comparing the BMI distribution between the NAFLD and non-NAFLD groups. (D) Bar chart comparing the triglyceride (TG) levels between the NAFLD and non-NAFLD groups. (E) Bar chart comparing the GGT (gamma-glutamyl transferase) levels between the NAFLD and non-NAFLD groups. (F) ROC curve to distinguish NAFLD from non-NAFLD based on BMI diagnostic performance. (G) ROC curve to distinguish NAFLD from non-NAFLD based on triglyceride (TG) diagnostic performance. (H) ROC curve to distinguish NAFLD from non-NAFLD based on GGT diagnostic performance. Please click here to view a larger version of this figure.

Figure 3: Variable selection and model evaluation for NAFLD diagnosis using machine learning models. (A) LASSO regression variable selection process: Variable selection based on LASSO regression (Least Absolute Shrinkage and Selection Operator). The horizontal axis represents the regularization parameter Log Lambda, and the vertical axis represents the variable regression coefficients. As Lambda increases, less important variables are gradually reduced to zero, ultimately selecting key variables. (B) Cross-validation error curve: LASSO regression cross-validation error curve. The horizontal axis represents the regularization parameter Log Lambda, and the vertical axis represents binomial deviance. The vertical line indicates the optimal Lambda value selected by the model and the corresponding number of variables. (C) Variable importance analysis: Importance scores of the features in the model, ranked from low to high. (D) Performance evaluation of multiple models: Performance of different models on the training, testing, and validation sets, including AUC values on the left and F1 scores on the right. The heatmap color gradient: From red to blue indicating performance differences, with deep red indicating higher performance scores. The green bar chart shows the average AUC and F1 scores for each model across the three datasets, with the bar length reflecting the overall performance level. (E) Training set ROC curve displaying the diagnostic performance of the GBM_grid_1_model_77 model in the training set. (F) Testing set ROC curve displaying the diagnostic performance of the GBM_grid_1_model_77 model in the testing set. (G) Training set confusion matrix: The confusion matrix of the predicted results for the GBM_grid_1_model_77 model in the training dataset, including accuracy indicators for the control and case groups. (H) Testing set confusion matrix: The confusion matrix of the predicted results for the GBM_grid_1_model_77 model in the testing dataset, including accuracy indicators for the control and case groups. Please click here to view a larger version of this figure.

Figure 4: GBM model variable importance and SHAP analysis results. (A) Variable importance ranking: Importance scores of the key variables in the GBM_grid_1_model_77 model. (B) SHAP value distribution plot: Distribution of the SHAP (Shapley Additive Explanations) values for the variables, showing the contribution size and direction of each variable to the model output. Purple represents lower variable values, while yellow represents higher variable values. (C) Single-sample SHAP analysis. It shows the contribution of SHAP values for a specific sample. (D–H) Partial dependence plots: Relationships between key variables (BMI, TG, GGT, DM, and Glu) and model prediction output. The horizontal axis represents the variable value, and the vertical axis represents the model's predicted response value, with the green line showing the marginal effect of the variable. (I) SHAP decision path analysis shows the SHAP value contributions of a positive prediction sample, where BMI, TG, and UA are the main positive contributing variables. (J) SHAP decision path analysis shows the SHAP value contributions of a negative prediction sample, where BMI and GGT are the main negative contributing variables.
| Characteristics | Sub-group | Overall | Non-NAFLD | NAFLD | P-value |
| Numbers | | 2677 | 1959 | 718 | / |
| Group | Non-NAFLD | 1959 (73.18) | 1959 (100.00) | 0 (0.00) | / |
| NAFLD | 718 (26.82) | 0 (0.00) | 718 (100.00) |
| Age,(years) | | 46.844 (17.781) | 46.845 (17.856) | 46.843 (17.586) | 0.9977 |
| CAP,(dB/m) | | 262.514 (61.791) | 233.617 (42.429) | 341.357 (28.768) | <0.0001 |
| Gender | Male | 1314 (49.08) | 890 (45.43) | 424 (59.05) | <0.0001 |
| Female | 1363 (50.92) | 1069 (54.57) | 294 (40.95) | |
| Hypertension | Yes | 861 (32.16) | 511 (26.08) | 350 (48.75) | <0.0001 |
| No | 1816 (67.84) | 1448(73.92) | 368 (51.25) | |
| Diabetes mellitus | Yes | 312 (11.65) | 140 (7.15) | 172 (23.96) | <0.0001 |
| No | 2365(88.35) | 1819(92.85) | 546 (76.04) | |
| Alb,(g/dL) | | 4.085 (0.322) | 4.100 (0.323) | 4.047 (0.315) | 0.0001 |
| Glb,(g/dL) | | 3.048 (0.409) | 3.040 (0.408) | 3.072 (0.409) | 0.0685 |
| TP,(g/dL) | | 7.138 (0.416) | 7.141 (0.418) | 7.131 (0.410) | 0.5634 |
| TC,(mg/dL) | | 187.803 (41.199) | 186.701 (40.351) | 190.809 (43.313) | 0.0223 |
| CO2,(mmol/L) | | 25.639 (2.478) | 25.685 (2.496) | 25.514 (2.423) | 0.1145 |
| Scr,(mg/dL) | | 0.878 (0.294) | 0.872 (0.280) | 0.897 (0.329) | 0.0535 |
| Glu,(mg/dL) | | 98.827 (28.174) | 94.588 (21.183) | 110.393 (39.422) | <0.0001 |
| TBil,(mg/dL) | | 0.469 (0.279) | 0.476 (0.291) | 0.450 (0.243) | 0.0309 |
| TG,(mg/dL) | | 141.224 (119.74) | 122.167 (103.103) | 193.219 (144.168) | <0.0001 |
| UA,(mg/dL) | | 5.432 (1.487) | 5.203 (1.383) | 6.058 (1.579) | <0.0001 |
| Na⁺,(mmol/L) | | 140.285 (2.704) | 140.263 (2.695) | 140.345 (2.729) | 0.4844 |
| K⁺,(mmol/L) | | 4.072 (0.356) | 4.063 (0.352) | 4.096 (0.366) | 0.0357 |
| Cl⁻,(mmol/L) | | 101.114 (2.758) | 101.187 (2.672) | 100.915 (2.971) | 0.0238 |
| BUN,(mg/dL) | | 14.595 (5.329) | 14.345 (5.227) | 15.276 (5.544) | 0.0001 |
| Ca,(mg/dL) | | 9.295 (0.355) | 9.299 (0.348) | 9.284 (0.374) | 0.3562 |
| AST,(u/L) | | 21.640 (12.418) | 20.753 (11.524) | 24.057 (14.315) | <0.0001 |
| ALT,(u/L) | | 22.419 (16.867) | 19.793 (13.967) | 29.585 (21.420) | <0.0001 |
| LDH,(IU/L) | | 155.782 (35.255) | 154.602 (36.464) | 159.001 (31.523) | 0.0042 |
| GGT,(IU/L) | | 30.581 (41.784) | 26.952 (39.685) | 40.481 (45.620) | <0.0001 |
| BMI,(kg/㎡) | | 29.572 (7.243) | 27.714 (6.322) | 34.643 (7.172) | <0.0001 |
| WBC,(1000 cells/uL) | | 7.243 (2.496) | 7.002 (2.562) | 7.900 (2.174) | <0.0001 |
| RBC,(million cells/uL) | | 4.767 (0.493) | 4.719 (0.480) | 4.898 (0.502) | <0.0001 |
| NEUT,(1000 cells/uL) | | 4.181 (1.696) | 4.019 (1.668) | 4.620 (1.693) | <0.0001 |
| LYM,(1000 cells/uL) | | 2.257 (1.438) | 2.207 (1.612) | 2.393 (0.771) | 0.0029 |
| MON,(1000 cells/uL) | | 0.574 (0.226) | 0.558 (0.227) | 0.619 (0.215) | <0.0001 |
| EOS,(1000 cells/uL) | | 0.198 (0.157) | 0.189 (0.152) | 0.223 (0.167) | <0.0001 |
| PLT,(1000 cells/uL) | | 244.401 (61.515) | 243.137 (61.239) | 247.850 (62.175) | 0.0791 |
| RDW | | 13.754 (1.234) | 13.706 (1.231) | 13.885 (1.230) | 0.0009 |
Table 1: Clinical characteristics at baseline of individuals with NAFLD or non-NAFLD. This table presents the clinical relevance scores for the variables initially selected by the LASSO model. Notes: P < 0.05 was considered statistically significant.
Supplementary Table 1: The AUC values and optimal cut-off points for each clinical variable. Please click here to download this file.
Supplementary Table 2: Results of assigning importance scores to the variables selected by LASSO using ChatGPT-4. ChatGPT-4 was utilized to evaluate and score these variables based on existing medical knowledge and literature. These AI-screened features were subsequently validated through the AutoML modeling process. Scores range from 1 to 10, with higher scores indicating greater clinical relevance.Please click here to download this file.