Research Article

Deep Learning of Lateral Thoracolumbar Radiographs and Clinical Risk Factors for Incident Vertebral Fracture: Single-Center Retrospective Cohort Study

39 views

DOI:

10.3791/71628

August 18th, 2026

In This Article

Summary

A deep learning score extracted from thoracolumbar lateral radiographs, combined with clinical risk factors, enabled accurate prediction of incident vertebral fractures within two years. The internally validated model showed better discrimination, calibration, reclassification, and decision benefit than the clinical model, supporting individualized risk stratification and early preventive management strategies.

Abstract

Early identification of patients at risk of incident vertebral fracture remains challenging because routine clinical risk assessment does not fully capture local spinal fragility. This single-center retrospective cohort study evaluated whether deep learning (DL) features extracted from baseline thoracolumbar lateral radiographs improve the prediction of incident vertebral fracture within 2 years when combined with clinical risk factors. A total of 2,173 patients were included and chronologically divided into a derivation cohort (n = 1,449) and an internal validation cohort (n = 724). DL features were derived from baseline radiographs, and LASSO-Cox regression was used to select predictors and build a clinical model, a DL model, and a combined model. Performance was assessed by bootstrap optimism correction, temporal internal validation, calibration, decision curve analysis, time-dependent net reclassification improvement (NRI), integrated discrimination improvement (IDI), and sensitivity analyses. Of 2,048 candidate DL features, 5 were retained to generate a DL score, which remained an independent predictor in the combined model (HR 1.64, 95% CI 1.34–2.01; P < 0.001). In internal validation, the combined model achieved a C-index of 0.759, a 2-year AUC of 0.774, and a 2-year Brier score of 0.077, all superior to the clinical model, with good calibration (intercept 0.012; slope 0.972). Compared with the clinical model, the combined model also improved reclassification (2-year NRI 0.316 in derivation and 0.241 in validation) and discrimination (2-year IDI 0.047 and 0.033, respectively; all P < 0.01), and provided greater net benefit on decision curve analysis. Sensitivity analyses were consistent with the primary results. Combining DL features from thoracolumbar lateral radiographs with clinical risk factors may enable more accurate individualized prediction of incident vertebral fracture within 2 years.

Introduction

Vertebral fracture is one of the most common types of osteoporotic fragility fractures, and is particularly common in the thoracolumbar region. It can lead to chronic pain, loss of height, kyphotic deformity, limited mobility, and increase the risk of refracture and adverse prognosis1. In clinical practice, a considerable proportion of patients lack typical symptoms before fracture occurrence, and many cases are identified only on follow-up imaging, suggesting that relying solely on symptoms or retrospective diagnosis makes it difficult to complete timely screening of high-risk populations2,3. Existing risk assessment mainly relies on information such as age, sex, body mass index, previous fragility fracture, diabetes, glucocorticoid exposure, and bone mineral density, which can reflect the background of systemic bone fragility, but it is difficult to fully characterize the local structural fragility and mechanical abnormalities of the thoracolumbar spine, and this is also a key difficulty that has long existed in the prediction of new vertebral fracture risk4. Thoracolumbar lateral radiography is one of the most commonly used and accessible spinal imaging examinations in clinical practice. It can not only show vertebral morphology, but may also contain occult phenotypes related to future fracture, such as endplate changes, sparse bone texture, mild wedging, and imbalance of alignment5. Previous studies have mostly focused on detecting existing vertebral fractures, diagnosing osteoporosis, or assessing risk using manual measurement indicators6,7. Recent evidence has further shown that deep learning-identified prevalent vertebral fracture and osteoporosis on lateral spine imaging, together with clinical risk factors, can improve prediction of incident fracture5; however, evidence remains limited for predicting incident vertebral fracture specifically in patients without target vertebral fracture at baseline using routine thoracolumbar lateral radiographs and local deep learning (DL) features. Artificial intelligence methods have been used for spinal imaging analysis, but studies directly targeting this specific clinical scenario remain limited, and systematic evaluation of calibration, net benefit from decision analysis, and temporal split validation is still insufficient in this setting8.

Therefore, it is difficult to answer a more clinically relevant question: can the features extracted by deep learning from routine thoracolumbar lateral X-ray provide independent and meaningful incremental information on the basis of clinical risk assessment9? Based on the above background, this study adopted a single-center retrospective cohort design, extracted deep learning features from thoracolumbar lateral X-ray, and combined them with clinical risk factors to construct a risk prediction model for incident vertebral fracture within 2 years, and evaluated the discrimination, calibration, robustness, and clinical value of the model through temporal internal validation, bootstrap optimism correction, and sensitivity analysis. This study focused on individualized risk warning under routine X-ray, integrating occult local imaging fragility phenotypes and systemic clinical susceptibility information into an interpretable prediction tool, so as to provide a basis for high-risk identification, intensified follow-up, and preventive intervention.

Protocol

This study was reviewed and approved by the Medical Ethics Committee of Shanghai Eighth People’s Hospital, Shanghai, China (approval number 2026-102-03-02). Because this study was a retrospective study and all data had been de-identified before analysis, the ethics committee waived informed consent from patients.

Study design:

Study type

This study was a single-center retrospective cohort study, and the study database was established using data from the hospital's picture archiving and communication system (PACS), radiology information system (RIS), and electronic medical record system. The study population comprised consecutive patients who underwent thoracolumbar lateral digital X-ray examination at the hospital. The inclusion period extended from January 1, 2018, to December 31, 2023, and the follow-up deadline was December 31, 2025. The study report followed the TRIPOD+AI and STROBE recommendations to ensure the reporting standardization of prediction model studies involving artificial intelligence and observational studies.

Study setting and case source

Cases were derived from the routine clinical diagnosis and treatment process of outpatients, emergency patients, and inpatients in the hospital. Imaging data were all derived from original DICOM files in PACS, and clinical data were derived from structured electronic medical records, laboratory systems, and prescription records. The date of the first thoracolumbar lateral X-ray examination that met the inclusion criteria during the study period was defined as the baseline date; when the same patient had multiple examinations meeting the criteria, only the earliest one was retained as the baseline examination to avoid repeated enrollment. All data were de-identified before analysis, and imaging and clinical information were matched using a unique study identification number.

Study population:

Inclusion criteria

The inclusion criteria were as follows: age 50 years or older; completion of standard standing thoracolumbar lateral digital X-ray examination in the hospital during the study period; baseline imaging in traceable DICOM format; complete visualization of the T10 to L4 vertebrae on baseline imaging; no existing vertebral fracture from T10 to L4 on review of baseline imaging; extractable prespecified baseline clinical variables from electronic medical records; at least 1 follow-up thoracolumbar X-ray, CT, or MRI examination within 24 months after baseline, or occurrence of an imaging-confirmed new vertebral fracture within 24 months.

Exclusion criteria

The exclusion criteria were as follows: vertebral fracture from T10 to L4 at baseline; a definite history of high-energy violent injury at baseline or during follow-up; primary or metastatic spinal tumor, spinal infection, or destructive bone disease; previous thoracolumbar internal fixation surgery, vertebroplasty, or kyphoplasty; scoliosis with a Cobb angle greater than 30° or obvious kyphotic deformity (including Scheuermann-type kyphotic deformity, if present) resulting in inability to accurately identify the endplates from T10 to L4; obvious motion artifact, abnormal exposure, metal occlusion, or insufficient display range on imaging; inability to confirm key baseline variables or outcome information from electronic medical records.

Retrospective cohort construction process

Study population screening was independently completed by two researchers according to the prespecified criteria, and disagreements were resolved by discussion to reach consensus. After case screening was completed, time-series grouping was performed according to the baseline date: patients enrolled from January 1, 2018, to December 31, 2021, constituted the derivation cohort for feature selection and model construction; patients enrolled from January 1, 2022, to December 31, 2023, constituted the internal validation cohort for model performance evaluation. Time splitting rather than random splitting can reduce the risk of information leakage and is closer to the real application scenario of the model in subsequent patients. The study population screening process is presented in the form of a flowchart.

Primary outcome and its determination:

Definition of the primary outcome

The primary outcome of this study was the first incident fragility vertebral fracture from T10 to L4 within 24 months after baseline. The prediction time window of the study was prespecified as 2 years, and the model output was the individual risk probability of incident vertebral fracture within 2 years.

Criteria for the determination of an incident vertebral fracture

Incident vertebral fracture was defined as follows: relative to baseline imaging, follow-up imaging showed a decrease of 20% or more in the anterior, middle, or posterior height of any vertebral body from T10 to L4, with an absolute height reduction of at least 4 mm, or the appearance of new endplate collapse or cortical interruption10. Outcome determination was made comprehensively based on follow-up thoracolumbar X-ray, CT, and MRI. Image reading was performed independently by 2 musculoskeletal radiologists, with 8 years and 12 years of relevant diagnostic experience, respectively, and neither had access to clinical data or model output results during image reading; if disagreement occurred, adjudication was made by 1 senior musculoskeletal radiologist with 18 years of experience. Vertebral fractures caused by tumor, infection, or high-energy violence were not counted as outcome events.

Follow-up start point, end point, and observation window

The follow-up start point was the date of the baseline thoracolumbar lateral X-ray examination. The follow-up endpoint was defined as the earliest of the following time points: the date of the first incident vertebral fracture, 24 months after baseline, the date of the last spinal imaging examination confirming no vertebral fracture, or the date of death. Fractures first appearing after 24 months were not included in the primary outcome. Patients without outcome events were treated as censored.

Collection of clinical data and definition of candidate clinical variables:

Demographic and general clinical data

Baseline clinical data were extracted from the electronic medical record system by two researchers according to a unified case report form, without reviewing the outcome determination results during extraction. The collected demographic and general clinical data included age, sex, height, weight, and body mass index. Age was defined as the actual age on the baseline date; weight and height were taken from the record closest to the baseline date within 30 days before or after the baseline date; body mass index was calculated as weight divided by height squared, in kilograms per square meter.

Medical history, medication use, and bone metabolism-related data

Based on clinical availability and model generalizability, the following candidate clinical risk factors were prespecified for inclusion: previous fragility fracture history, type 2 diabetes mellitus, rheumatoid arthritis, chronic oral glucocorticoid use, and baseline anti-osteoporosis treatment. Standardized baseline bone mineral density measurements and the FRAX score were not prespecified candidate predictors because they were not uniformly available as standardized baseline variables across the whole cohort; several FRAX-related clinical factors were instead considered separately as individual candidate variables. Previous fragility fracture history, diagnosis of underlying diseases, and medication information were all derived from electronic medical records, discharge records, and prescription systems before baseline, and all variables were required to have existed before baseline to ensure that the predictors temporally preceded the outcome event.

Definition criteria for clinical variables

Previous fragility fracture history was defined as a fracture occurring after the age of 40 years, caused by low-energy injury, and clearly recorded in the medical record; skull, facial bone, finger bone, and toe bone fractures were not included in this definition. Type 2 diabetes mellitus was defined as a clear diagnosis recorded before baseline, or long-term use of hypoglycemic drugs. Rheumatoid arthritis was defined as a clear diagnosis made by a rheumatology specialist in the medical record. Chronic oral glucocorticoid use was defined as a prednisone equivalent dose of not less than 5 mg/d for not less than 3 months within 1 year before baseline. Baseline anti-osteoporosis treatment was defined as continuous use of any of bisphosphonates, denosumab, teriparatide, raloxifene, calcitonin, alfacalcidol, or calcitriol within 3 months before baseline, for a duration of not less than 8 weeks. Age and body mass index were treated as continuous variables in modeling and were not artificially categorized.

Imaging data acquisition and image preprocessing

Thoracolumbar lateral X-ray acquisition protocol

All baseline images were standard standing thoracolumbar lateral X-rays acquired by the hospital's digital radiography system. During examination, patients assumed a natural standing position, with both upper limbs flexed forward to reduce shoulder overlap, and the imaging range covered T10 to L4. Automatic exposure control was used for examination, with a tube voltage range of 80–95 kV and a source-to-image distance of 110 cm. For the same patient, when multiple eligible lateral radiographs were available on the baseline date, the one with a complete display range and the best image quality was selected as the analysis object.

Image inclusion criteria and quality control

Baseline images were required to meet the following quality requirements: complete visualization of the T10 to L4 vertebrae and their upper and lower endplates; clear vertebral anterior and posterior margins, endplates, and cortical boundaries; no obvious motion artifact; no severe overexposure or underexposure; no large-area metal occlusion; and no obvious morphological distortion caused by body position rotation. Images with severe degenerative change or osteophytes that precluded reliable identification of vertebral margins or endplates were also excluded. Two musculoskeletal radiologists performed quality review of all baseline images, and any image failing to meet any key quality criterion was excluded.

Image preprocessing and standardization

All DICOM images were anonymized before analysis. The preprocessing steps included unifying image orientation, resampling to a spatial resolution of 0.30 mm × 0.30 mm, truncating grayscale values between the 0.5th percentile and the 99.5th percentile, and standardizing pixel values to the 0–1 interval using the min-max normalization method. The above preprocessing workflow was kept consistent in the derivation cohort and the validation cohort, and was all automatically completed by prespecified scripts to reduce bias caused by manual operations.

Deep learning imaging feature extraction:

Region of interest determination

The region of interest was the lateral projection region of the spine between the upper endplate of T10 and the lower endplate of L4. One musculoskeletal radiologist with 8 years of experience completed rectangular box annotation of all baseline images in ITK-SNAP software, with the anterior boundary set 5 mm anterior to the vertebral anterior margin and the posterior boundary set 5 mm posterior to the vertebral posterior margin11; another musculoskeletal radiologist with 12 years of experience reviewed the images case by case. The ROI was a region-level rectangular box rather than a strict vertebral contour segmentation; therefore, common marginal osteophytes were not separately removed and could be partially included if they fell within the prespecified boundary, whereas cases with degenerative change severe enough to obscure the vertebral margins or endplates had already been excluded during image quality review. To evaluate the reproducibility of region annotation, 50 images were randomly selected and re-annotated by the same radiologist after 4 weeks, and independently re-annotated by the second radiologist, for subsequent feature stability analysis. After ROI cropping, all images were uniformly resized to 224 × 224 pixels.

Deep learning model architecture and feature extraction process

This study used the ResNet50 convolutional neural network as the deep learning feature extractor. The network parameters were initialized with ImageNet pretrained weights, and self-supervised domain adaptation was performed on all baseline ROI images in the derivation cohort, without using outcome labels during the adaptation process. Specifically, a contrastive self-supervised task was used, in which two independently augmented views generated from the same ROI image were treated as a positive pair, whereas views from different patients within the same mini-batch were treated as negative pairs, so that the encoder could adapt to the distribution of the study images. Model training used the AdamW optimizer, with an initial learning rate set at 1 × 10^-4, a batch size of 64, and 200 training epochs; during training, data augmentation was performed with ±5° rotation, 0.9–1.1-fold scaling, translation of no more than 10 pixels, and contrast perturbation of ±10%12. These augmentations were used to generate paired views for the self-supervised task, and only unlabeled images from the derivation cohort were used at this stage. After domain adaptation, no outcome-supervised fine-tuning was performed, and the adapted backbone encoder was fixed for feature extraction. After completion of domain adaptation, the 2,048-dimensional vector output from the global average pooling layer was extracted as the candidate deep learning features for each patient.

Imaging feature screening and dimensionality reduction

First, the intraclass correlation coefficient of features was calculated based on the 50 images with repeated annotation, and features with both intraobserver and interobserver ICC not lower than 0.80 were retained to ensure the stability of features against slight ROI variation. Subsequently, the retained features were Z-score standardized in the derivation cohort, zero-variance features were removed, and for features with an absolute pairwise correlation coefficient greater than 0.90, only one of them was retained. Finally, LASSO-Cox regression was used for feature selection, and the penalty parameter was determined by 10-fold cross-validation according to the 1-SE criterion. Features with non-zero regression coefficients were weighted and summed according to their coefficients to construct the deep learning score (DL score)13. After this scoring formula was determined in the derivation cohort, it was fixed unchanged and directly applied to the internal validation cohort.

Preprocessing and integration of candidate predictors:

Missing data handling and data standardization

All candidate clinical variables were obtained from structured medical record fields. Variables with a missing rate exceeding 20% were excluded from the modeling process. The remaining missing values were handled using multiple imputation by chained equations, generating 10 imputed datasets; the imputation model incorporated all candidate predictors, the outcome indicator variable, and the Nelson-Aalen cumulative hazard estimate to preserve time-to-event outcome information as much as possible. Continuous clinical variables and the DL score were standardized using the mean and standard deviation of the derivation cohort, and the same transformation parameters were applied to the validation cohort; binary variables were uniformly coded as 0 or 1.

Clinical risk factor selection

The prespecification of candidate clinical risk factors was based on clinical interpretability, previous evidence, and data availability, and univariable P-value screening was not used. The candidate clinical variables entered into LASSO-Cox selection were age, sex, body mass index, previous fragility fracture history, type 2 diabetes mellitus, rheumatoid arthritis, chronic oral glucocorticoid use, and baseline anti-osteoporosis treatment; height and weight were collected descriptively and used to derive body mass index, but were not entered separately into modeling. LASSO-Cox regression was performed separately in the 10 imputed datasets of the derivation cohort, and the penalty parameter was selected using 10-fold cross-validation; variables with non-zero coefficients in at least 7 imputed datasets entered the final clinical model. Age and body mass index were both tested for nonlinear relationships using restricted cubic splines; if the nonlinear term was not statistically significant, the linear form was retained. Multicollinearity was evaluated by the variance inflation factor, and variables with a variance inflation factor greater than 5 were not retained simultaneously.

Construction of the combined predictor set

To avoid overfitting caused by directly entering high-dimensional imaging features into the model, the deep learning information was first compressed into a single continuous variable, the DL score, and then jointly entered into combined modeling together with the selected clinical risk factors. No interaction terms were prespecified in the combined model, so as to maintain model parsimony and interpretability. The final combined predictor set consisted of the DL score and the retained clinical variables.

Risk prediction model construction:

Modeling strategy

In the derivation cohort, the clinical model, the deep learning model, and the combined model were established separately. The models used Cox proportional hazards regression, with the first incident fragility vertebral fracture within 24 months after baseline as the study endpoint, and the censoring rules are described in the above follow-up definition. To control overfitting, the complexity of the combined model was restricted before modeling, and a relatively high event-per-parameter ratio was maintained as much as possible. The final regression coefficients and standard errors of each model were estimated separately in the 10 imputed datasets and then pooled using Rubin's rules. The baseline hazard function was estimated according to the Breslow method, and the individual 2-year risk probability was calculated.

Clinical model construction

The clinical model included the clinical risk factors retained after LASSO-Cox selection. All continuous variables were kept in continuous form and were not dichotomized. After model fitting, the proportional hazards assumption was tested using Schoenfeld residuals; for variables not satisfying the proportional hazards assumption, an interaction term with ln(time) was added for correction. The clinical model was used to characterize the predictive ability of traditional clinical information for incident vertebral fracture.

Construction of the deep learning imaging model

The deep learning model was established as a Cox proportional hazards model using the DL score as the only predictor, to quantify the predictive ability of deep learning features from baseline thoracolumbar lateral X-ray for the risk of incident vertebral fracture within 2 years. This model did not introduce any clinical information, and thus served as an imaging unimodal model for comparison with the other models.

Combined model construction

The combined model further added the DL score on the basis of the clinical model, and constructed a comprehensive prediction model based on deep learning features from thoracolumbar lateral X-ray combined with clinical risk factors. After the combined model was established, a 2-year risk nomogram was drawn according to its regression coefficients for individualized risk estimation and clinical application display.

Internal validation and performance evaluation of the model:

Internal validation method

Internal validation adopted a temporally separated single-center internal validation strategy. All models established in the derivation cohort were directly applied to the validation cohort enrolled from January 1, 2022, to December 31, 2023, after parameters were fixed, without refitting. In addition, 1,000 bootstrap resamples were performed within the derivation cohort to obtain optimism-corrected performance estimates, so as to evaluate model stability.

Discrimination evaluation

Model discrimination was evaluated by the Harrell concordance index and the 2-year time-dependent AUC calculated based on the inverse probability of censoring weighting method, both with 95% confidence intervals reported. Higher discrimination indicates that the model is better able to distinguish individuals who will and will not develop incident vertebral fractures in the future. Differences in discrimination between models were calculated using the bootstrap method with 95% confidence intervals.

Calibration evaluation

Model calibration was evaluated using the 2-year risk calibration curve, calibration intercept, calibration slope, and 2-year Brier score. The calibration curve was plotted based on deciles of predicted risk and was bootstrap-corrected. A calibration intercept close to 0, a calibration slope close to 1, and a lower Brier score indicate good agreement between the predicted risk and the actually observed risk.

Evaluation of clinical application value

The clinical application value of the model was evaluated by a 2-year decision curve analysis, comparing the net benefit under different threshold probabilities. The threshold probability range was prespecified as 0.05–0.30 to cover the risk interval that may be used clinically for intensified follow-up, further bone assessment, or intervention management14. A model with a higher net benefit was considered to have better clinical decision support value.

Model comparison and determination of the best model

The clinical model, deep learning model, and combined model were comprehensively compared by discrimination, calibration, Brier score, and decision curve. The gain of the combined model relative to the clinical model was further quantified using the 2-year time-dependent net reclassification improvement and integrated discrimination improvement. The best model was prespecified as the model that simultaneously had higher discrimination, good calibration, lower prediction error, and greater net benefit.

Statistical Analysis:

Continuous variables were initially evaluated for distribution pattern using the Shapiro-Wilk test; those conforming to a normal distribution were presented as mean ± standard deviation, whereas those with a skewed distribution were reported as median and interquartile range; categorical variables were presented as number of cases and percentage. Comparisons of baseline characteristics between the derivation cohort and the validation cohort were performed using the independent-samples t test, Mann-Whitney U test, χ2 test, or Fisher's exact test, respectively. Baseline comparisons were used only to describe cohort characteristics and were not used as a basis for variable selection. All statistical tests were two-sided, and P < 0.05 was considered statistically significant. Statistical analyses were completed in R software, mainly using the survival, glmnet, mice, rms, timeROC, and rmda packages; image preprocessing and deep learning analysis were completed in the Python and PyTorch environment. To evaluate the robustness of the results, a complete-case analysis was additionally performed as a sensitivity analysis.

Results

Retrospective cohort construction process and baseline characteristics of the cohorts

During the study period, thoracolumbar lateral X-ray records were retrieved, and 6,114 patients were included for screening after deduplication. After stepwise exclusion of patients aged < 50 years, those with existing fractures at baseline, and those with insufficient follow-up, a total of 2,173 patients were finally included, including 1,449 in the derivation cohort and 724 in the internal validation cohort (Figure 1). The baseline characteristic distributions of the derivation cohort and the internal validation cohort were generally balanced, and there were no statistically significant differences in age, sex, body mass index, or major clinical risk factors (all P > 0.05). The median follow-up time in the two cohorts was 23.4 months and 23.1 months, respectively; there were 131 and 63 incident vertebral fractures events, respectively; and the 2-year cumulative incidence was 9.21% and 8.91%, respectively, with no statistically significant difference (P = 0.812) (Table 1).

Clinical risk factor selection, imaging feature screening, and risk prediction model construction

After LASSO-Cox selection, age, female sex, body mass index, previous fragility fracture history, type 2 diabetes mellitus, and chronic oral glucocorticoid use reached the prespecified inclusion frequency threshold; after stepwise screening of 2048 deep learning features, 5 features with non-zero coefficients were retained at λ1se to construct the DL score (Figure 2A–C). Based on the selected clinical variables and the DL score, the clinical model, deep learning model, and combined model were further established. Multivariable Cox regression showed that the above clinical variables were all associated with the risk of incident vertebral fracture within 2 years (all P < 0.05), and after the DL score was added to the clinical model, it remained an independent predictor in the combined model (HR = 1.64, 95% CI 1.34–2.01, P < 0.001) (Table 2). Accordingly, a nomogram of the combined model was drawn for individualized estimation of the 2-year risk of incident vertebral fracture; the higher the total score, the higher the predicted risk (Figure 2D).

Internal validation and performance evaluation of the model

After bootstrap optimism correction in the derivation cohort, the combined model still maintained the best predictive performance. Internal validation showed that the C-index and AUC₂y of the combined model were 0.759 and 0.774, respectively, both higher than those of the clinical model; its Brier₂y was the lowest (0.077), and the calibration intercept was close to 0, and the calibration slope was close to 1, indicating that this model had good discrimination and calibration (Table 3). In the derivation cohort, both the apparent calibration curve and the bootstrap bias-corrected curve were close to the ideal line. In the internal validation cohort, the predicted 2-year risk was generally consistent with the Kaplan-Meier observed risk, and the decile calibration points were distributed near the ideal line, indicating that the combined model had good 2-year risk calibration (Figure 3A, B).

Model comparison and evaluation of clinical application value

Compared with the clinical model, the combined model achieved significant net reclassification improvement and discrimination improvement in both the derivation cohort and the internal validation cohort, with NRI₂y values of 0.316 and 0.241, respectively, and IDI₂y values of 0.047 and 0.033, respectively (all P < 0.01) (Table 4). In the derivation cohort and the internal validation cohort, the combined model generally achieved the highest net benefit within the prespecified threshold probability range of 0.05 – 0.30, and its decision curve was mostly above Treat-all and Treat-none, indicating that it had better clinical application value (Figure 4A, B).

Sensitivity analysis results

The complete-case sensitivity analysis showed that the conclusions of the primary analysis remained basically stable. In both the derivation cohort and the internal validation cohort, the C-index and AUC₂y of the combined model were higher than those of the clinical model, and the Brier₂y was lower; its calibration intercept and calibration slope in the internal validation cohort were 0.019 and 0.964, respectively, suggesting that the model had good robustness (Table 5). During follow-up, 27 deaths in the derivation cohort and 13 deaths in the internal validation cohort were recorded. In the Fine–Gray competing-risk sensitivity analysis treating death as a competing event, the DL score remained independently associated with incident vertebral fracture in the combined model (subdistribution HR = 1.58, 95% CI 1.28–1.95, P < 0.001), and the overall conclusions were unchanged.

In summary, the combined model integrating the deep learning score from baseline thoracolumbar lateral radiographs with selected clinical risk factors showed the best overall performance for predicting incident vertebral fracture within 2 years. Compared with the clinical model, it demonstrated higher discrimination, better calibration, lower prediction error, improved reclassification, and greater net benefit in both the derivation and internal validation cohorts. The independent predictive value of the deep learning score and the consistency of findings in complete-case and competing-risk sensitivity analyses further supported the robustness of the main results.

DATA AVAILABILITY:

The raw data have been uploaded as Supplementary file 1.

Lateral thoracolumbar radiograph study flowchart: patient inclusion and exclusion criteria pathway.
Figure 1. Flowchart of study population screening. For the same patient, when multiple examinations met the eligibility criteria, only the earliest one was retained as the baseline examination. Each exclusion reason was applied sequentially according to the prespecified order, and each patient was counted only once for exclusion. Please click here to view a larger version of this figure.

Risk assessment diagram with bar graph, LASSO plots, nomogram for fracture prediction modeling.
Figure 2. LASSO-Cox selection of clinical risk factors, deep learning features, and a nomogram of the combined model. (A) Inclusion frequency of candidate clinical variables in 10 imputed datasets, with the dashed line indicating the 70% threshold. (B) LASSO-Cox coefficient paths of deep learning features. (C) Partial likelihood deviance curve from 10-fold cross-validation, with the vertical dashed lines indicating λmin and λ1se, respectively. (D) Nomogram for 2-year risk in the combined model, each predictor corresponds to a certain number of points, and the points are summed to obtain the total score, which is further converted into the individual 2-year risk of incident vertebral fracture. DL score, deep learning score. Please click here to view a larger version of this figure.

Calibration curves comparing predicted vs observed 2-year vertebral fracture risk, derivation and validation cohorts, data analysis chart.
Figure 3. 2-year risk calibration curves of the combined model in the derivation cohort and the internal validation cohort. (A) Derivation cohort. (B) Internal validation cohort. The calibration points were generated according to deciles of predicted risk, and the observed risk was estimated using the Kaplan-Meier method. Please click here to view a larger version of this figure.

Threshold probability vs net benefit graph; derivation, validation cohorts; model comparison.
Figure 4. Decision curve analysis of the three models in the derivation cohort and the internal validation cohort. (A) Derivation cohort. (B) Internal validation cohort. The horizontal axis represents threshold probability, and the vertical axis represents net benefit. Treat-all indicates intervention for all, and Treat-none indicates intervention for none. Please click here to view a larger version of this figure.

Variable NameMissing Values, n (%)Derivation Cohort (n=1449)Internal Validation Cohort (n=724)P
Baseline characteristics
Sample size, n1449724
Age, years0 (0.00)68.41 ± 8.3768.96 ± 8.560.155
Female, n (%)0 (0.00)962 (66.39%)463 (63.95%)0.259
Height, cm16 (0.74)158.42 ± 7.91157.98 ± 8.160.232
Weight, kg21 (0.97)59.76 ± 9.8859.21 ± 10.140.23
Body mass index, kg/m²28 (1.29)23.77 ± 3.2823.69 ± 3.340.597
Previous fragility fracture history, n (%)0 (0.00)171 (11.80%)96 (13.26%)0.329
Type 2 diabetes mellitus, n (%)0 (0.00)303 (20.91%)158 (21.82%)0.624
Rheumatoid arthritis, n (%)0 (0.00)49 (3.38%)29 (4.01%)0.461
Chronic oral glucocorticoid use, n (%)0 (0.00)65 (4.49%)38 (5.25%)0.43
Baseline anti-osteoporosis treatment, n (%)0 (0.00)131 (9.04%)75 (10.36%)0.323
Follow-up and outcome description
Follow-up time, months0 (0.00)23.4 [18.7, 24.0]23.1 [18.4, 24.0]0.341
Number of incident vertebral fracture events, n0 (0.00)13163
2-year cumulative incidence of incident vertebral fracture, % (95% CI)9.21 (7.82, 10.60)8.91 (6.79, 11.03)0.812

Table 1: Baseline characteristics and outcomes of the two cohorts. The missing values column was based on the original observed data, and multiple imputation was used only for modeling. Continuous variables are presented as x̄ ± s or M[IQR] according to distribution, and between-group comparisons were performed using the independent-samples t test or Mann-Whitney U test; categorical variables are presented as n (%), and between-group comparisons were performed using the χ2 test. The 2-year cumulative incidence of incident vertebral fracture was estimated by the Kaplan-Meier method and reported with 95% CI; between-group comparison was performed using the log-rank test. P values were used only to describe differences in cohort composition between the two cohorts and were not used for predictor selection.

PredictorβHR95% CIP
Clinical model
Age (per 1 SD increase)0.281.331.10–1.600.003
Female (yes vs no)0.261.291.02–1.630.031
Body mass index (per 1 SD increase)−0.190.830.70–0.980.03
Previous fragility fracture history (yes vs no)0.661.931.38–2.71<0.001
Type 2 diabetes mellitus (yes vs no)0.311.361.06–1.750.016
Chronic oral glucocorticoid use (yes vs no)0.491.631.14–2.330.008
Deep learning model
DL score (per 1 SD increase)0.581.781.46–2.17<0.001
Combined model
Age (per 1 SD increase)0.221.251.07–1.460.004
Female (yes vs no)0.231.261.01–1.560.04
Body mass index (per 1 SD increase)−0.180.840.72–0.980.031
Previous fragility fracture history (yes vs no)0.591.81.27–2.560.001
Type 2 diabetes mellitus (yes vs no)0.271.311.01–1.700.044
Chronic oral glucocorticoid use (yes vs no)0.421.531.05–2.210.026
DL score (per 1 SD increase)0.51.641.34–2.01<0.001

Table 2: Predictors and Cox regression results of the three models. Parameter estimates of the clinical model and the combined model were pooled from 10 imputed datasets according to Rubin's rules, and P values were obtained using the Wald test. Continuous variables and the DL score entered the models as standardized values, and HR corresponded to a per-1 SD increase; the reference category for binary variables was uniformly defined as "no" or "none". The DL score was a composite score obtained by weighting deep learning features. The 2-year baseline survival rates S₀ (2 years) of the three models were 0.9387, 0.9194, and 0.9413, respectively. The 2-year risk of the combined model was calculated as: 2 - yearrisk = 1 - [S0(2 years)]exp(LP).

ModelApparent C (95% CI)Corrected CValidation C (95% CI)ΔC (95% CI)Apparent AUC₂y (95% CI)Corrected AUC₂yValidation AUC₂y (95% CI)ΔAUC₂y (95% CI)Apparent Brier₂yCorrected Brier₂yValidation Brier₂yValidation InterceptValidation Slope
Clinical model0.702 (0.657–0.747)0.6910.687 (0.619–0.754)Ref0.711 (0.665–0.757)0.70.694 (0.626–0.762)Ref0.0810.0820.0820.0730.901
Deep learning model0.734 (0.691–0.777)0.7220.713 (0.648–0.778)0.026 (−0.018–0.070)0.743 (0.698–0.789)0.7310.722 (0.658–0.786)0.028 (−0.016–0.072)0.0790.080.080.0580.843
Combined model0.787 (0.748–0.826)0.7730.759 (0.699–0.819)0.072 (0.030–0.114)0.799 (0.758–0.841)0.7850.774 (0.715–0.833)0.080 (0.038–0.122)0.0750.0760.0770.0120.972

Table 3: Predictive performance, optimism-corrected performance, and internal validation results of the three models. The corrected results are point estimates after 1,000 bootstrap optimism corrections. ΔC and ΔAUC₂y are the differences relative to the clinical model. Larger C and AUC₂y values and smaller Brier₂y values indicate better model performance; a calibration intercept closer to 0 and a calibration slope closer to 1 indicate better calibration. C, Harrell's concordance index; AUC₂y, 2-year time-dependent area under the receiver operating characteristic curve; Brier₂y, 2-year Brier score.

CohortNRI₂y95% CIPIDI₂y95% CIP
Derivation cohort0.3160.174–0.463<0.0010.0470.024–0.073<0.001
Internal validation cohort0.2410.058–0.3890.0090.0330.009–0.0580.007

Table 4: 2-year NRI and IDI of the combined model relative to the clinical model. Positive values of NRI₂y and IDI₂y indicate that the combined model has better incremental predictive value than the clinical model. NRI₂y and IDI₂y were both calculated based on the 2-year time-dependent method, and censored data were handled using the inverse probability of censoring weighting method; 95% CI was obtained by 1,000 bootstrap resamples, and P values were two-sided. NRI₂y, 2-year net reclassification improvement; IDI₂y, 2-year integrated discrimination improvement.

ModelDerivation nDerivation EventsDerivation C (95% CI)Derivation AUC₂y (95% CI)Derivation Brier₂yValidation nValidation EventsValidation C (95% CI)Validation AUC₂y (95% CI)Validation Brier₂yValidation InterceptValidation Slope
Clinical model14311290.699 (0.654–0.744)0.707 (0.661–0.752)0.082714620.681 (0.613–0.749)0.690 (0.622–0.759)0.0830.0840.892
Combined model14311290.783 (0.744–0.822)0.795 (0.753–0.837)0.076714620.753 (0.692–0.814)0.769 (0.709–0.829)0.0780.0190.964

Table 5: Complete-case sensitivity analysis. Complete cases were defined as patients with original observed values for all variables required by the corresponding model. The sensitivity analysis used complete-case analysis without multiple imputation. The 95% CI was obtained by 1,000 bootstrap resamples. C, Harrell's concordance index; AUC₂y, 2-year time-dependent area under the receiver operating characteristic curve; Brier₂y, 2-year Brier score.

Supplementary File 1: Raw data Please click here to download this file.

Discussion

The combined model still maintained the optimal performance after optimism correction and temporal internal validation, suggesting that the deep learning features from thoracolumbar lateral X-ray were not a simple repetition of clinical information, but could provide independent and verifiable incremental information for the risk assessment of incident vertebral fracture within 2 years. Its significance lies in incorporating the background of systemic fragility and local spinal structural fragility into the same prediction framework. Age, female sex, low body mass index, previous fragility fracture, diabetes, and glucocorticoid exposure reflect bone mass loss, impaired bone quality, insufficient muscular support, and susceptibility to refracture, and determine the patient's overall baseline fracture risk15; deep learning features are more likely to capture vertebral endplate morphology, slight wedging, sparse bone texture, cortical boundary changes, and abnormal mechanical distribution in the thoracolumbar region that are difficult to quantify stably by routine image reading, thus supplementing fragility information at the local imaging level16. The two types of information correspond to different pathological levels, and after combination, discrimination, calibration, prediction error, reclassification ability, and clinical net benefit were all improved, and this consistency supports that the model improvement was not accidental. Traditional risk models relying only on clinical variables are convenient for application, but it is difficult for them to identify local heterogeneity of the vertebrae17. Assessment strategies represented by bone mineral density or FRAX are more oriented toward systemic fracture tendency, and may not fully reflect the immediate structural fragility of the thoracolumbar region18. Previous artificial intelligence studies have mostly focused on the detection of existing vertebral fractures or the classification of osteoporosis, and are still one step away from clinical early warning19. The current results are closer to the real decision-making scenario, indicating that the occult phenotypes contained in routine X-ray, after being extracted by deep learning, can substantially enhance clinical risk stratification.

In vertebral fracture risk assessment, CT, MRI, bone mineral density-based assessment, and other imaging analysis methods each have their own applicable scenarios. CT more directly depicts vertebral morphology, endplate changes, and cortical bone destruction, and MRI has greater advantages in the assessment of bone marrow edema, soft-tissue involvement, and acute fractures, but both are inferior to thoracolumbar lateral radiographs in terms of examination cost, accessibility, and routine follow-up availability, making them difficult to use as early risk stratification tools on a large scale and with a low threshold. Bone mineral density measurement and FRAX are more suitable for reflecting the background of systemic bone fragility and have important reference value for the overall fracture tendency, but they are relatively limited in reflecting local structural fragility of the thoracolumbar region, mild wedging, subtle endplate abnormalities, and local mechanical imbalance. Existing radiomics methods can extract predefined quantitative features from radiographs, CT, or MRI and have potential in risk assessment, but they usually rely on manually predefined feature spaces and relatively strict segmentation procedures. Compared with these methods, the present study chose to construct a model based on routine thoracolumbar lateral radiographs, with the focus not on replacing CT, MRI, or bone mineral density assessment, but on supplementing, on the basis of the imaging modality most readily available in daily clinical practice, the occult local fragility information that is difficult to capture by traditional clinical assessment, thereby providing a more generalizable risk stratification pathway for the early identification of incident vertebral fracture.

The clinical variables entering the final model had clear pathophysiological implications, suggesting that this prediction framework was not the result of chance selection. Increasing age, female sex, and low body mass index correspond to bone mass loss, weakened muscular support, and increased susceptibility to falls, constituting the basic background of vertebral fragility. Previous fragility fracture history indicates persistent systemic bone fragility in the individual and is an important marker of refracture. Even when bone mineral density is not significantly reduced in patients with type 2 diabetes mellitus, deposition of advanced glycation end products, abnormal bone turnover, and microstructural impairment may still weaken vertebral mechanical strength20. Long-term oral glucocorticoid use inhibits bone formation, promotes bone resorption, impairs trabecular integrity, and leads to increased fracture risk21. After screening by stability, correlation, and penalized regression, only a small number of features were retained from the deep learning features to construct the DL score, indicating that the model captured imaging information that was stable and related to the outcome. These features are difficult to correspond one by one to a single manual indicator, and are more likely to comprehensively reflect subtle pre-collapse changes of the endplates, slight imbalance in vertebral morphology, sparse bone texture, changes in cortical contour, and abnormal local stress distribution in the thoracolumbar region. Therefore, they still retained independent predictive value after adjustment for clinical variables22. Existing epidemiological evidence has confirmed that the above clinical factors are closely associated with fragility fracture, and the results of the present study are basically consistent with this. Compared with traditional manual measurements or predefined radiomics features, deep learning does not require features to be prespecified, and is more suitable for identifying occult and complex fragility phenotypes in X-ray23. Rheumatoid arthritis and baseline anti-osteoporosis treatment did not enter the final model, which may be related to the lower prevalence of the former and treatment indication bias in the latter24. It can thus be seen that this model was established on the basis of complementary integration of the clinical risk spectrum and occult fragility phenotypes on X-ray, rather than a simple stacking of variables.

After bootstrap optimism correction, temporal internal validation, and complete-case sensitivity analysis, the advantage of the combined model remained stable, indicating that its predictive ability did not arise from within-sample fitting but had good internal validity. Temporal split validation is closer to the real application scenario than random splitting, and can more rigorously test the performance of the model in subsequent patients; optimism correction helps identify the risk of overfitting, and therefore the persistence of superiority after correction more strongly supports the robustness of the results. The calibration curve was close to the ideal line, the validation intercept was close to zero, and the slope was close to one, indicating that the model output was not merely a ranking score, but an absolute risk probability that was relatively consistent with the actual event occurrence level. This has greater clinical significance for determining follow-up intensity, further bone assessment, and the timing of preventive intervention. The higher net benefit within the prespecified threshold range indicates that, after adding deep learning features from X-ray, the model improvement was reflected not only in statistical indices, but also in potential benefit at the decision-making level25. The nomogram transformed the combined model into an interpretable individualized tool, which is conducive to completing risk stratification on the basis of routine thoracolumbar lateral X-ray examination26. Many previous artificial intelligence prediction studies mainly reported discrimination, paid insufficient attention to calibration, overfitting control, and clinical net benefit, and also lacked temporal validation or sensitivity analysis, thereby limiting transferability in real-world scenarios27,28. The complete chain of evidence formed around discrimination, calibration, prediction error, decision curve, and sensitivity analysis can better support the clinical translation of this combined model as a risk stratification tool for incident vertebral fracture.

This study was a single-center retrospective cohort study, and all cases were derived from hospital patients who underwent thoracolumbar lateral X-ray examination and completed imaging follow-up. The sample composition was influenced by referral pattern, examination indications, and follow-up adherence, and there was selection bias; therefore, caution is warranted when generalizing the results to other centers, community screening populations, or different equipment conditions. During the study period, baseline radiographs were acquired using the hospital's digital radiography system from a single vendor rather than multiple radiography systems/vendors, which reduced inter-vendor technical heterogeneity but may also limit generalizability to other imaging platforms. In particular, because outcome ascertainment required follow-up imaging, patients without imaging follow-up within 24 months were excluded, which may have preferentially retained patients with more symptoms, greater healthcare utilization, or higher baseline risk and may have increased the observed event rate. In addition, because follow-up imaging was obtained in routine clinical practice rather than under a fixed protocol, censoring may not have been completely non-informative, and the Cox-based risk estimates may still have been influenced by the follow-up imaging process. Although temporal internal validation, bootstrap optimism correction, and a complete-case sensitivity analysis were performed, independent external validation has not yet been conducted, and the cross-center stability and generalizability of the model remain to be confirmed. This study relied on routine lateral X-ray, which has the advantage of easy acquisition and dissemination, but compared with CT, MRI, or bone mineral density testing, its representation of bone microstructure, bone mass status, and adjacent tissue information remains limited; although deep learning features can improve predictive performance, their specific imaging and biological meanings are still not sufficiently intuitive. In addition, no dedicated feature-attribution or saliency analysis was performed; therefore, the related biological interpretations should be regarded as hypothesis-generating rather than directly validated. Candidate variables were mainly derived from structured medical records and routine clinical data, and did not include fall history, physical function, nutritional status, bone metabolism laboratory indices, or standardized bone mineral density measurements; therefore, residual confounding may still exist. Moreover, BMD- or FRAX-based models were not evaluated in the present study; therefore, the incremental value of the DL score was established only relative to the prespecified clinical model. Future studies should conduct external validation in multiple centers, with different equipment and in different clinical settings, and explore integration with bone mineral density, laboratory indices, and other imaging modalities, so as to improve the generalizability, interpretability, and practical application value of the model.

Disclosures

The authors declare that they have no conflicts of interest.

Acknowledgements

The authors thank the staff of the study hospital for their support in image retrieval, data extraction, and data management. The authors also thank all clinicians and radiologic technologists involved in patient care and imaging acquisition. This study was financially supported by the Medical Research Project in Xuhui District in 2024(SHXH202405).

Materials

List of materials used in this article
NameCompanyCatalog NumberComments
glmnet packageCRANN/AUsed for LASSO-Cox regression analysis.
ITK-SNAPUniversity of Pennsylvania / ITK-SNAP ProjectN/AUsed for ROI annotation of baseline images.
mice packageCRANN/AUsed for multiple imputation.
PythonPython Software Foundationversion 3.10Used for image preprocessing and deep learning analysis.
PyTorchPyTorch Foundation / Linux Foundationversion 2.1Used for deep learning model development and feature extraction.
R versionR Foundation for Statistical Computingversion 4.3.2Used for statistical analysis.
rmda packageCRANN/AUsed for decision curve analysis.
rms packageCRANN/AUsed for model development and calibration analysis.
survival packageCRANN/AUsed for Cox proportional hazards regression analysis.
timeROC packageCRANN/AUsed for time-dependent AUC analysis.

References

  1. Daskalakis II, Bastian JD, Mavrogenis AF, Tosounidis TH. Osteoporotic vertebral fractures: an update. SICOT J. 2025;11:40.
  2. Na D et al. Underdiagnosis and underreporting of vertebral fractures on chest radiographs in men aged over 50 years or postmenopausal women with and without type 2 diabetes mellitus: a retrospective cohort study. BMC Med Imaging. 2022;22(1):81.
  3. Urrutia J, Besa P, Piza C. Incidental identification of vertebral compression fractures in patients over 60 years old using computed tomography scans showing the entire thoraco-lumbar spine. Arch Orthop Trauma Surg. 2019;139(11):1497-1503.
  4. Zerikly R, Demetriou EW. Use of Fracture Risk Assessment Tool in clinical practice and Fracture Risk Assessment Tool future directions. Women's Health (Lond). 2024;20:17455057241231387.
  5. Hong N et al. Deep learning-based identification of vertebral fracture and osteoporosis in lateral spine radiographs and DXA vertebral fracture assessment to predict incident fracture. J Bone Miner Res. 2025;40(5):628-638.
  6. Hong N et al. Deep-Learning-Based Detection of Vertebral Fracture and Osteoporosis Using Lateral Spine X-Ray Radiography. J Bone Miner Res. 2023;38(6):887-895.
  7. Johansson L et al. Grade 1 Vertebral Fractures Identified by Densitometric Lateral Spine Imaging Predict Incident Major Osteoporotic Fracture Independently of Clinical Risk Factors and Bone Mineral Density in Older Women. J Bone Miner Res. 2020;35(10):1942-1951.
  8. Li Y et al. Machine learning value in the diagnosis of vertebral fractures: A systematic review and meta-analysis. Eur J Radiol. 2024;181:111714.
  9. Kong SH et al. Development of a Spine X-Ray-Based Fracture Prediction Model Using a Deep Learning Algorithm. Endocrinol Metab (Seoul). 2022;37(4):674-683.
  10. Lunt M et al. Defining incident vertebral deformities in population studies: a comparison of morphometric criteria. Osteoporos Int. 2002;13(10):809-815.
  11. Da Mutten R et al. Whole Spine Segmentation Using Object Detection and Semantic Segmentation. Neurospine. 2024;21(1):57-67.
  12. Xiao W, Chen R. A study on ACCC surface defect classification method using ResNet18 with integrated SE attention mechanism. Appl Sci. 2026;16(4):1899.
  13. Liu F, Zhang DB, Cheng SH, Gu GS. A radiomics and deep learning nomogram developed and validated for predicting no-collapse survival in patients with osteonecrosis after multiple drilling. BMC Med Inform Decis Mak. 2025;25(1):26.
  14. Yokota T et al. Internal validation of an 11-yr prediction model for new vertebral fractures using the vertebral bone quality score: a prospective cohort study. JBMR Plus. 2025;9(11):ziaf155.
  15. Chen W, Mao M, Fang J, Xie Y, Rui Y. Fracture risk assessment in diabetes mellitus. Front Endocrinol (Lausanne). 2022;13:961761.
  16. Kong SH. Incorporating Artificial Intelligence into Fracture Risk Assessment: Using Clinical Imaging to Predict the Unpredictable. Endocrinol Metab (Seoul). 2025;40(4):499-507.
  17. Schini M et al. An overview of the use of the fracture risk assessment tool (FRAX) in osteoporosis. J Endocrinol Invest. 2024;47(3):501-511.
  18. LeBoff MS et al. The clinician's guide to prevention and treatment of osteoporosis. Osteoporos Int. 2022;33(10):2049-2102.
  19. Gu Y, Wang Y, Li M, Wang R. Current applications of deep learning in vertebral fracture diagnosis. Osteoporos Int. 2025;36(11):2071-2082.
  20. Cavati G et al. Role of Advanced Glycation End-Products and Oxidative Stress in Type-2-Diabetes-Induced Bone Fragility and Implications on Fracture Risk Stratification. Antioxidants (Basel). 2023;12(4):928.
  21. Hofbauer LC, Compston JE, Saag KG, Rauner M, Tsourdi E. Glucocorticoid-induced osteoporosis: novel concepts and clinical implications. Lancet Diabetes Endocrinol. 2025;13(11):964-979.
  22. Saravi B et al. Integrating radiomics with clinical data for enhanced prediction of vertebral fracture risk. Front Bioeng Biotechnol. 2024;12:1485364.
  23. Zhang J et al. Differentiation of acute and chronic vertebral compression fractures using conventional CT based on deep transfer learning features and hand-crafted radiomics features. BMC Musculoskelet Disord. 2023;24(1):165.
  24. McGrath LJ et al. Using negative control outcomes to assess the comparability of treatment groups among women with osteoporosis in the United States. Pharmacoepidemiol Drug Saf. 2020;29(8):854-863.
  25. Piovani D, Sokou R, Tsantes AG, Vitello AS, Bonovas S. Optimizing Clinical Decision Making with Decision Curve Analysis: Insights for Clinical Investigators. Healthcare (Basel). 2023;11(16):2244.
  26. Nguyen HT et al. A predictive nomogram for selective screening of asymptomatic vertebral fractures: The Vietnam Osteoporosis Study. Osteoporos Sarcopenia. 2025;11(1):9-14.
  27. Hu Y et al. Beyond Comparing Machine Learning and Logistic Regression in Clinical Prediction Modelling: Shifting from Model Debate to Data Quality. J Med Internet Res. 2025;27:e77721.
  28. Groot OQ et al. Availability and reporting quality of external validations of machine-learning prediction models with orthopedic surgical outcomes: a systematic review. Acta Orthop. 2021;92(4):385-393.

Reprints and Permissions

Tags

LASSO Cox RegressionModel ValidationRisk PredictionDecision Curve AnalysisNet Reclassification Improvement