Research Article

Prediction of Chronic Low Back Pain Risk Based on Dietary Trace Elements Using Multiple Machine Learning Models and Dual Interpretability Frameworks

DOI:

10.3791/70624

May 26th, 2026

In This Article

Summary

Loading...
$$\rightleftharpoonup{xx}$$ $$\longleftharp{xx}$$, $$\longrightharp{xx}$$,

This study investigates associations between dietary trace elements and chronic low back pain using NHANES data. Interpretable machine learning models are applied to evaluate predictive performance and identify nutrients associated with reduced risk, providing insights into potential dietary factors linked to low back pain.

Abstract

Loading...
$$\rightleftharpoonup{xx}$$ $$\longleftharp{xx}$$, $$\longrightharp{xx}$$,

The relationship between dietary trace elements and the risk of chronic low back pain (LBP) remains unclear. Data from the National Health and Nutrition Examination Survey (NHANES) 2001–2004 and 2009–2010 cycles were analyzed. Multicollinearity was assessed using Spearman’s correlation and variance inflation factors (VIF), and feature selection was performed using the Boruta algorithm. Six machine learning models were subsequently developed, with SHAP and LIME applied to enhance interpretability and identify key associations between dietary trace elements and LBP risk. Among the evaluated models, Random Forest demonstrated the best predictive performance, both when incorporating demographic variables and when using dietary trace elements alone. Interpretability analyses consistently identified several dietary components—including moisture, theobromine, calcium, caffeine, sodium, and vitamin C—as inversely associated with LBP risk. Although model discrimination was moderate (AUC ≈ 0.60–0.72), these findings reveal clinically relevant patterns that may support population-level risk stratification and highlight potentially modifiable dietary factors for future investigation.

Introduction

Loading...
$$\rightleftharpoonup{xx}$$ $$\longleftharp{xx}$$, $$\longrightharp{xx}$$,

Chronic low back pain (LBP) is among the most prevalent musculoskeletal disorders globally and has significant socioeconomic costs associated with disability, healthcare, and loss of quality of life. Epidemiological data indicate that 7–10% of adults report LBP globally, with lifetime prevalence sometimes reaching 84%1. It is most common in low/middle-income countries, where work-related factors, poor access to health services, and socio-economic inequity increase risk, with prevalence often exceeding 20% among working-age adults2. Adolescents and the elderly, particularly, are more vulnerable to structural and degenerative changes such as intervertebral disc degeneration and osteoporosis3.

LBP is a multifactorial disease influenced by biomechanical, psychosocial, genetic, and nutritional factors, necessitating analytical methods capable of simultaneously capturing the complex interactions among these diverse variables. Recent systematic reviews continue to highlight that the etiology of LBP remains enigmatic, influenced by biomechanical, psychosocial, and genetic factors, and that it is thus difficult to pinpoint with specificity. More than 85% of cases of ‘LBP are nonspecific’ - making the development of targeted prevention and treatment challenging4,5. Although multimodal therapies encompassing a combination of physical and pharmacological methods tend to provide symptomatic relief, many patients remain recurrent, a sign that points to an urgent need for new, preventable, and modifiable risk factors for LBP6.

Traditional statistical methods are limited in addressing this complexity, as they often rely on linear assumptions and fail to adequately model high-dimensional interactions and nonlinear relationships inherent in multifactorial data such as NHANES dietary profiles. In contrast, machine learning approaches offer superior utility by handling large numbers of variables concurrently, detecting subtle patterns and interactions, and providing robust predictions for diseases like LBP where multiple factors converge.

Dietary micronutrients, including vitamins, minerals, and trace elements, are necessary for physiological balance and help reduce the risk of chronic disease. The antioxidant vitamins E (α-tocopherol) and C reduce oxidative damage present in cardiovascular and neurodegenerative diseases7,8. B-complex vitamins, such as thiamine (vitamin B1), pyridoxine (vitamin B6), folate, and cobalamin (vitamin B12), are also immunoregulatory and facilitate cellular energy metabolism; deficiencies of which have been linked to cancer, infertility, and depressed mood9,10,11. Minerals, particularly magnesium, zinc, copper, and selenium, participate in enzymatic and anti-inflammatory pathways that prevent sarcopenia, endometriosis, and postoperative nutritional discharge after bariatric surgery12. Narrative reviews highlight their importance in preventing obesity-related comorbidities and skin conditions such as seborrheic dermatitis and alopecia areata via lipid modulation and immunologic roles13,14. Overall, a signal that the best approach to community health may be adequate micronutrient intake, though interventional trials largely show heterogeneous outcomes based on contextual factors15,16.

Furthermore, machine learning applications incorporating nutritional data have shown efficacy in elucidating risk factors for other musculoskeletal diseases, such as predicting arthritis risk from survey data17 and identifying lifestyle (including dietary) contributors to LBP and related conditions in population-based cohorts18.

Despite a wealth of links between micronutrients and chronic disease, the relationship between micronutrients in food and LBP risk remains poorly defined, and the mechanisms of action are uncertain. While observational studies frequently report deficiencies in vitamin D, magnesium, and antioxidants among individuals with LBP, meta-analyses of supplementation trials—particularly those involving vitamin D—have yielded limited and inconsistent evidence of pain relief19,20.

A systematic examination of existing literature reveals key limitations: reliance on cross-sectional and case-control designs that cannot establish causality21; inadequate control for confounders, including demographics, lifestyle, and comorbidities20,22; small and non-representative samples; variable definitions of micronutrient status and pain outcomes; and a focus on isolated nutrients without considering dietary synergies or antagonisms23. These shortcomings have hindered the development of clear, actionable insights into nutritional modulation of LBP.

For a more explicit hypothesis structure, it is proposed that micronutrients exert protective effects against LBP through a triad of mechanistically linked pathways: antioxidant and anti-inflammatory actions that attenuate oxidative stress and neuroinflammation (vitamins C, E, selenium, zinc, magnesium); support for neural repair and mitigation of homocysteine-induced damage (B-complex vitamins); and enhancement of bone and muscle homeostasis via mineral signaling and vitamin D-mediated calcium regulation (vitamin D, magnesium, calcium). Such a framework highlights the bidirectional interplay between micronutrient status, systemic inflammation, and pain chronicity, as evidenced by associations with pro-inflammatory diets24.

Cross-sectional studies indicate that pro-inflammatory diets lacking micronutrients produce greater levels of abdominal or chronic pain, suggesting a bidirectional association in which a lack of micronutrient inputs aggravates systemic inflammation and pain24.

To address these issues, the developed multiple ML models that incorporate a wide array of diet micronutrients (vitamins (E, A, B1, B6, B12, C, K), minerals (calcium, magnesium, iron, zinc, copper, selenium) and other dietary elements including fiber, polyunsaturated fatty acids, potassium and caffeine) with NHANES data for predicting LBP risk. This ML-based strategy is particularly suited to the multifactorial nature of LBP, as it overcomes the limitations of traditional statistics by modeling complex, nonlinear associations among nutritional and other risk factors simultaneously. This more exhaustive exploration can handle nonlinear associations and provides a robust equation for key LBP predictors. Using SHAP values for global interpretability, they can see how combinations of nutrients affect LBP risk, which opens the door to personalized interventions and nutrigenomic studies.

Access restricted. Please log in or start a trial to view this content.

Protocol

Loading...
$$\rightleftharpoonup{xx}$$ $$\longleftharp{xx}$$, $$\longrightharp{xx}$$,

This investigation was approved by the National Center for Health Statistics (NCHS) Ethics Review Board, and informed consent was obtained from all participants. Publicly available NHANES data were used in compliance with relevant guidelines.

Study population

The National Health and Nutrition Examination Survey (NHANES), conducted by the NCHS, provides nationally representative data on the health and nutritional status of the non-institutionalized U.S. population. Data from the 2001–2004 and 2009–2010 cycles were retrieved. Participants aged ≥18 years were included, while individuals with missing data on chronic LBP, dietary micronutrients, education level, or key covariates—including poverty-income ratio (PIR), body mass index (BMI), smoking status, hypertension, diabetes, and alcohol consumption—were excluded. The final analytical cohort comprised 8,710 participants (Figure 1).

Assessment of dietary nutrients

Dietary intake data were obtained using two 24-hour dietary recall interviews. The first interview was conducted at a Mobile Examination Center, and the second was completed via telephone within several days. The Automated Multiple-Pass Method was applied to improve recall accuracy and minimize reporting bias. Detailed AMPM procedures are described in the NHANES dietary methodology documentation.

Assessment of chronic LBP

Chronic LBP status was determined using questionnaire data. Participants were classified as having LBP if they reported back pain lasting ≥3 months.

Covariates

Demographic characteristics included age, sex, race/ethnicity (Mexican American, Other Hispanic, Non-Hispanic White, Non-Hispanic Black, Other Race), education (<9th grade, 9–11th), degree (high school diploma/GED, some college/AA degree and college graduate), family income-to-poverty ratio (PIR), body mass index (BMI), smoking status as well as drinking status and history of hypertension and diabetes. Hypertension was defined based on self-reported physician diagnosis or current use of antihypertensive medication. Diabetes was defined using self-reported physician diagnosis, elevated glucose levels after a 2-hour glucose tolerance test (≥11.1 mmol/L), or glycated hemoglobin ≥6.5% or fasting glucose ≥7.0 mmol/L. Prediabetes was defined as being told by a physician that you had prediabetes, 2 hr OGTT glucose tolerance values of 7.8–11.1 mmol/L, fasting glucose of 6.1–6.9 mmol/L, or an HbA1c value of between 5.7%–6.4%. The status of smoking was defined as (1) Non-current before the last month, or quit for more than a year, and (2) smoking currently (self-reporting recent) is defined as smoking more than 2 cigarettes/day upon regaining the habit or prior to quitting. As regards drinking status, two definitions exist: (1) lifetime nondrinkers with no prior alcohol consumption (<12 total), and (2) current drinkers who report drinking ≥12 times per year, or drinking on at least six occasions over the last 12 months. BMI calculated as weight ÷ height2.

Feature preprocessing and selection for machine learning

A total of 52 variables, including 45 continuous and 7 categorical variables, were included. All preprocessing and feature selection procedures—including multicollinearity removal, SMOTE application, and Boruta-based feature selection—were performed exclusively on the training set to prevent data leakage. To eliminate multicollinearity, the variables with correlation coefficients greater than 0.9 were removed first, and then the variables with VIFs greater than 3.

To improve class imbalance and the recognition of minority (read: more difficult) classes, the Synthetic Minority Over-sampling Technique (SMOTE) was applied (Supplementary Figure 1), which creates distances between minority-class samples, calculates the k-nearest neighbors for those distances, and generates synthetic data in random directions to balance classes. All variables used were then standardized.

Feature selection was completed with Boruta on Random Forests, generating so-called “shadow features” and “testing” them versus the input features over 500 iterations. Those classified as “confirmed” were retained to construct the model. For detailed information on hyperparameter tuning, please refer to Supplementary Table 1.

Statistical analysis

All analyses were conducted in accordance with NHANES analytical guidelines. Continuous variables were reported as mean ± standard deviations (SD), while categorical variables were expressed as count and percentage. Between-group differences were tested using chi-square or Student’s t-tests where appropriate.

To prevent overfitting, the data were randomly partitioned into a training set7 and a test set. Using MLR3, six machine learning algorithms were built: Random Forest which constructs many decision trees and joins their predictions, having strong generalization power and consistency; LightGBM based on gradient-boosted decision trees, which is tuned for efficiency and scale; K-KNN sort samples according to how close they are to their neighbors K-KNN typically has better results on small non-linear datasets; Naive Bayes, which by utilizing Bayes’ theorem gives simplicity and is robust; SVM which possesses ideal hyperplanes for the classification task, even for samples in high dimensionality; XGBoost is a form of gradient boosting ensemble that possesses high performance, even with very large incomplete datasets.

The model was assessed using accuracy, F-beta, area under the ROC curves (AUC-ROC), the sensitivity, specificity, and AUC-PR. The AUC-ROC is the main metric they use, while the other metrics give deeper insights into performance. They validate their models using tenfold cross-validation to ensure stability. Differences in their models' performance were evaluated using ANOVA and the Kruskal–Wallis H test.

Feature interpretability was investigated using SHapley Additive exPlanations (SHAP) and Local Interpretable Model-agnostic Explanations (LIME). SHAP derives estimates of each feature's contribution to model predictions using Shapley values, based on cooperative game theory, to represent both linear and interactive effects. LIME derives feature influence by approximating complex models with simpler, interpretable surrogates (e.g., Lasso regressions) to help explain individual feature influence. All analyses were performed in IBM SPSS Statistics version 24.0 and in R version 4.3.0. A two-tailed p-value < 0.05 is regarded as statistically significant.

Access restricted. Please log in or start a trial to view this content.

Results

Loading...
$$\rightleftharpoonup{xx}$$ $$\longleftharp{xx}$$, $$\longrightharp{xx}$$,

Baseline characteristics by chronic LBP status

Baseline characteristics stratified by chronic LBP status are described in Table 1. The final analytic cohort included 8,710 participants from NHANES 2001–2004 and 2009–2010, with a mean age of 49.13 years (SD = 18.43). Of these, 4,528 (51.99%) were women and 4,182 (48.01%) were men. Overall, 3,838 participants reported having LBP (mean age 48.62 years; SD = 17.69). Compared to participants without LBP, those with...

Access restricted. Please log in or start a trial to view this content.

Discussion

Loading...
$$\rightleftharpoonup{xx}$$ $$\longleftharp{xx}$$, $$\longrightharp{xx}$$,

This study used NHANES data and applied several algorithms, finding an association between dietary micronutrients and chronic LBP. Importantly, given the cross-sectional observational design, these findings reflect predictive associations rather than causal relationships, and the machine learning models cannot establish causality between micronutrient intake and LBP risk. After feature selection and collinearity testing, the Random Forest appeared to give the best predictions, both with demographics included and just die...

Access restricted. Please log in or start a trial to view this content.

Disclosures

Loading...
$$\rightleftharpoonup{xx}$$ $$\longleftharp{xx}$$, $$\longrightharp{xx}$$,

The authors have no conflicts of interest to declare.

Acknowledgements

Loading...
$$\rightleftharpoonup{xx}$$ $$\longleftharp{xx}$$, $$\longrightharp{xx}$$,

This work was supported by Hangzhou Project Foundation for Agriculture and Social Development (20241029Y119), Xiaoshan District Science and Technology Plan Guidance Project (2025316), and Zhejiang Province Traditional Chinese Medicine Science and Technology Project (2024ZR152).

Access restricted. Please log in or start a trial to view this content.

Materials

List of materials used in this article
NameCompanyCatalog NumberComments
Automated Multiple-Pass Method (AMPM)National Center for Health Statistics (NCHS), CDCN/ADietary assessment system used in NHANES 24-hour dietary recall interviews to improve reporting accuracy
Boruta AlgorithmOpen-source (R package: Boruta)N/AFeature selection method based on random forests; used to identify relevant variables through shadow feature comparison
Computer-Assisted Personal Interviewing (CAPI)National Center for Health Statistics (NCHS), CDCN/AInterview system used by trained personnel to collect questionnaire data
National Health and Nutrition Examination Survey (NHANES)National Center for Health Statistics (NCHS), CDCN/ANationwide survey providing demographic, dietary, and health-related data
Prostate Conditions Questionnaire (KIQ)National Center for Health Statistics (NCHS), CDCN/AQuestionnaire source used to assess back pain information
Random Forest AlgorithmOpen-source (e.g., R / Python implementations)N/AMachine learning model used in feature selection and classification
Synthetic Minority Over-sampling Technique (SMOTE)Open-source (e.g., DMwR / imbalanced-learn)N/AOversampling technique used to address class imbalance by generating synthetic minority samples

Reprints and Permissions

Request permission to reuse the text or figures of this JoVE article

Request Permission

Tags

Risk PredictionRandom ForestSHAP AnalysisLIME InterpretabilityFeature SelectionNHANES DataPopulation Risk Stratification

Related Articles