A subscription to JoVE is required to view this content. Sign in or start your free trial.

Research Article

Assessment of Malnutrition in Crohn's Disease Patients: A Novel Risk Prediction Model with Dynamic Optimization Potential and Effectiveness Validation

93 views

⸱

DOI:

10.3791/70247

⸱

June 22nd, 2026

 ,  ,  ,  ,  ,  ,  , 

* These authors contributed equally

In This Article

Summary

This study constructed and validated a novel malnutrition risk prediction model for patients with Crohn’s disease using meta-analysis, multivariable logistic regression, and machine learning. The model demonstrated good predictive performance and clinical utility for guiding personalized nutritional interventions.

Abstract

This study aimed to construct and validate a malnutrition risk prediction model combining multivariable logistic regression and machine learning for patients with Crohn’s disease (CD), with the goal of improving the precision of malnutrition risk identification through integration of inflammatory markers and disease characteristics. PubMed, Web of Science, Cochrane Library, Embase, and China National Knowledge Infrastructure (CNKI) were systematically searched to identify risk factors associated with malnutrition in patients with CD. High-quality studies using the Global Leadership Initiative on Malnutrition (GLIM) 2019 criteria, European Society for Clinical Nutrition and Metabolism (ESPEN) 2015 criteria, or Malnutrition Universal Screening Tool (MUST) criteria were included in the meta-analysis, while the study cohort applied the ESPEN 2015 criteria exclusively to ensure consistent outcome definition. The prediction model was developed using data from 800 patients with CD from the Inflammatory Bowel Disease Cohort Database (IBDCD) and validated using bootstrap resampling and an independent non-overlapping hold-out subset of 280 patients from the same database. Model performance was evaluated using the area under the receiver operating characteristic curve (AUC), Hosmer–Lemeshow test, and Brier score. No significant baseline differences were observed between the training and validation cohorts. The model achieved an AUC of 0.987 in the training cohort and 0.967 in the validation cohort, demonstrating good discrimination and calibration. Decision curve analysis further demonstrated clinically meaningful net benefits across relevant threshold probabilities. This model effectively identifies malnutrition risk in patients with CD and may support personalized nutritional intervention, optimize clinical decision-making, and improve patient outcomes and quality of life. Future multicenter studies are required to further validate the model's generalizability and to evaluate the integration of socioeconomic factors for further optimization.

Introduction

Crohn’s disease (CD) is a chronic, progressive inflammatory bowel disease (IBD) with a multifactorial etiology, and its global burden has increased substantially in recent years1. Epidemiological studies indicate that although CD remains more prevalent in Western countries, the incidence and prevalence of CD in Asian populations, particularly in China, have increased markedly over the past three decades1.

Malnutrition is one of the most common and clinically significant complications in patients with CD.^2 Approximately one-third to one-half of patients with CD experience moderate to severe malnutrition2. Malnutrition is associated with prolonged hospitalization, increased surgical requirements, higher rates of postoperative complications, accelerated intestinal fibrosis, increased glucocorticoid dependency, reduced therapeutic response, and decreased health-related quality of life (HRQoL)3,4. In addition, malnutrition contributes to immune dysfunction, thereby exacerbating intestinal inflammation and creating a vicious cycle that imposes substantial socioeconomic and healthcare burdens3,4.

Several nutritional assessment tools are currently used in clinical practice; however, their applicability to patients with CD remains limited5. The Patient-Generated Subjective Global Assessment (PG-SGA), although recommended for nutritional screening in patients with chronic disease, relies heavily on subjective clinical judgment and demonstrates limited consistency among clinicians when applied to CD populations6. In addition, PG-SGA lacks specificity for intestinal malabsorption. The Nutritional Risk Screening 2002 (NRS 2002), which incorporates body mass index (BMI)-based thresholds, has shown relatively low sensitivity in Asian populations and may contribute to underdiagnosis7. Similarly, the Mini Nutritional Assessment-Short Form (MNA-SF), originally developed for elderly populations, demonstrates relatively high false-positive rates in younger adults with CD.

Importantly, most currently available nutritional assessment tools are based on static scoring systems and do not adequately capture the dynamic fluctuations in nutritional status associated with disease activity in CD8. As CD is characterized by alternating periods of remission and relapse, nutritional status may deteriorate or improve over time in parallel with inflammatory activity9. Conventional assessment approaches lack the ability to perform dynamic risk stratification and may therefore fail to identify early subclinical deterioration in nutritional status. Furthermore, many existing tools do not incorporate disease-specific biomarkers or inflammatory indicators, such as C-reactive protein (CRP), serum albumin, and fecal calprotectin, thereby limiting their predictive accuracy in complex clinical settings10.

Recent advances in artificial intelligence and machine learning (ML) have introduced new opportunities for improving nutritional risk assessment in chronic inflammatory diseases11. Compared with traditional regression-based approaches, ML algorithms can identify nonlinear associations and complex interactions among variables, making them particularly suitable for heterogeneous clinical datasets with multidimensional predictors. Emerging evidence suggests that ML-based nutritional risk prediction models may outperform conventional statistical approaches in predicting nutritional deterioration in patients with IBD12. In addition to improving predictive performance, ML approaches may reduce multicollinearity through regularization techniques and optimize predictor selection to improve model interpretability and generalizability13,14.

The increasing availability of large-scale public databases has further strengthened nutrition-related research in CD. International databases, including the IBD Biobank, National Health and Nutrition Examination Survey (NHANES), and UK Biobank, contain extensive longitudinal clinical data encompassing nutritional indices, inflammatory biomarkers, genomic information, and lifestyle characteristics15,16,17. Integration of these multidimensional datasets enables cross-population validation and improves the external applicability of predictive models. Meta-analyses have demonstrated that prediction models developed using multicenter integrated datasets exhibit greater stability and reproducibility than models derived from single-center cohorts.

This study hypothesized that integrating multidimensional predictive variables, including inflammatory markers, disease characteristics, and demographic and socioeconomic factors, combined with multivariable logistic regression and machine learning algorithms would improve the accuracy and efficiency of malnutrition risk prediction in patients with CD. The resulting prediction model may facilitate personalized nutritional intervention, optimize clinical decision-making, and improve patient outcomes and quality of life.

Access restricted. Please log in or start a trial to view this content.

Protocol

All patients whose data were entered into the Inflammatory Bowel Disease Cohort Database (IBDCD) provided written informed consent for the use of de-identified clinical data for research purposes at the time of database enrollment. The same ethical approval and consent framework applied to all data extracted from the database, including both the development and validation cohorts. No identifiable patient information was used in this study, and all data were de-identified and encrypted for storage. As this study involved retrospective analysis of de-identified data from an established database, the ethics committee waived the requirement for additional patient consent for this specific analysis. The research tools used in the protocol are listed in the Table of Materials.

1. Screening of influencing factors in the meta-analysis

  1. Literature search strategy
    A systematic search strategy was conducted across PubMed, Web of Science, Cochrane Library, Embase, and China National Knowledge Infrastructure (CNKI). The English search strategy was defined as follows:

    ("Crohn Disease"[Mesh] OR "CD"[tiab] OR "Crohn*"[tiab]) AND ("Malnutrition"[Mesh] OR "Nutritional Status"[Mesh] OR "Nutritional Risk"[tiab]) AND ("Risk Factors"[Mesh] OR "Predict*"[tiab])

    The search period covered database inception through March 31, 2025. In addition, manual screening of reference lists from included studies was performed to identify potentially missed articles. This strategy enabled a comprehensive literature review of malnutrition-related risk and predictive factors in patients with Crohn’s disease (CD), with clearly defined databases, search terms, time frame, and supplementary manual retrieval procedures to ensure completeness of the literature collection.
  2. Inclusion and exclusion criteria
    Inclusion criteria required that study populations consist of adult patients (aged ≥18 years) diagnosed with Crohn’s disease according to the 2018 Consensus on the Diagnosis and Treatment of Inflammatory Bowel Disease. Given the clinical significance of malnutrition in CD, emphasis was placed on studies defining malnutrition using the European Society for Clinical Nutrition and Metabolism (ESPEN) 2015 criteria, Global Leadership Initiative on Malnutrition (GLIM) 2019 criteria, or Malnutrition Universal Screening Tool (MUST) score ≥2. A clear distinction was maintained between literature-defined diagnostic criteria and the exclusive application of the ESPEN 2015 criteria for outcome definition within the study cohort.

    Eligible study designs included cohort, case-control, cross-sectional, and mixed-methods studies to ensure comprehensive evaluation of malnutrition risk factors from multiple perspectives. Exclusion criteria included animal studies, review articles, and conference abstracts because of insufficient data support or lack of peer review. Studies lacking essential statistical information, including odds ratios (ORs) and corresponding 95% confidence intervals (CIs), were also excluded.

2. Literature screening and data extraction

Literature screening and data extraction were independently conducted by two researchers to ensure completeness and accuracy. The initial stage involved screening titles and abstracts to exclude studies that clearly failed to meet the inclusion criteria, thereby narrowing the dataset to potentially relevant, high-quality studies.

Subsequently, full-text screening was performed for studies considered potentially eligible. This stage involved detailed evaluation of study design, methodology, results, and adherence to predefined inclusion criteria to ensure that only studies meeting all requirements were included in the final analysis.

Data extraction was performed using standardized forms to systematically record key information, including author names, publication year, country, sample size, Montreal classification subtype distribution, average disease duration, definitions of malnutrition, assessment tools, and extracted ORs with corresponding 95% CIs for influencing factors associated with malnutrition risk.

Any discrepancies or uncertainties arising during screening and data extraction were resolved through internal discussion. When consensus could not be reached, a third-party expert was consulted to arbitrate, minimizing subjective bias and ensuring objective decision-making. This rigorous process enhanced methodological transparency and provided a robust foundation for the identification of key factors associated with malnutrition risk in CD.

3. Quality assessment

Rigorous quality assessment criteria were applied to ensure the scientific validity and reliability of the included studies. Cohort and case-control studies were evaluated using the Newcastle–Ottawa Scale (NOS), which assesses study quality across three domains: selection, comparability, and exposure. Only studies with NOS scores ≥7 were classified as high quality and included in the analysis.

Cross-sectional studies were evaluated using the Agency for Healthcare Research and Quality (AHRQ) assessment tool, which examines participant selection, sample size adequacy, validity of data collection methods, and appropriateness of statistical analyses. Studies with AHRQ scores ≥8 were considered eligible for inclusion.

Application of these stringent quality assessment criteria minimized potential bias related to study quality variability and ensured that the predictive model was supported by a reliable evidence base.

4. Prediction model construction

  1. Data source
    The Inflammatory Bowel Disease Cohort Database (IBDCD) is a multicenter prospective cohort database containing clinical data from patients with Crohn’s disease collected between January 2018 and March 2025, including demographic information, clinical characteristics, laboratory indicators, treatment regimens, and nutritional status assessments. The model development dataset included 800 patients with Crohn’s disease derived from the IBDCD database, consisting of 520 patients in the training cohort and 280 patients in the validation cohort.

    Inclusion criteria required a confirmed diagnosis of Crohn’s disease for at least 6 months and availability of complete clinical data, including identified influencing factors and nutritional assessment data derived from the meta-analysis. Patients with comorbid conditions potentially affecting nutritional status, including malignancies or chronic kidney disease, were excluded to ensure data accuracy and model validity.
  2. Variable definition and assignment
    Malnutrition was diagnosed strictly according to the ESPEN 2015 criteria, which include three core diagnostic components: unintentional weight loss (≥5% within 3 months or ≥10% within 6 months), reduced food intake or absorption (≥25% reduction for ≥14 days), and decreased muscle mass assessed through physical examination and anthropometric measurements, including mid-arm muscle circumference. A low body mass index (BMI < 18.5 kg/m2) was considered a supportive rather than a definitive diagnostic indicator. Malnutrition was confirmed only when at least one core criterion was present, while low BMI served as a supplementary indicator. Patients presenting with low BMI without evidence of the core criteria were not classified as malnourished. This definition avoided circularity between low BMI, a predictive variable, and malnutrition, an outcome.

    The prevalence of malnutrition was 42.5% in the training cohort (n = 520) and 40.4% in the validation cohort (n = 280). Based on the meta-analysis results, five key influencing factors were selected as predictive variables, including disease activity defined as C-reactive protein (CRP >10 mg/L; Yes = 1, No = 0), small bowel involvement defined by Montreal classification (LL1/LL3 versus LL2 colonic involvement; Yes = 1, No = 0), biologic use (Yes = 1, No = 0), history of intestinal resection (Yes = 1, No = 0), and low BMI (<18.5 kg/m2; Yes = 1, No = 0). Age and sex were collected as baseline demographic characteristics but were not included as predictive variables in the final model because they were not statistically significant during preliminary analysis.

    A total of 17 high-quality studies were included in the meta-analysis. The pooled ORs and corresponding 95% CIs for the five predictive factors were as follows: elevated CRP (OR = 4.72, 95% CI: 3.21–6.95), small bowel involvement (OR = 2.89, 95% CI: 1.93–4.33), biologic use (OR = 0.39, 95% CI: 0.15–1.01), history of intestinal resection (OR = 6.17, 95% CI: 2.35–16.18), and low BMI (OR = 3.56, 95% CI: 2.41–5.27).
  3. Model construction method
    The pooled ORs and corresponding 95% CIs derived from the meta-analysis were transformed into log-OR values and used to derive the regression coefficients (β) of the multivariable logistic regression model. Consistency between the meta-analysis results and model coefficients was verified using Pearson correlation analysis (r = 0.98, P < 0.001). Based on the identified influencing factors, a multivariable logistic regression model was constructed according to the following equation:

    "logit"(P) = α + β1 X1 + β2 X2 + β3 X3 + β4 X4 + β5 X5

    where P represents the probability of malnutrition occurrence and β values represent the natural logarithm of the pooled OR values derived from the meta-analysis.

    Using R software (version 4.2.1) and the “rms” package, the model was translated into a clinically applicable nomogram. Original risk scores (theoretical range: 0–325.1) were linearly scaled to a 0–100 range to improve clinical readability. Based on total nomogram scores, patients were stratified into low-risk (≤20 points), moderate-risk (21–40 points), and high-risk (>40 points) categories using percentile-based calibration relative to the cohort score distribution.

    Machine learning algorithms, including random forest (RF) and gradient boosting decision tree (GBDT), were used for feature optimization. A stacking strategy combined RF and GBDT as base learners with logistic regression as the meta-classifier. Hyperparameters were optimized using 5-fold cross-validation, with RF configured as n_estimators = 200 and max_depth = 10, and GBDT configured as n_estimators = 150 and learning_rate = 0.1. The Synthetic Minority Oversampling Technique (SMOTE) was applied to address minor class imbalance within the training dataset.

5. Model validation

  1. Validation dataset
    A dual validation strategy was employed. Internal validation was performed using bootstrap resampling with 1,000 repetitions. Hold-out validation was conducted using an independent non-overlapping subset of 280 patients extracted from the IBDCD database, excluding patients included in the training cohort. All patients in the validation cohort were diagnosed with Crohn’s disease according to the 2018 Inflammatory Bowel Disease Consensus, with prospectively collected data spanning January 2018 to March 2025.

    The validation cohort followed the same ascertainment procedures as the training cohort. Malnutrition was assessed using the ESPEN 2015 criteria, and all predictive variables, including CRP, lesion location, biologic use, history of intestinal resection, and BMI, were measured within 72 h of hospital admission. No significant differences were observed between the training and validation cohorts for baseline characteristics, including age, sex, and disease duration (P > 0.05), supporting the reliability and clinical applicability of the validation results.

    Additional performance metrics were calculated to further evaluate model performance. In the training cohort, sensitivity was 92.3%, specificity was 94.1%, positive predictive value (PPV) was 90.5%, and negative predictive value (NPV) was 95.2%. In the validation cohort, sensitivity was 90.2%, specificity was 92.8%, PPV was 88.7%, and NPV was 93.9%. The Brier score was 0.087 in the training cohort and 0.102 in the validation cohort, indicating acceptable prediction error.
  2. Validation metrics
    Model discrimination was evaluated using the area under the receiver operating characteristic (ROC) curve (AUC). Calibration was assessed using the Hosmer–Lemeshow test, where P > 0.05 indicated no significant difference between predicted probabilities and observed outcomes, and the Brier score, where values <0.20 indicated acceptable prediction error. Decision curve analysis (DCA) was additionally performed to evaluate clinical net benefit across a range of threshold probabilities.

    This validation framework adhered to the Transparent Reporting of a multivariable prediction model for Individual Prognosis or Diagnosis (TRIPOD) statement guidelines.

6. Statistical analysis

Meta-analysis was conducted using RevMan version 5.4, and pooled ORs with corresponding 95% CIs were calculated. Heterogeneity was assessed using the I2 statistic; values>50% indicate substantial heterogeneity. Random-effects models were applied when significant heterogeneity was present; otherwise, fixed-effects models were used. Sensitivity analysis was performed by sequentially excluding individual studies to evaluate the stability of the results. Publication bias was assessed using funnel plots and Egger’s test.

Basic statistical analyses were conducted using SPSS version 26.0. The “pROC”, “rms”, and “glmnet” packages in R software (version 4.2.1) were used for model construction and validation.

Access restricted. Please log in or start a trial to view this content.

Results

Basic characteristics of the study population
This meta-analysis included 17 high-quality studies evaluating malnutrition in patients with Crohn’s disease, as summarized in Supplementary Table 1. The studies were published between 2009 and 2025, with a median sample size of 502 patients (range: 175–773). Study designs included prospective cohort, retrospective cohort, case-control, cross-sectional, and mixed-methods studies, with mixed-methods studies accounting for 17.6% (3/17) of t...

Access restricted. Please log in or start a trial to view this content.

Discussion

Crohn’s disease (CD) is a chronic, recurrent inflammatory bowel disease that can affect any part of the gastrointestinal tract1,35,36,37. The etiology of CD remains unclear and is thought to involve multiple factors, including genetic susceptibility, environmental exposures, and immune dysregulation. Common clinical manifestations include abdominal pain, diarrhea, weight loss, and fistula...

Access restricted. Please log in or start a trial to view this content.

Disclosures

The authors declare no conflicts of interest.

Materials

List of materials used in this article
NameCompanyCatalog NumberComments
RevMan softwareCochrane Collaboration‌5.4The professional meta-analysis software developed by the Cochrane collaboration is mainly used for the formulation of systematic reviews and meta-analyses, data entry, and visualization of results. ‌
R softwareThe R Project for Statistical Computing4.2.1R is a branch of the S language that was widely used in the field of statistics and was born around 1980. It can be regarded as an implementation of the S language. The S language, developed by AT&T Bell LABS, is an interpretive language used for data exploration, statistical analysis and graphing

Reprints and Permissions

Explore More Articles

Malnutrition RiskMachine LearningLogistic RegressionInflammatory MarkersESPEN CriteriaGLIM CriteriaNutritional InterventionModel Validation