This study aims to develop and externally validate a web-based system integrating machine learning models for early diagnosis and clinical phenotyping of pneumonia-associated ARDS to facilitate precision treatment.
Research Article
This study aims to develop and externally validate a web-based system integrating machine learning models for early diagnosis and clinical phenotyping of pneumonia-associated ARDS to facilitate precision treatment.
Acute respiratory distress syndrome (ARDS) is a highly heterogeneous disease with clinical manifestations that may overlap with severe pneumonia, posing challenges for accurate differentiation. Therefore, early prediction and bedside rapid subtype clustering of ARDS patients are urgently needed. This study aims to develop a web-based system, which includes validated models of early bedside diagnosis and clinical subgroup classification, for predicting the development and phenotypes of pneumonia-associated ARDS. Diagnostic and subgroup models were developed and validated from the two large databases, Medical Information Mart for Intensive Care IV (MIMIC-IV) and Telehealth Intensive Care Unit (eICU) and were incorporated into a web-based prediction system. Data from patients with pneumonia hospitalized for more than 24 h between 2008 and 2019 were analyzed. The MIMIC-IV derivation cohort included 24,987 patients with pneumonia (14,121 with pneumonia-associated ARDS); the eICU verification cohort included 20,676 patients with pneumonia (9946 with pneumonia-associated ARDS). In diagnosis, the stacking method of machine learning performed best with an AUC of 0.919, an accuracy of 70.00%, a precision of 69.88% and a recall of 82.27% in the MIMIC-IV derivation cohort. The AUC, accuracy, precision, and recall of the eICU validation cohort were 0.915, 70.87%, 69.70% and 69.70% respectively. Pneumonia-associated ARDS was classified into three clinical phenotypes with different clinical characteristics and outcomes, all of which responded differently to treatment. Among patients in clusters 0 and 1, the in-hospital mortality rates were higher among those who received early corticosteroid treatment than among those who did not, whereas among patients in cluster 2, the in-hospital mortality rate was lower among those who received corticosteroids than among those who did not. We performed a web transformation of the diagnosis prediction and clinical subgroup classification of pneumonia-associated ARDS. Our web-based models of early bedside diagnosis and clinical subgroup classification of pneumonia-associated ARDS may assist clinicians in diagnosing and treating the disease and in promoting individualized precision treatment.
Acute respiratory failure, especially acute respiratory distress syndrome (ARDS) after lung infection, is a common, devastating problem encountered in critically ill patients. Studies have shown that the incidence of ARDS is as high as 10% among patients in intensive care unit (ICU)1, and the mortality rate is approximately 40%2,3. Severe pneumonia is widely considered to be the main cause of ARDS4. Because the clinical symptoms of severe pneumonia and ARDS are similar, it is often difficult to distinguish ARDS from severe pneumonia. Therefore, early prediction of the development of ARDS in cases of pneumonia may reduce the incidence of ARDS and rates of mortality5. In addition, because ARDS is a highly heterogeneous disease6, early and correct subgroup classification of ARDS can enable precision medicine. Such classification is also one of the main directions of research on respiratory critical illness throughout the world7, the goal is to improve the effectiveness of targeted subgroup intervention.
At present, the independent risk factors for ARDS have been studied extensively8, but few have focused on the prediction of ARDS in patients with pneumonia. Moreover, no studies have been conducted on specific clinical phenotypes of pneumonia-associated ARDS; most phenotype studies have focused on the entire population of patients with ARDS. Calfee et al.9 classified ARDS into hyperinflammatory and hypoinflammatory phenotypes. To define these biological phenotypes, plasma biomarkers must be used as categorically defining variables, but such an investigation is not readily available at the bedside. In one study, readily available clinical indicators were used to classify ARDS into three clinical phenotypes, whose responses to randomized interventions were also assessed, but the study was based on the entire population of patients with ARDS10. Defining the clinical phenotypes of ARDS according to different causes, such as pneumonia, may yield more refined and accurate results.
The aims of this study were to use machine learning to construct a predictive model of pneumonia-associated ARDS with early clinical data; to use early and easily available clinical data to classify pneumonia-associated ARDS into clinical phenotypes and to explore differences in their clinical characteristics, outcomes, and treatment responses; and to implement the prediction model and classification model as a Web-based application that would assist clinicians in the diagnosis and treatment of pneumonia-associated ARDS and promote further research.
Access restricted. Please log in or start a trial to view this content.
This study accessed the Medical Information Mart for Intensive Care IV (MIMIC-IV) Database11(Version 1.0, PhysioNet: https://physionet.org/content/mimiciv/1.0/) and Telehealth Intensive Care Unit (eICU) Database12(Version 2.0, PhysioNet: https://physionet.org/content/eicu-crd/2.0/) after completing the Protecting Human Research Participants examination (Record ID: 44151052). This study was conducted in accordance with the principles of the Declaration of Helsinki (2013), and patients had provided consent for their data to be captured in the two databases. Ethical approval was waived for this study because the data in the eICU and MIMIC-IV databases were fully anonymized (no personal identifiers retained).
Materials and tools
Data Sources: MIMIC-IV Database: Version 1.0, single-center open-access registry containing 76,540 ICU admissions (2008-2019), accessed through PhysioNet. eICU Database: Version 2.0, multicenter database containing >200,000 electronic medical records from 335 units at 208 U.S. hospitals (2014-2015), accessed through PhysioNet. Software & Execution Environment: RapidMiner Studio: Version 9.10.001 (execution environment: Windows 10 Pro 64-bit), used for model building (classification/clustering) and feature selection; IBM SPSS Statistics: Version 23.0 (execution environment: Windows 10 Pro 64-bit), used for statistical analysis and missing value imputation; Java Development Kit (JDK): Version Java SE 8u381 (execution environment: Windows 10 Pro 64-bit), used for Web application development; supporting IDE: Eclipse IDE 2023-09; Apache Tomcat: Version 9.0.85, used for deploying the Web-based application.
Study design and settings
The concept of this study is illustrated in Supplementary Figure 1. This study analyzed data from two large sources, the Medical Information Mart for Intensive Care IV (MIMIC-IV) and the Telehealth Intensive Care Unit (eICU) databases, to construct a predictive model of pneumonia-associated ARDS and classify affected patients according to clinical phenotypes. This study followed the guidelines of the Transparent Reporting of a Multivariable Prediction Model for Individual Prognosis or Diagnosis (TRIPOD).
Study samples
This study's training and testing data were obtained from MIMIC-IV. The source of external verification was the multicenter eICU database. This study used the International Classification of Diseases (Ninth and Tenth Revision) to extract data about patients with pneumonia and pneumonia-associated ARDS. Patients who were admitted for less than 24 h were excluded.
Predictors
After evaluating data availability and the rate of missing clinical variables in the MIMIC-IV and eICU datasets, 52 variables for pneumonia-associated ARDS were selected as candidate training variables to be used in machine learning, including patients' demographic characteristics, laboratory findings, and scores of illness severity. This study used a weight-by-correlation algorithm (based on the correlation weight between variables and outcomes) to screen for key predictors (variables that are ultimately included in the models) of the 52 candidate variables and then selected subgroup clustering factors from the key predictors.
To do this, open RapidMiner Studio, create a new process, and add the Weight by Correlation operator by clicking on Process > Operators > Feature Selection > Weight by Correlation. Set the parameter: Correlation threshold = 0.04 (retain variables with weight > 0.04). To force the model to account for early stages in the disease and timeliness, the maximum and minimum values of laboratory indicators were obtained within 24 h after admission. Because the Acute Physiology Score III (APSIII) was not included in the eICU database, scores on the Acute Physiology and Chronic Health Evaluation IV (APACHE IV) were used instead. This study cleaned the data and interpolated the missing values by using the multiple imputation method and standardized the input variables when developing prediction and sub-phenotyping models.
Statistical analysis
Open SPSS 23.0, import the preprocessed dataset, and select Analyze > Descriptive Statistics > Explore. Check Normality Plots with Tests to perform the Kolmogorov-Smirnov test for continuous variables. Skewed distributed variables are reported as median (interquartile range, IQR); categorical variables are reported as count (percentage, %). To compare continuous variables between groups, the Mann-Whitney U test or the Kruskal-Wallis's test was used, as appropriate. To compare categorical variables between groups, Pearson's chi-squared test or Fisher's exact test was used, as appropriate.
Model development and web presentation
For model development, this study divided the MIMIC-IV derivation cohort into a training set (90% of the full sample randomly chosen for model development and hyperparameter tuning) and a test set (10% of the full sample randomly chosen for internal testing); this study performed external verification in eICU. This study implemented machine learning with five methods: decision tree, logistic regression, naive bayes, stacking (Base learners were the logistic regression, naive bayes, and random forest methods; the stacking learner was the decision tree method), and random forest. Decision tree: click on Process > Operators > Modeling > Classification > Decision Tree with Algorithm = C4.5, Minimum leaf size = 5. For logistic regression: click on Process > Operators > Modeling > Classification > Logistic Regression with Regularization strength = 0.1, Maximum iterations = 100. For Naive Bayes: click on Process > Operators > Modeling > Classification > Naive Bayes with Kernel type = Gaussian. For Random forest: click on Process > Operators > Modeling > Classification > Random Forest with Number of trees = 100, Maximum depth = 20. For Stacking: click on Process > Operators > Modeling > Ensemble > Stacking with Base learners = Logistic Regression + Naive Bayes + Random Forest; Stacking learner = Decision Tree. This study measured the prediction performance of the developed model by computing the area under the receiver operating characteristic curve (AUC), accuracy, precision, and recall.
In constructing the classification of pneumonia-associated ARDS clinical subgroups, this study used k-means clustering with key predictors. To determine the optimal number of cluster k, the Gap statistics' value in the Gap statistics plot was examined, which suggested the optimal number of clusters. This study analyzed the clinical characteristics and outcomes of different phenotypes, and we compared the differences in rates of in-hospital mortality among patients with different clusters who did and did not receive early low- to medium-dose corticosteroid treatment.
In RapidMiner Studio, select File > Export Model and save the diagnosis prediction and subgroup classification models as .model files. Use JDK Java SE 8u381 and Eclipse IDE 2023-09 to write Web pages in Java language; import the .model files into the project. This study used the Java language to develop a Web page in which the diagnostic model and clinical subgroup clustering model were embedded, and we uploaded the model online.
Access restricted. Please log in or start a trial to view this content.
Participants
The MIMIC-IV database included data from 24,987 patients with pneumonia, of whom 14,121 had pneumonia-associated ARDS (Table 1). The eICU database included data from 20,676 patients with pneumonia, of whom 9946 had pneumonia-associated ARDS (Supplementary Table 1).
Establishment and verification of pneumonia-associated ARDS prediction model
We used the data of the MIMIC-IV cohort to construct a diagnostic ...
Access restricted. Please log in or start a trial to view this content.
As we know, this is the first diagnostic model and clinical subgroup classification model using machine learning to report ARDS in pneumonia patients, and the largest study to report the diagnosis and clinical subgroup classification of pneumonia-associated ARDS. In this study, we derived and validated two machine learning-based models and translated them into web-based applications for clinical practice and subsequent research. In the eICU validation cohort, the prediction of which patients with pneumonia would develop ...
Access restricted. Please log in or start a trial to view this content.
The authors declare that they have no competing interests.
| Name | Company | Catalog Number | Comments |
|---|---|---|---|
| Apache Tomcat | Apache software foundation | Version 9.0.85 | |
| Eclipse IDE | Eclipse | 2023-09 | |
| Java Development Kit | Java | Version Java SE 8u381 | |
| RapidMiner Studio | Altair Engineering Inc. | Version 9.10.001 | |
| SPSS Statistics | IBM | Version 23.0 |
Access restricted. Please log in or start a trial to view this content.
Request permission to reuse the text or figures of this JoVE article
Request Permission