The study protocol was reviewed and approved by the Ethics Committee of Santai County Hospital of Traditional Chinese Medicine in Mianyang, Sichuan Province (Approval No. 2025003). The study was conducted in accordance with the principles of the Declaration of Helsinki. Written informed consent was obtained from each patient or from a first-degree relative before inclusion in the study. This retrospective cohort study reviewed the clinical records of elderly patients hospitalized for acute exacerbation of chronic obstructive pulmonary disease (AECOPD) in the Department of Respiratory and Critical Care Medicine of Santai County Traditional Chinese Medicine Hospital between May 2023 and May 2025. The study was designed to identify routinely available clinical predictors of pulmonary hypertension (PH) and to develop a nomogram for individualized PH risk estimation. The laboratory instruments, echocardiography system, blood collection supplies, data extraction tools, and statistical software used in this workflow were listed in the Table of Materials.
1. Patient screening and eligibility assessment
Hospitalized patients with a discharge diagnosis of AECOPD during the study period were screened through the inpatient electronic medical record system. Each record was reviewed to confirm that the patient was at least 60 years old and that the diagnosis of chronic obstructive pulmonary disease was consistent with the 2021 Global Initiative for Chronic Obstructive Lung Disease diagnostic criteria15. AECOPD was defined as an acute worsening of respiratory symptoms that required hospitalization and additional treatment. For patients with more than one AECOPD admission during the study period, only the first eligible hospitalization was included.
Patients were included when all of the following criteria were met: age of at least 60 years; hospitalization for AECOPD between May 2023 and May 2025; available transthoracic echocardiography during the index hospitalization; available first fasting venous blood sample after admission; and complete clinical, laboratory, and echocardiographic information required for model construction. Patients were excluded if they had severe hepatic insufficiency, end-stage renal disease, active malignancy, pulmonary hypertension attributable to other known causes, primary left-heart disease, connective tissue disease such as systemic sclerosis, active pulmonary tuberculosis, or missing key variables required for outcome classification or model construction.
Before records were removed because of incomplete information, age, sex, smoking history, drinking history, drug allergy history, hypertension, diabetes mellitus, coronary heart disease, asthma, bronchiectasis, pulmonary infection, emphysema, respiratory failure, heart failure, echocardiographic PH status, complete blood count indices, albumin, creatinine, high-sensitivity C-reactive protein, D-dimer, fibrinogen, B-type natriuretic peptide, and neutrophil-to-lymphocyte ratio were checked. The number of screened and excluded records, exclusion reasons, and final included cases were recorded to construct the patient selection flow diagram (Figure 1).

Figure 1: Patient screening and cohort construction. The flow diagram shows record identification, eligibility assessment, exclusion, and final grouping of elderly patients hospitalized with AECOPD. Patients were classified into the non-PH or PH group according to transthoracic echocardiographic findings. Abbreviations: AECOPD, acute exacerbation of chronic obstructive pulmonary disease; PH, pulmonary hypertension. Please click here to view a larger version of this figure.
2. Clinical grouping and diagnostic definition of pulmonary hypertension
Included patients were classified into a non-PH group or a PH group according to transthoracic echocardiographic findings during hospitalization. The diagnostic framework was based on the 2015 European Society of Cardiology/European Respiratory Society guidelines for pulmonary hypertension16. Resting transthoracic echocardiography was used as the diagnostic basis in this retrospective dataset. Examinations were performed by trained echocardiography physicians using standard parasternal, apical, and subcostal views. Echocardiographic interpretability was reviewed before group assignment because poor acoustic windows are common in patients with chronic obstructive pulmonary disease. A record was considered suitable for PH classification when the tricuspid regurgitation Doppler signal was measurable or when right ventricular size, wall thickness, and systolic function could be evaluated from interpretable cardiac views.
Pulmonary arterial systolic pressure was estimated using peak tricuspid regurgitation velocity and estimated right atrial pressure. Right atrial pressure was estimated from the inferior vena cava diameter and inspiratory collapse when both measurements were available. PH was defined as an estimated pulmonary arterial systolic pressure of at least 40 mmHg. When pulmonary arterial systolic pressure could not be reliably estimated, PH classification required indirect echocardiographic evidence of right ventricular involvement, including right ventricular enlargement, right ventricular hypertrophy, right ventricular systolic dysfunction, interventricular septal flattening, or pulmonary artery enlargement. Echocardiographic records with unclear images, untraceable tricuspid regurgitation Doppler signals, or inconsistent right-heart findings were reviewed by a second echocardiography physician before the final group assignment. Records that remained uninterpretable after review were excluded from the final analysis. In the final cohort, 230 elderly patients met the eligibility criteria; 120 patients were assigned to the non-PH group, and 110 patients were assigned to the PH group.
3. Extraction of demographic and clinical variables
All clinical data were retrieved from the inpatient electronic medical record system using a standardized data collection form. Sex, age, smoking history, drinking history, drug allergy history, hypertension, diabetes mellitus, coronary heart disease, asthma, bronchiectasis, pulmonary infection, emphysema, respiratory failure, and heart failure were recorded. Smoking history was defined as lifetime consumption of at least 100 cigarettes, and drinking history was defined as regular alcohol consumption at least two times per week.
Pulmonary infection was diagnosed by integrating respiratory symptoms, physical signs, chest imaging findings, and, when available, microbiological evidence. Heart failure was diagnosed according to clinical manifestations, physical signs, B-type natriuretic peptide levels, and echocardiographic findings. Comorbidity information was cross-checked against discharge diagnoses, progress notes, imaging reports, laboratory reports, and consultation records, when available.
4. Collection and processing of laboratory variables
The first fasting venous blood sample collected on the morning after hospital admission was used as the laboratory data source. This time point corresponded to the first morning blood draw after admission and before routine daytime treatment adjustment, when this information was available in the medical record. Complete blood count indices, biochemical indices, coagulation markers, inflammatory markers, and cardiac stress markers were extracted from the hospital laboratory information system.
Neutrophil count, lymphocyte count, platelet count, mean platelet volume, and red-cell distribution width coefficient of variation were recorded from the complete blood count report. The neutrophil-to-lymphocyte ratio was calculated as follows:

Albumin, creatinine, D-dimer, fibrinogen, high-sensitivity C-reactive protein, and B-type natriuretic peptide were recorded from the corresponding laboratory reports. Consistent measurement units were used across the dataset: albumin in g/L, creatinine in µmol/L, high-sensitivity C-reactive protein in mg/L, D-dimer in µg/mL, fibrinogen in g/L, neutrophil count, lymphocyte count, and platelet count in ×109/L, mean platelet volume in fL, and B-type natriuretic peptide in pg/mL.
5. Data cleaning and preparation
The dataset was inspected before statistical analysis. Each included record was checked for eligibility, group assignment, laboratory completeness, echocardiographic interpretability, and availability of candidate predictors. Direct personal identifiers were removed, and each patient was assigned a study identification number. Duplicate records, repeated hospitalizations, inconsistent group labels, impossible dates, and implausible laboratory values were checked against the original electronic medical record and laboratory report. No imputation was performed for key outcome or predictor variables because the final model was based on complete-case analysis.
Continuous variables were reviewed for distributional characteristics. Normally distributed variables were summarized as mean ± standard deviation; non-normally distributed variables as median with interquartile range; and categorical variables as counts and percentages. Before multivariable modeling, clinically related variables were examined for redundancy. Neutrophil count and lymphocyte count were reviewed together with the neutrophil-to-lymphocyte ratio; pulmonary infection was reviewed together with neutrophil count, high-sensitivity C-reactive protein, and the neutrophil-to-lymphocyte ratio; heart failure was reviewed together with B-type natriuretic peptide and echocardiographic cardiac findings; D-dimer and fibrinogen were reviewed as coagulation-related variables; and platelet count, mean platelet volume, and red-cell distribution width coefficient of variation were reviewed as hematological variables.
6. Univariate comparison between groups
Demographic characteristics, comorbidities, laboratory indicators, and echocardiographic grouping variables were compared between the non-PH and PH groups. An independent-sample t-test was used for normally distributed continuous variables. The Mann-Whitney U test was used for non-normally distributed continuous variables. The χ2 test or Fisher’s exact test was used for categorical variables, depending on expected cell counts. A two-sided P-value less than 0.05 was considered statistically significant. Variables with significant between-group differences were considered for multivariable evaluation together with clinically relevant predictors.
7. Multivariable logistic regression analysis
Binary logistic regression was used to identify independent predictors associated with PH in elderly patients with AECOPD. PH status was used as the dependent variable. Continuous predictors were entered in their original measurement units, and categorical predictors were coded as binary variables. Candidate variables were selected according to univariate results and clinical relevance. Variables with P < 0.05 in the univariate comparison were first identified and then reviewed for redundancy, interpretability, and clinical overlap.
The initial candidate variables included high-sensitivity C-reactive protein, D-dimer, red-cell distribution width coefficient of variation, platelet count, mean platelet volume, neutrophil-to-lymphocyte ratio, B-type natriuretic peptide, coronary heart disease, and asthma history. Neutrophil count and lymphocyte count were not entered together with neutrophil-to-lymphocyte ratio because they were components of the ratio. Pulmonary infection and heart failure were evaluated during redundancy assessment. To examine whether their exclusion affected model interpretation, a sensitivity analysis was performed by adding pulmonary infection and heart failure to the primary predictor set. Regression estimates, discrimination, calibration, and clinical net benefit were compared between the primary and sensitivity models.
Multicollinearity among candidate predictors was assessed using variance inflation factors. Categorical variables were coded before calculation, and continuous variables were entered in the same units used for regression modeling. A variance inflation factor below 5 was considered to indicate no severe multicollinearity. Regression coefficients, standard errors, Wald χ2 values, odds ratios, 95% confidence intervals, P-values, and variance inflation factors were reported for variables included in the multivariable model.
8. Nomogram construction
The nomogram was constructed from the independent predictors retained in the multivariable logistic regression model. The final predictors were D-dimer, neutrophil-to-lymphocyte ratio, B-type natriuretic peptide, and asthma history. The logistic regression model was converted into a point-based nomogram using R software and the rms package. A data distribution object was generated for the modeling dataset, the logistic regression model was fitted, and regression coefficients were used to assign points to each predictor. For each patient, the points for D-dimer, neutrophil-to-lymphocyte ratio, B-type natriuretic peptide, and asthma history were summed to generate a total score, which was then mapped to the estimated probability of PH.
The nomogram was interpreted using standard point-based procedures described in previous clinical prediction model studies17,18. The patient’s D-dimer value, neutrophil-to-lymphocyte ratio, B-type natriuretic peptide value, and asthma status were located on their corresponding axes. A vertical line was drawn from each predictor value to the point scale. The points were added to obtain the total score, and a vertical line was drawn from the total score scale to the predicted probability axis to obtain the individualized estimated probability of PH. The illustrative patient used in the manuscript was selected from values within or close to the observed clinical distribution of the cohort.
9. Evaluation of model performance
Model discrimination was evaluated using receiver operating characteristic curve analysis and the area under the curve. The area under the curve was reported with a 95% confidence interval. The nomogram was compared with individual predictors and with a model combining quantitative laboratory markers. Calibration was assessed by comparing predicted probabilities with observed PH outcomes. A calibration curve was generated using bootstrap resampling. Calibration intercept, calibration slope, mean absolute error, Brier score, and the Hosmer-Lemeshow test were reported. Bootstrap-corrected calibration intercept and bootstrap-corrected calibration slope were also calculated . The Hosmer-Lemeshow test was interpreted together with the calibration curve and quantitative calibration indices.
Clinical utility was evaluated using decision curve analysis. Net benefit was plotted across threshold probabilities and compared with treat-all and treat-none reference strategies. The threshold probability range in which the model provided higher net benefit than the reference strategies was identified.
10. Internal validation and reproducibility
Internal validation was performed using bootstrap resampling with 1,000 replicates. In each bootstrap replicate, the model was refitted and tested to estimate optimism in model performance. Optimism-corrected discrimination and calibration indices were calculated and reported together with the apparent model performance.
Statistical analyses were performed using standard statistical software and R version 4.5.2. R packages were used for logistic regression modeling, nomogram construction, receiver operating characteristic curve analysis, calibration assessment, internal validation, and decision curve analysis. The nomogram was constructed using the rms package in R.