The retrospective review of both cohorts was approved by the Taikang Tongji (Wuhan) Hospital (approval No. 128/06/2025). All data were anonymized before analysis, and informed consent was waived. Clinical trial registration was not required because both cohorts were retrospective, all exposures and follow-up strategies had occurred as part of routine clinical care before data extraction, and no prospective enrollment or research-directed assignment of interventions was undertaken.
Study design
This study used a two-stage retrospective cohort design. The cohort treated between January 2022 and January 2023 was used for model development and bootstrap internal validation, whereas the cohort treated between January and December 2024 was used to retrospectively compare outcomes between patients who received model-informed, risk-guided follow-up and those who received routine, fixed-interval follow-up. The second cohort evaluated the documented care pathway and did not constitute external validation of model performance. Figure 1 presents patient selection for both retrospective cohorts.
Study population
Retrospective cohort
The model-development cohort comprised 124 patients with early gastrointestinal cancer who underwent ESD in the Department of Gastroenterology between January 2022 and January 2023. The inclusion criteria were: (1) pathologically confirmed early esophageal, gastric, or colorectal cancer, classified as T1a or T1b without vascular invasion; (2) first-time ESD with complete R0 resection; (3) postoperative survival of at least 6 months; and (4) complete clinical and nursing records, including the specified dietary, medication, psychological, and follow-up information. The exclusion criteria were: (1) distant metastasis or multiple primary malignant tumors; (2) neoadjuvant chemoradiotherapy; (3) a severe postoperative complication requiring reoperation within 3 months; (4) cognitive impairment or mental illness that precluded follow-up cooperation; and (5) loss to follow-up or incomplete data. No formal a priori sample-size calculation was performed because this was a retrospective model-development cohort, and all eligible records available during the prespecified period were included. For binary clinical prediction models, however, sample-size adequacy depends on the number of outcome events, the number of candidate predictor parameters, anticipated model performance, and the need to limit overfitting, rather than on the total sample size alone. The cohort contained 22 recurrence events across seven predictor parameters, corresponding to 3.1 events per parameter; given this low event-to-parameter ratio, the model was considered vulnerable to coefficient instability and overfitting11,12.
Retrospective comparative cohort
The retrospective comparative cohort comprised 84 patients with early gastrointestinal cancer who underwent ESD between January and December 2024 and had complete documentation of the follow-up strategy received and the prespecified 12-month outcomes. The inclusion criteria were: (1) pathologically confirmed early esophageal, gastric, or colorectal cancer; (2) ESD performed during the prespecified study period; (3) complete documentation of the follow-up strategy received; and (4) complete 12-month outcome data. The exclusion criteria were: (1) distant metastasis or multiple primary malignant tumors; (2) neoadjuvant chemoradiotherapy; (3) a severe postoperative complication requiring reoperation within 3 months; (4) cognitive impairment or mental illness that precluded follow-up participation; and (5) incomplete follow-up or outcome data. Based on the follow-up strategy documented in the medical and nursing records, 42 patients were classified into the model-informed, risk-guided follow-up group and 42 into the routine, fixed-interval follow-up group. No random allocation, allocation concealment, or research-directed assignment was performed. The cohort size was determined by the number of eligible records with complete 12-month outcome data during the prespecified period; consequently, no prospective power-based sample size calculation was applicable.
Data collection
Retrospective data extraction
Baseline variables included age, sex, body mass index (BMI), and pathological type. Nursing-relevant postoperative indicators were abstracted using the prespecified definitions in the database: high-salt or spicy food intake at least twice per week during the first 3 postoperative months; incomplete daily wound-exudate documentation during the first postoperative month; a Hospital Anxiety and Depression Scale score (HADS)13 ≥11 at 1 month; at least two missed doses of sucralfate per week during the first 3 postoperative months; at least two missed follow-up appointments during the first 6 months; smoking at least 10 cigarettes per day for at least 6 months; and a mean fasting blood glucose concentration ≥7.0 mmol/L during the first 3 postoperative months among patients with diabetes. The outcome was recurrence within 12 months, defined as local recurrence or metachronous cancer confirmed by endoscopic biopsy and histopathology. Because the assessment windows for several postoperative indicators overlapped with the recurrence-observation period, temporal precedence could not be verified for every patient; these variables were therefore analyzed as postoperative correlates of recurrence, and the resulting model was not interpreted as a baseline prognostic tool.
Model development and internal validation
Seven binary nursing-relevant variables—high-salt or spicy food intake, incomplete wound-exudate documentation, HADS score ≥11, irregular sucralfate use during the first 3 postoperative months, at least two missed follow-up appointments, smoking exposure, and uncontrolled diabetes—were entered simultaneously into a multivariable logistic regression model. Logistic regression was selected because the outcome was binary and the study sought an interpretable probability equation that could be applied in routine records. Alternative statistical or machine-learning algorithms were not compared because only 22 recurrence events were available, making reliable algorithm development and tuning unlikely. Multicollinearity was assessed using variance inflation factors, with values >5 regarded as unacceptable.
Model discrimination was assessed using the area under the receiver operating characteristic curve, with the 95% confidence interval calculated using the DeLong method. The operating threshold was selected by maximizing the Youden index, after which sensitivity and specificity were calculated. Calibration was evaluated using the Hosmer-Lemeshow goodness-of-fit test, the Brier score, and a calibration plot based on predicted risk quintiles. Apparent and bootstrap-corrected calibration were displayed, with optimism estimated from 1,000 bootstrap resamples in which the full model was refitted; the same bootstrap procedure was used to estimate optimism in the AUC14. Regression coefficients, odds ratios, confidence intervals, the intercept, and the complete probability equation were reported to permit independent calculation of the model score.
Intervention protocol
Risk-guided follow-up group: Medical and nursing records indicated that patients in this group had received follow-up and nursing support linked to the model-derived recurrence probability. The probabilities were categorized using the operational thresholds retained in the original institutional pathway: low risk, <20%; intermediate risk, 20% to <40%; and high risk, ≥40%. These thresholds reflected local clinical practice and multidisciplinary review rather than cutoffs derived from a published guideline, previous validation study, or formal Delphi consensus; they were used to determine nursing intensity and were distinct from the ROC-derived threshold of 0.1484, which was used solely to summarize model sensitivity and specificity.
The pathway was implemented through routine follow-up by gastroenterology nursing staff. Low-risk patients received one annual review, a standardized booklet covering diet, daily lifestyle, and warning symptoms, and brief reinforcement messages every 3 months. Intermediate-risk patients received semi-annual review, including one outpatient endoscopic reassessment, together with monthly nurse-led telephone contacts, review of food diaries, individualized dietary plans, and reinforcement of medication adherence. High-risk patients received an eight-session structured cognitive-behavioral support program delivered once weekly, video-based wound-care instruction, daily recording of bleeding, pain, abdominal distension, and other warning symptoms, weekly submission of the records to the nursing team, monthly family meetings to clarify supervision responsibilities, and coordinated diabetes management with endocrinology consultation and scheduled glucose monitoring.
The archived pathway described the psychological component as cognitive behavioral therapy (CBT) but did not retain the provider’s mental-health credentials or a treatment manual; it is therefore reported conservatively as structured cognitive-behavioral support rather than formal therapist-delivered CBT. The risk-guided follow-up pathway was developed primarily as a local institutional pathway, drawing on the hospital’s routine post-ESD procedures and published recommendations supporting scheduled endoscopic surveillance after curative ESD15,16. The published guidance informed the general surveillance framework but did not specify the model-based probability thresholds or the nursing actions assigned to each risk tier. These elements were developed locally by mapping nursing-relevant indicators to corresponding care actions and were reviewed by the hospital’s multidisciplinary clinical team without a formal Delphi process. Delivery was documented in the existing nursing and follow-up records; no separate fidelity checklist, prespecified fidelity threshold, or independent fidelity assessment was retained (Table 1).
Routine follow-up group: Records indicated that patients in this group had received fixed-interval follow-up, including endoscopic examinations at 6 months, 1 year, and 2 years after surgery, as well as routine health education without model-informed, individualized guidance.
Outcome indicators
Model performance indicators
Model performance was summarized using the AUC and its 95% confidence interval; the selected probability threshold; sensitivity, specificity, and the Youden index; the Hosmer-Lemeshow statistic and P value; the Brier score; bootstrap-estimated optimism; and the optimism-corrected AUC.
Comparative outcome in the retrospective comparative cohort
The principal comparative outcome was recurrence confirmed by endoscopic biopsy and histopathology within 12 months after ESD. Results were reported as absolute numbers and percentages, together with the risk difference, relative risk, odds ratio, and their 95% confidence intervals. Exact recurrence dates were not retained in the 2024 analysis dataset; therefore, recurrence-free survival, Kaplan-Meier curves, Cox regression, and recurrence-time distributions were not analyzed.
Nursing management quality indicators
Nursing compliance was evaluated at 1, 6, and 12 months and included dietary compliance, defined as high-salt or spicy food intake fewer than twice per week; wound-care execution, defined as completion of standardized daily wound records; medication adherence, defined as fewer than two missed doses per week; and follow-up adherence, defined as no missed scheduled appointments.
Psychological status was assessed at the same time points using the Chinese version of the 14-item Hospital Anxiety and Depression Scale, which comprises 7-item anxiety and depression subscales, each scored from 0 to 21, with higher scores indicating greater symptom burden. Anxiety, depression, and total scores were analyzed as continuous variables. The Chinese version of the HADS has demonstrated satisfactory psychometric properties, including construct validity, internal consistency, and concurrent validity, in a multicenter sample of Chinese cancer patients13.
Self-management was assessed using the Gastrointestinal Disease Self-Management Scale recorded in the original study database. The retained scale comprised four scored domains—dietary management, symptom monitoring, medication management, and emotional regulation—and a total score calculated as the sum of the four domain scores; higher scores indicated better self-management. The original study documentation recorded a Cronbach’s α of 0.85. Item-level responses, the original instrument manual, and an independent external validation reference were not retained; the scale was therefore treated as a study instrument rather than described as a fully externally validated measure.
Nursing satisfaction was assessed at 12 months using the hospital-developed 100-point satisfaction scale recorded in the study database. Scores were categorized as very satisfied (≥90), satisfied (80–89), fair (60-79), or dissatisfied (<60). The original study documentation reported a Cronbach’s α of 0.87, but item-level responses and the original scale-development records were not retained; therefore, satisfaction findings were considered exploratory.
Medical resource efficiency indicators
Resource use was evaluated over the 12-month observation period from a hospital direct medical resource perspective. Examination costs were expressed in nominal Chinese yuan as recorded in the hospital database and included charges for endoscopic, imaging, and laboratory examinations; no inflation adjustment or discounting was applied because the analysis covered a single 12-month follow-up period. Hospitalizations attributable to recurrence or postoperative complications were identified from the electronic medical records. Follow-up time was analyzed using the composite variable retained in the database, which combined healthcare provider follow-up time with patient travel and consultation time; the individual time components were not stored separately. Because cost and utilization variables may be skewed, between-group mean differences and 95% confidence intervals were estimated using 10,000 nonparametric bootstrap resamples.
Statistical methods
Data were analyzed using SPSS and R. Continuous single-time-point variables were assessed for normality using the Shapiro-Wilk test and summarized as mean ± standard deviation or median (interquartile range), as appropriate. Between-group comparisons used independent-samples t tests or Mann-Whitney U tests. Categorical variables were summarized as n (%) and compared using Pearson’s χ2 test or Fisher’s exact test when expected cell counts were <5. Ordered satisfaction categories were compared using the Mann-Whitney rank-sum test, while the overall satisfaction rate was compared using Fisher’s exact test.
Repeated continuous outcomes were analyzed using generalized estimating equations with a Gaussian distribution, identity link, exchangeable working correlation, and robust standard errors; fixed effects comprised group, time, and the group-by-time interaction. Repeated binary compliance outcomes were analyzed using generalized estimating equations with a binomial distribution and logit link. Group-by-time interaction P values were adjusted within each outcome family using the Benjamini-Hochberg false-discovery-rate procedure, and 12-month between-group differences or odds ratios were reported with 95% confidence intervals.
One missing 6-month self-management total score in the routine-follow-up group was reconstructed as 72 because all four domain scores were available and summed exactly to that total. Recurrence comparisons were reported using the risk difference, relative risk, and odds ratio with 95% confidence intervals. Resource-use mean differences and confidence intervals were estimated using 10,000 nonparametric bootstrap resamples. The seven-variable logistic model was evaluated using the procedures described above, including 1,000 bootstrap resamples for internal validation. Because the comparative cohort was nonrandomized and contained only 18 recurrences, its results were interpreted as unadjusted observational associations rather than causal effect estimates. All tests were two-sided, and p < 0.05 was considered statistically significant.