需要JoVE订阅才能观看此内容。 请登录或开始免费试用

方法文章

基于机器学习模型的三种淋巴结分期系统在结直肠印戒细胞癌中的预测性能比较

1.2K 次观看

DOI:

10.3791/67941

2025年4月18日

* These authors contributed equally

本文内容

摘要

本研究利用机器学习模型和竞争风险分析,评估了结直肠印戒细胞癌患者的预后系统。研究发现,与pN分期相比,阳性淋巴结的对数比(LODDS)是更优的预后预测指标,具有良好的预测性能,可通过可靠的生存预测工具辅助临床决策。

摘要

淋巴结状态是患者重要的预后预测指标;然而,结直肠印戒细胞癌(SRCC)的预后研究受到的关注较为有限。本研究结合机器学习模型(随机森林、XGBoost 和神经网络)与竞争风险模型,探讨阳性淋巴结的对数比值(LODDS)、淋巴结比率(LNR)以及 pN 分期在 SRCC 患者中的预后预测能力。相关数据来源于监测、流行病学和最终结果(SEER)数据库。对于机器学习模型,首先通过单变量和多变量 Cox 回归分析确定癌症特异性生存(CSS)的预后因素,随后应用三种机器学习方法——XGBoost、RF 和 NN——以确定最优的淋巴结分期系统。在竞争风险模型中,采用单变量和多变量竞争风险分析识别预后因素,并构建列线图以预测 SRCC 患者的预后。采用受试者工作特征曲线下面积(AUC-ROC)和校准曲线评估模型的预测性能。本研究共纳入 2,409 例 SRCC 患者。为验证模型的有效性,另纳入一个包含 15,122 例非 SRCC 结直肠癌患者的队列进行外部验证。机器学习模型与竞争风险列线图在预测生存结局方面均表现出较强的性能。与 pN 分期相比,LODDS 分期系统展现出更优的预后判断能力。经评估,机器学习模型与竞争风险模型均具有优异的预测性能,表现为良好的区分度、校准度和可解释性。本研究结果可能有助于指导患者的临床决策。

引言

结直肠癌(CRC)是全球第三大最常见的恶性肿瘤1,2,3。印戒细胞癌(SRCC)是CRC的一种罕见亚型,约占病例的1%,其特征是细胞内含有大量黏液,导致细胞核被挤压至一侧1,2,4。SRCC常发生于较年轻的患者,女性发病率较高,且在诊断时多处于晚期肿瘤阶段。与结直肠腺癌相比,SRCC分化程度更差,远处转移风险更高,5年生存率仅为12%–20%5,6。建立准确有效的SRCC预后预测模型,对于优化治疗策略和改善临床结局至关重要。

本研究旨在利用先进的统计方法(包括机器学习(ML)和竞争风险模型)为印戒细胞癌(SRCC)患者构建一个稳健的预后模型。这些方法能够处理临床数据中的复杂关系,提供个体化的风险评估,并在预测准确性方面优于传统方法。机器学习模型(如随机森林、XGBoost和神经网络)在处理高维数据和识别复杂模式方面表现出色。已有研究表明,人工智能模型在结直肠癌生存结局预测中具有显著效果,凸显了机器学习在临床应用中的潜力7,8。与机器学习相辅相成,竞争风险模型可针对多种事件类型(如癌症特异性死亡与其他死因)进行分析,从而优化生存分析。与Kaplan-Meier估计器等传统方法不同,竞争风险模型能够在存在竞争风险的情况下准确估计事件的边际概率,提供更精确的生存评估8。将机器学习与竞争风险分析相结合,可提升预测性能,为SRCC个体化预后工具的开发提供有力框架9,10,11

淋巴结转移显著影响结直肠癌(CRC)患者的预后和复发。尽管TNM分期系统中的N分期评估至关重要,但淋巴结检出不足——在48%至63%的病例中均有报道——可能导致疾病被低估。为解决这一问题,已提出淋巴结转移率(LNR)和阳性淋巴结对数比(LODDS)等替代方法。LNR是指阳性淋巴结(PLNs)与总淋巴结(TLNs)数量的比值,受TLN数量影响较小,是结直肠癌的预后因素。LODDS是阳性淋巴结与阴性淋巴结(NLNs)数量比值的对数形式,在胃部印戒细胞癌和结直肠癌中均显示出更优的预测能力10,11。机器学习在肿瘤学中的应用日益广泛,相关模型已在多种癌症(包括乳腺癌、前列腺癌和肺癌)中改善了风险分层和预后预测12,13,14。然而,其在结直肠印戒细胞癌中的应用仍较为有限。

本研究旨在通过将阳性淋巴结的对数比(LODDS)与机器学习及竞争风险模型相结合,构建一个全面的预后预测工具。通过评估LODDS的预后价值并利用先进的预测技术,本研究旨在加强临床决策能力,改善SRCC患者的治疗结局。

访问受限。请登录或开始试用以查看此内容。

方案

本研究未涉及伦理审批及参与同意。研究使用的数据来源于数据库。我们纳入了2004年至2015年间诊断为结直肠印戒细胞癌的患者,以及其他类型的结直肠癌患者。排除标准包括生存时间少于一个月的患者、临床病理信息不完整的患者,以及死因不明确或未指明的病例。

1. 数据获取

  1. 下载SEER。从SEER数据库网站(http://seer.cancer.gov/about/overview.html)获取统计软件8.4.3版本。登录软件后,点击 Case List Session > Data 并选择 Incidence SEER Research Plus Data, 17 Registries, Nov 2021 Sub (2000-2019) database.
  2. 点击 Selection > Edit and choose {Race, Sex, Year Dx. Year of diagnosis} = '2004', '2005', '2006', '2007', '2008', '2009', '2010', '2011', '2012', '2013', '2014', '2015' AND {Site and Morphology. Site recode ICD-O-3/WHO 2008} = '8490/3'.
  3. 然后点击 Table并在可用变量界面中,选择 Age recode with single ages and 100+, Sex, Marital, Site recode ICD-O-3/WHO 2008, CS tumor size, Regional nodes_examined(1988+), Regional nodes_positive(1988+), Derived AJCC Stage Group, 6th ed (2004-2015), Derived AJCC T, 6th ed (2004-2015), Derived AJCC N, 6th ed (2004-2015), Derived AJCC M, 6th ed (2004-2015), CEA, Radiation recode, Chemotherapy recode (yes, no/unk), SEER cause-specific death classification, Vital status recode (study cutoff used), Survival months, Year of diagnosis.
  4. 最后,点击 Output命名数据,并点击 Execute 以输出并保存数据。详细的纳入流程如图1所示。 Figure 1.
  5. 从结直肠癌患者中下载数据,排除印戒细胞癌病例,用于后续外部验证。点击“选择” > Edit and choose {Race, Sex, Year Dx. Year of diagnosis} = '2004', '2005', '2006', '2007', '2008', '2009', '2010', '2011', '2012', '2013', '2014', '2015' AND {Primary Site - labeled} = 'C18-C20'重复步骤 1.3 和 1.4,以获取临床病理学信息,并排除 {Site and Morphology. Site recode ICD-O-3/WHO 2008} = '8490/3' 从下载的文件中。
  6. 为便于比较,需对若干变量进行处理。采用淋巴结比率(LNR)与阳性淋巴结对数比(LODDS)两种方法对淋巴结状态进行分类。
    1. 将淋巴结比率(LNR)定义为阳性淋巴结(PLNs)数量与总淋巴结(TLNs)数量之比。使用以下公式计算对数阳性淋巴结比值(LODDS):
      loge(number of PLNs + 0.5) / (number of negative lymph nodes (NLNs) + 0.5)
      此处添加 0.5 是为了防止结果出现无穷值。LNR、LODDS 和肿瘤大小的截断值是使用 X-tile 软件(版本 3.6.1),基于最小 P 值法确定的。
  7. 打开 X-tile 软件,点击 File > Open并选择数据文件将其导入软件。数据加载完成后,映射变量:Censor 对应生存状态,Survival time 对应生存时间,marker1 为待分析变量,确保数据匹配正确。
  8. 然后,点击 Do > Kaplan-Meier > Marker1 执行Kaplan-Meier生存分析并生成生存曲线。根据Kaplan-Meier生存曲线的分离程度、统计显著性(例如p值)以及临床相关性,确定最佳截断值,最终记录或导出分析结果。
    1. 将 LNR 分为三组:LNR 1(≤0.16), LNR 2 (0.16 - 0.78), and LNR 3 (≥ 0.78)。根据LODDS将患者分为三组:LODDS 1(≤ -1.44)、LODDS 2(-1.44 - 0.86)和LODDS 3(≥ 0.86)。将肿瘤大小分为三类:≤ 3.5 cm、3.5 - 5.5 cm和≥ 5.5 cm。将年龄从连续变量转换为分类变量。将患者初诊时的年龄分类为≥ 60岁和< 60岁。根据印戒细胞癌(SRCC)肿瘤的分布情况,将肿瘤位置分类为右半结肠、左半结肠和直肠。右半结肠包括盲肠、升结肠、肝曲和横结肠,而左半结肠包括脾曲、降结肠、乙状结肠和直肠乙状结肠交界处。在本研究中,将总共2409例符合条件的SRCC患者数据随机分配至训练队列。≤ -1.44), LODDS 2 (-1.44 - 0.86), and LODDS 3 (≥ 0.86).
    2. 将肿瘤大小分为三类: ≤ 3.5 cm, 3.5 - 5.5 cm, and ≥ 5.5 cm。将年龄从连续变量转换为分类变量。将患者' 初诊时年龄分为 ≥60 years and <60岁。根据印戒细胞癌(SRCC)肿瘤的分布情况,将肿瘤位置分类为右半结肠、左半结肠和直肠。右半结肠包括盲肠、升结肠、肝曲和横结肠,而左半结肠包括脾曲、降结肠、乙状结肠和直乙交界处。
  9. 在本研究中,将2409例符合条件的SRCC患者数据按7:3的比例随机分配至训练队列(N = 1686)和验证队列(N = 723)。使用以下代码进行随机拆分,源数据.csv来自SEER数据库。随机拆分后生成的文件将用于后续分析。
    library(caret)
    data <- read.csv("data.csv")
    set.seed(123)
    train_indices <- createDataPartition(data$variable, p = 0.7, list = FALSE)
    train_data <- data[train_indices, ]
    test_data <- data[-train_indices, ]
    write.csv(train_data, "traindata.csv", row.names = FALSE)
    write.csv(test_data, "testdata.csv", row.names = FALSE)

2. 机器学习模型的构建与验证

  1. 下载 RStudio (2024.04.2+764) 和 R 软件 (4.4.1)。打开 RStudio 以运行 R 软件。点击 新建文件,并选择 R 脚本 以创建新的 R 编程界面。在代码编辑器中输入相关代码,然后点击 运行 以执行代码。
  2. 使用以下代码通过 Cox 回归分析筛选机器学习模型中包含的变量。此外,探讨 LODDS、LNR 和 pN 分期对 SRCC 患者癌症特异性生存(CSS)的影响。traindata.csv 是从 SEER 数据库获取的数据。
    library("survival")
    library("survminer")
    library("rms")
    library("dplyr")
    data <- read.csv("traindata.csv")
    data$time=as.numeric(data$time)
    data$status=as.numeric(data$status)
    variables <- c("Sex", "Age", "Race", "Marital", "Stage", "T", "N", "M","Tumor_size", "LNR", "LODDS", "CEA","Radiation", "Chemotherapy", "Site")
    data <- data %>%
    mutate(across(all_of(variables), as.factor))
    cox=coxph(Surv(time, status) ~ data$T, data = data)
    cox$coefficients
    pval=anova(cox)$Pr[2]
    clean_data=data[,c(1:12, 14:18)]
    get_coxVariable=function(your_data,index){cox_list=c() k=1
    for (i in 1:index) {mod=coxph(Surv(time, status) ~ your_data[,i],data=your_data) pval=anova(mod)$Pr[2] print(pval) print(colnames(your_data)[i]) if (pval<0.05) {cox_list[k]=colnames(your_data)[i] k=k+1}}return(cox_list)}
    variable_select=get_coxVariable(clean_data,15)
    for(i in 1:15){print(variable_select[i])}
    for (var in variable_select) {formula <- as.formula(paste("Surv(time, status) ~", var))cox_model <- coxph(formula, data = data) print(summary(cox_model))
    ggforest(cox)
    variables <- c("Sex", "Age", "Race", "Marital", "Stage", "T", "N", "M", "Tumor_size", "LNR", "LODDS", "Chemotherapy")
    data <- data %>%
    mutate(across(all_of(variables), as.factor))
    cox=coxph(Surv(time, status) ~ Sex+Age+Race+Marital+T+N+M+Tumor_size+LNR+
    LODDS+Chemotherapy,data = data)
    ggforest(cox,data = data)
    ggplot_forest <- ggforest(cox, data = data)
  3. 使用以下代码比较三种淋巴结分期系统(LODDS、LNR 和 pN 分期)在训练集、验证集和外部验证集中的预后预测能力。
    library(rms)
    library(survival)
    library(survminer)
    library(riskRegression)
    library(gt)
    train_data <- read.csv("train_data123.csv")
    validation_data <- read.csv("test_data123.csv")
    dd <- datadist(train_data)
    options(datadist = "dd")
    model_LNR <- cph(Surv(time, status) ~ LNR, data = train_data, x = TRUE, y = TRUE)
    model_LODDS <- cph(Surv(time, status) ~ LODDS, data = train_data, x = TRUE, y = TRUE)
    model_pN <- cph(Surv(time, status) ~ N, data = train_data, x = TRUE, y = TRUE)
    calculate_performance <- function(model, data) {pred <- predict(model, newdata = data) c_index_result <- concordance(Surv(data$time, data$status) ~ pred) c_index <- c_index_result$concordance aic <- AIC(model) bic <- BIC(model) return(c(C_index = round(c_index, 3), AIC = round(aic, 2), BIC = round(bic, 2)))}
    calculate_performance <- function(model, data) {pred <- predict(model, newdata = data, type = "lp") concordance_result <- concordancefit(Surv(data$time, data$status), x = pred) c_index <- concordance_result$concordance ci_lower <- c_index - 1.96 * sqrt(concordance_result$var) ci_upper <- c_index + 1.96 * sqrt(concordance_result$var) aic <- AIC(model) bic <- BIC(model) return(c(C_Index = round(c_index, 3), CI_Lower = round(ci_lower, 3), CI_Upper = round(ci_upper, 3), AIC = round(aic, 2), BIC = round(bic, 2)))}
    train_LNR <- calculate_performance(model_LNR, train_data)
    train_LODDS <- calculate_performance(model_LODDS, train_data)
    train_pN <- calculate_performance(model_pN, train_data)
    model_LNR_val <- cph(Surv(time, status) ~ LNR, data = validation_data, x = TRUE, y = TRUE)
    model_LODDS_val <- cph(Surv(time, status) ~ LODDS, data = validation_data, x = TRUE, y = TRUE)
    model_pN_val <- cph(Surv(time, status) ~ N, data = validation_data, x = TRUE, y = TRUE)
    val_LNR <- calculate_performance(model_LNR_val, validation_data)
    val_LODDS <- calculate_performance(model_LODDS_val, validation_data)
    val_pN <- calculate_performance(model_pN_val, validation_data)
    results <- data.frame(Variable = c("LNR", "LODDS", "pN"), Training_C_Index = c(paste(train_LNR["C_Index"], "(", train_LNR["CI_Lower"], ", ", train_LNR["CI_Upper"], ")", sep = ""), paste(train_LODDS["C_Index"], "(", train_LODDS["CI_Lower"], ", ", train_LODDS["CI_Upper"], ")", sep = ""), paste(train_pN["C_Index"], "(", train_pN["CI_Lower"], ", ", train_pN["CI_Upper"], ")", sep = "")), Training_AIC = c(train_LNR["AIC"], train_LODDS["AIC"], train_pN["AIC"]), Training_BIC = c(train_LNR["BIC"], train_LODDS["BIC"], train_pN["BIC"]), Validation_C_Index = c(paste(val_LNR["C_Index"], "(", val_LNR["CI_Lower"], ", ", val_LNR["CI_Upper"], ")", sep = ""), paste(val_LODDS["C_Index"], "(", val_LODDS["CI_Lower"], ", ", val_LODDS["CI_Upper"], ")", sep = ""), paste(val_pN["C_Index"], "(", val_pN["CI_Lower"], ", ", val_pN["CI_Upper"], ")", sep = "")), Validation_AIC = c(val_LNR["AIC"], val_LODDS["AIC"], val_pN["AIC"]), Validation_BIC = c(val_LNR["BIC"], val_LODDS["BIC"], val_pN["BIC"]))
    results_table <- gt(results) %>%
    tab_header(title = "三种淋巴结分期系统的预测性能") %>%
    cols_label(Variable = "变量",Training_C_Index = "C指数 (95% CI) (训练集)", Training_AIC = "AIC (训练集)", Training_BIC = "BIC (训练集)", Validation_C_Index = "C指数 (95% CI) (验证集)", Validation_AIC = "AIC (验证集)", Validation_BIC = "BIC (验证集)")
    write.csv(results, "prediction_performance.csv", row.names = FALSE)
  4. 使用以下代码构建 XGBoost 模型,并生成变量相对重要性的柱状图,从而比较三种淋巴结系统的重要性。同样地,生成 ROC 曲线和校准曲线。数据来源于 SEER 数据库。
    library(xgboost)
    library(caret)
    library(pROC)
    train_data <- read.csv("train_data.csv")
    test_data <- read.csv("test_data.csv")
    train_matrix <- xgb.DMatrix(data = as.matrix(train_data[, c('Age', 'T', 'N', 'M', 'LODDS', 'Chemotherapy')]), label = train_data$status)
    test_matrix <- xgb.DMatrix(data = as.matrix(test_data[, c('Age', 'T', 'N', 'M', 'LODDS', 'Chemotherapy')]), label = test_data$status)
    params <- list(booster = "gbtree", objective = "binary:logistic", eval_metric = "auc", eta = 0.1, max_depth = 6, subsample = 0.8, colsample_bytree = 0.8)
    xgb_model <- xgb.train(params = params, data = train_matrix, nrounds = 100, watchlist= list(train = train_matrix), verbose = 1)
    pred_probs <- predict(xgb_model, newdata = test_matrix)
    pred_labels <- ifelse(pred_probs > 0.5, 1, 0)
    conf_matrix <- confusionMatrix(as.factor(pred_labels), as.factor(test_data$status))
    roc_curve <- roc(test_data$status, pred_probs)
    auc_value <- auc(roc_curve)
    ci_auc <- ci.auc(roc_curve)
    sensitivity <- conf_matrix$byClass["Sensitivity"]
    specificity <- conf_matrix$byClass["Specificity"]
    accuracy <- conf_matrix$overall["Accuracy"]
    ppv <- conf_matrix$byClass["Pos Pred Value"]
    npv <- conf_matrix$byClass["Neg Pred Value"]
    result_table <- data.frame(Model = "XGBoost", AUC = sprintf("%.3f (%.3f-%.3f)", auc_value, ci_auc[1], ci_auc[3]), Sensitivity = sprintf("%.3f", sensitivity), Specificity = sprintf("%.3f", specificity), Accuracy = sprintf("%.3f", accuracy), PPV = sprintf("%.3f", ppv), NPV = sprintf("%.3f", npv))
    write.csv(result_table, "xgboost_model_performance.csv", row.names = FALSE)
    roc_df <- data.frame(FPR = 1 - roc_curve$specificities, TPR = roc_curve$sensitivities)
    roc_plot <- ggplot(roc_df, aes(x = FPR, y = TPR)) +geom_line(color = "steelblue", size = 1.2) + geom_abline(intercept = 0, slope = 1, linetype = "dashed", color = "gray") + annotate("text", x = 0.9, y = 0.2, label = paste("AUC =", round(auc_value, 3)), size = 5, color = "black") + labs(title = "XGBoost 模型的 ROC 曲线", x = "假阳性率", y = "真阳性率") + theme_minimal() + theme(panel.border = element_rect(color = "black", fill = NA, size = 1))
    calibration_data <- data.frame(Status = as.factor(test_data$status), pred_probs = pred_probs)
    calib_model <- calibration(Status ~ pred_probs, data = calibration_data, class = "1", cuts = 5)
    ggplot(calib_model$data, aes(x = midpoint, y = Percent)) + geom_line(color = "steelblue", size = 1) + geom_point(color = "red", size = 2) + geom_abline(intercept = 0, slope = 1, linetype = "dashed", color = "black") +labs(title = "XGBoost 模型的校准曲线", x = "预测概率", y = "观测比例") + theme_minimal() + theme(panel.border = element_rect(color = "black", fill = NA, size = 0.5))
  5. 使用以下代码构建随机森林(RF)模型,并生成变量相对重要性的柱状图,从而比较三种淋巴结系统的重要性。同样地,生成 ROC 曲线和校准曲线。数据来源于 SEER 数据库。library(randomForest)
    library(dplyr)
    library(ggplot2)
    library(pROC)
    library(caret)
    library(rms)
    trainset <- read.csv("train_data.csv")
    testsed <- read.csv("test_data.csv")
    trainset$status=factor(trainset$status)
    variables1 <- c("Age", "T", "N", "M", "LODDS", "Chemotherapy")
    trainset <- trainset %>%
    mutate(across(all_of(variables1), as.numeric))
    testsed$status=factor(testsed$status)
    testsed <- testsed %>%
    mutate(across(all_of(variables1), as.numeric))
    RF=randomForest(trainset$status ~ Age + T + N + M + LODDS + Chemotherapy, data=trainset,ntree=100,importance=TRUE,proximity=TRUE)
    imp=importance(RF)
    varImpPlot(RF)
    impvar=rownames(imp)[order(imp[,4],decreasing = TRUE)]
    importance_df <- as.data.frame(imp)
    importance_df$Variables <- rownames(importance_df)
    importance_plot <- ggplot(importance_df, aes(x = reorder(Variables, MeanDecreaseAccuracy), y = MeanDecreaseAccuracy)) +geom_bar(stat = "identity", fill = "steelblue") +coord_flip() + labs(title = "变量重要性", x = "变量", y = "平均准确度下降") + theme_minimal()
    pred_probs <- predict(RF, testset, type = "prob")[,2]
    roc_obj <- roc(testset$status, pred_probs)
    auc_value <- auc(roc_obj)
    roc_plot <- ggplot() +geom_line(aes(x = 1 - roc_obj$specificities, y = roc_obj$sensitivities), color = "steelblue", size = 1.2) +geom_abline(intercept = 0, slope = 1, linetype = "dashed", color = "gray") + annotate("text", x = 0.8, y = 0.2, label = paste("AUC =", round(auc_value, 3)), color = "black", size = 5, hjust = 0) + labs(title = "随机森林模型的 ROC 曲线", x = "假阳性率", y = "真阳性率") +theme_minimal() + theme(panel.border = element_rect(color = "black", fill = NA, size = 1))
    calibration_data <- data.frame(pred_probs = pred_probs, status = testsed$status)
    calib_model <- calibration(status ~ pred_probs, data = calibration_data, class = "1", cuts = 5)
    calib_df <- as.data.frame(calib_model[["data"]])
    calib_df$mid <- calib_df$midpoint
    calib_df$Percent <- calib_df$Percent
    calibration_plot <- ggplot(calib_df, aes(x = mid, y = Percent)) + geom_line(color = "steelblue", size = 1.2) + geom_point(color = "steelblue", size = 3) + geom_abline(intercept = 0, slope = 1, linetype = "dashed", color = "black", size = 0.8) + labs(title = "随机森林的校准曲线", x = "预测概率", y = "实际概率") + theme_minimal() + theme(panel.border = element_rect(color = "black", fill = NA, size = 1), plot.title = element_text(hjust = -0.05, vjust = -1.5, face = "bold", size = 12) )
    rf_probs <- predict(RF, newdata=testsed, type="prob")[, 2]
    rf_auc <- roc(testsed$status, rf_probs)
    auc_value <- auc(rf_auc)
    ci_auc <- ci.auc(rf_auc)
    rf_predictions <- predict(RF, newdata=testsed)
    conf_matrix <- confusionMatrix(rf_predictions, testsed$status)
    sensitivity <- conf_matrix$byClass["Sensitivity"]
    specificity <- conf_matrix$byClass["Specificity"]
    accuracy <- conf_matrix$overall["Accuracy"]
    ppv <- conf_matrix$byClass["Pos Pred Value"]
    npv <- conf_matrix$byClass["Neg Pred Value"]
    result_table <- data.frame(Model = "RF", AUC = sprintf("%.3f (%.3f-%.3f)", auc_value, ci_auc[1], ci_auc[3]), Sensitivity = sprintf("%.3f", sensitivity), Specificity = sprintf("%.3f", specificity), Accuracy = sprintf("%.3f", accuracy), PPV = sprintf("%.3f", ppv), NPV = sprintf("%.3f", npv))
    write.csv(result_table, "RF_model_performance.csv", row.names = FALSE)
  6. 使用以下代码构建神经网络(NN)模型,并生成变量相对重要性的柱状图,从而比较三种淋巴结系统的重要性。同样地,生成 ROC 曲线和校准曲线。数据来源于 SEER 数据库。library(nnet)
    library(caret)
    library(pROC)
    library(ggplot2)
    train_data <- read.csv("train_data.csv")
    test_data <- read.csv("test_data.csv")
    train_data$status <- as.factor(train_data$status)
    test_data$status <- as.factor(test_data$status)
    features <- c("Age", "T", "N", "M", "LODDS", "Chemotherapy")
    x_train <- train_data[, features]
    y_train <- train_data$status
    x_test <- test_data[, features]
    y_test <- test_data$status
    nn_model <- nnet(status ~ Age + T + N + M + LODDS + Chemotherapy, data = train_data, size = 5, decay = 0.01, maxit = 200)
    pred_probs <- predict(nn_model, newdata = x_test, type = "raw")
    pred_labels <- ifelse(pred_probs > 0.5, 1, 0)
    roc_curve <- roc(as.numeric(y_test), pred_probs)
    auc_value <- auc(roc_curve)
    auc_ci <- ci.auc(roc_curve)
    auc_text <- paste0(round(auc_value, 3), " (", round(auc_ci[1], 3), "-", round(auc_ci[3], 3), ")")
    conf_matrix <- confusionMatrix(as.factor(pred_labels), y_test)
    accuracy <- conf_matrix$overall["Accuracy"]
    sensitivity <- conf_matrix$byClass["Sensitivity"]
    specificity <- conf_matrix$byClass["Specificity"]
    ppv <- conf_matrix$byClass["Pos Pred Value"]
    npv <- conf_matrix$byClass["Neg Pred Value"]
    performance_table <- data.frame(Metric = c("AUC (95% CI)", "Accuracy", "Sensitivity", "Specificity", "PPV", "NPV"),Value = c(auc_text, round(accuracy, 3), round(sensitivity, 3), round(specificity, 3), round(ppv, 3), round(npv, 3)))
    write.csv(performance_table, "NN_performance_table.csv", row.names = FALSE)
    roc_curve <- roc(y_test, pred_probs)
    auc_value <- auc(roc_curve)
    roc_plot <- ggplot() + geom_line(aes(x = 1 - roc_curve$specificities, y = roc_curve$sensitivities), color = "steelblue", size = 1.2) +geom_abline(intercept = 0, slope = 1, linetype = "dashed", color = "gray") + annotate("text", x = 0.8, y = 0.2, label = paste("AUC =", round(auc_value, 3)), color = "black", size = 5, hjust = 0) + labs(title = "神经网络模型的 ROC 曲线", x = "假阳性率", y = "真阳性率") + theme_minimal() + theme(panel.border = element_rect(color = "black", fill = NA, size = 1))
    calibration_data <- data.frame(pred_probs = pred_probs, status = as.numeric(y_test) - 1)
    calibration_data$pred_probs <- as.numeric(calibration_data$pred_probs)
    calibration_data$calibration_bin <- cut(calibration_data$pred_probs, breaks = seq(0, 1, by = 0.2), include.lowest = TRUE)
    calibration_summary <- aggregate(status ~ calibration_bin, data = calibration_data, FUN = mean)
    calibration_summary$pred_mean <- aggregate(pred_probs ~ calibration_bin, data = calibration_data, FUN = mean)$pred_probs
    calibration_plot <- ggplot(calibration_summary, aes(x = pred_mean, y = status)) + geom_line(color = "steelblue", size = 1.2) + geom_point(color = "red", size = 3) + geom_abline(intercept = 0, slope = 1, linetype = "dashed", color = "black", size = 0.8) + labs(title = "神经网络的校准曲线", x = "预测概率", y = "实际概率") + theme_minimal() + theme(panel.border = element_rect(color = "black", fill = NA, size = 1))
    nn_var_importance <- varImp(nn_model)
    importance_df <- data.frame(Feature = rownames(nn_var_importance), Importance = nn_var_importance$Overall )
    importance_plot <- ggplot(importance_df, aes(x = reorder(Feature, Importance), y = Importance)) + geom_bar(stat = "identity", fill = "steelblue") + coord_flip() + labs(title = "神经网络的变量重要性", x = "特征", y = "重要性") + theme_minimal()

3. 竞争风险模型的构建与验证

  1. 使用以下代码进行单因素分析并绘制累积发生率函数(CIF)曲线。data.csv 是来自 SEER 数据库的数据。后续图像的保存方法与此步骤相同。将代码中的 Site 逐一替换为其他因素,以对所有因素进行单因素分析。
    library(tidycmprsk)
    library(gtsummary)
    library(ggplot2)
    library(ggsurvfit)
    library(ggprism)
    aa <- read.csv("data.csv")
    cif2 <- tidycmprsk::cuminc(Surv(time, Status1) ~Site, data = aa)
    tidy(cif2,times = c(12,24,36,48,60))
    tbl_cuminc(cif2, times =c(12,24,36,48,60), outcomes = c("CSS", "OSS"),estimate_fun = NULL, label_header = "**{time/12}-year cuminc**") %>%
    add_p() %>%
    add_n(location = "level")
    cuminc_plot <- ggcuminc(cif2, outcome = c("CSS", "OSS"), size = 1.5) + labs(x = "time") +add_quantile(y_value = 0.20, size = 1) + scale_x_continuous(breaks = seq(0, 84, by = 12), limits = c(0, 84)) +scale_y_continuous(label = scales::percent, breaks = seq(0, 1, by = 0.2), limits = c(0, 1)) + theme_prism() + theme(legend.position = c(0.2, 0.8), panel.grid = element_blank(),panel.grid.major.y = element_line(colour = "grey80")) + theme(legend.spacing.x = unit(0.1, "cm"), legend.spacing.y = unit(0.01, "cm")) + theme(axis.ticks.length.x = unit(-0.2, "cm"), axis.ticks.x = element_line(color = "black", size = 1, lineend = 1)) + theme(axis.ticks.length.y = unit(-0.2, "cm"), axis.ticks.y = element_line(color = "black", size = 1, lineend = 1))
  2. 使用以下代码进行多变量分析与可视化。data1.csv 来源于先前代码的运行结果。运行代码后,点击 Export然后点击 Save as PDF, and finally click Save 保存图像。
    library(tidycmprsk)
    library(gtsummary)
    aa <-read.csv('data1.csv')
    for (i in names(aa)[c(1:16, 19)]){aa[,i] <- as.factor(aa[,i])}
    mul1<tidycmprsk::crr(Surv(time,Status1)~Sex+Age+Race+Marital+T+N+M+Tumor_size+LNR+LODDS+Radiation+Chemotherapy+Site,data=aa, failcode=1,cencode=0)
    table2 <- mul1 %>%
    gtsummary::tbl_regression(exponentiate = TRUE) %>%
    add_n(location = "level");table2
    table_df <- as_tibble(table2)
    tab <- table2$table_body
    tab1 <- tab[,c(12,19,20,22:29)]
  3. Use the following code to plot the nomogram, ROC curve, and calibration curve. After training the model using data from the training cohort, use the validation and external validation cohorts data to validate the model.library(QHScrnomo). The external cohort data consists of samples of colorectal cancer other than ring cell carcinoma, which were selected in step 1.4.
    library(rms)
    library(timeROC)
    library(survival)
    aa <-read.csv('data3.csv')
    for (i in names(aa)[c(1:16, 19)]){aa[,i] <- as.factor(aa[,i])}
    dd <- datadist(aa)
    options(datadist = "dd")
    mul <- cph(Surv(time, Status1 == 1) ~ T + N + M + LODDS + Site, data = aa, x = TRUE, y = TRUE, surv = TRUE)
    m3 <- crr.fit(mul, failcode = 1, cencode = 0)
    nomo <-Newlabels(fit = m3, labels =c(T="T", N= "N", M = "M", LODDS = "LODDS", Site = "Site"))
    nomo<Newlevels(fit=nomo,list(T=c("T1","T2","T3","T4"),N=
    c("N0","N1","N2"),M=c("M0","M1"),LODDS=c
    ("LODDS1","LODDS2","LODDS3"),Site=
    c("RSC","LSC","Rectum")))
    nomogram.crr(fit =nomo , lp = F, xfrac = 0.3, fun.at =seq(from=0, to=1, by= 0.1) , failtime =c(12,36,60), funlabel = c("1-year CSS Cumulative Incidence","3-year CSS Cumulative Incidence","5-year CSS Cumulative Incidence"))
    time_points <- c(12, 36, 60)
    pred_risks_list <- lapply(time_points, function(time_point) {predict(m3, newdata = aa, type = "risk", time = time_point)})
    pred_risks_df <- data.frame(do.call(cbind, pred_risks_list))
    colnames(pred_risks_df) <- paste("risk_at", time_points, "months", sep = "_")
    roc_1year <- timeROC(T = aa$time, delta = ifelse(aa$Status1 == "CSS", 1, 0), marker = pred_risks_df$risk_at_12_months, cause = 1, times = 12, iid = TRUE)
    roc_3year <- timeROC(T = aa$time, delta = ifelse(aa$Status1 == "CSS", 1, 0), marker = pred_risks_df$risk_at_36_months, cause = 1, times = 36, iid = TRUE)
    roc_5year <- timeROC(T = aa$time, delta = ifelse(aa$Status1 == "CSS", 1, 0), marker = pred_risks_df$risk_at_60_months, cause = 1,times = 60, iid = TRUE)
    legend("bottomright",legend = c("1 year CSS", "3 year CSS", "5 year CSS"), col = c("#BF1D2D", "#262626", "#397FC7"), lwd = 2)
    sas.cmprsk(m3,time = 36)
    set.seed(123)
    aa$pro <- tenf.crr(m3,time = 36)
    cindex(prob = aa$pro, fstatus = aa$Status1, ftime = aa$time, type = "crr", failcode = 1, cencode = 0, tol = 1e-20)
    groupci(x=aa$pro, ftime = aa$time, fstatus = aa$Status1, failcode = 1, cencode = 0, ci = TRUE, g = 5, m = 1000, u = 36, xlab = "Predicted Probability ", ylab = "Actual Probability", lty=1, lwd=2, col="#262626",xlim=c(0,1.0), ylim=c(0,1.0), add =TRUE)

访问受限。请登录或开始试用以查看此内容。

结果

患者特征
本研究聚焦于结直肠印戒细胞癌(SRCC)患者,数据来源于2004年至2015年期间的SEER数据库。排除标准包括生存时间少于一个月的患者、临床病理信息不完整的患者,以及死亡原因不明确或未指明的病例。共有2409例符合纳入标准的结直肠SRCC患者,被随机分为训练队列(N = 1686)和验证队列(N = 723)。使用R软件对训练队列和验证队列的人口统计学及临床参数进行分析,结果如表1所示。在所有纳入的患者中,大多数年龄超过60岁,男女患者数量相近。大多数患者为白人。超过一半的患者(56%)已婚。大多数肿瘤分级为III-IV级(76%)。大多数患者(82%)的肿瘤大小超过3.5厘米,且多数患者属于LODDS1组(42%)。在整个队列中,接受化疗的患者比例较高(53%)。原发肿瘤主要位于右半结肠(67%)。经随机分组后,两组之间的基线特征在统计学上无显著差异。

机器学习模型中包含的预后临床因素的确定
我们首先通过Cox回归分析筛选出对机器学习模型具有显著性的变...

访问受限。请登录或开始试用以查看此内容。

讨论

结直肠印戒细胞癌(SRCC)是一种罕见且特殊的结直肠癌亚型,预后较差。因此,需要更加关注SRCC患者的预后情况。准确预测SRCC患者的生存率对于判断其预后并制定个体化治疗决策至关重要。在本研究中,我们探讨了SRCC患者临床特征与预后之间的关系,并从SEER数据库中确定了适用于SRCC患者的最优淋巴结分期系统。据我们所知,这是首次通过综合运用机器学习与竞争风险分析方法确定结直肠SRCC患者适宜的淋巴结分期系统,并构建用于预后预测的列线图的研究。

结直肠癌(CRC)患者转移性淋巴结(LN)的数量是评估预后和复发的重要指标。准确的淋巴结分期在确定印戒细胞癌(SRCC)患者的治疗策略和预后中起关键作用。淋巴结比率(LNR)和阳性淋巴结数的对数比值(LODDS)是用于评估胃癌(GC)淋巴结受累情况的替代方法,可改进分期系统并提供更精确的预后信息10,13。本研究利用SEER数据库,揭示了LODDS、LNR和pN分期与SRCC患者癌症特异性生存(CSS)之间的相关性。通过AUC...

访问受限。请登录或开始试用以查看此内容。

披露

作者无须披露任何财务利益冲突。

材料

本文使用的材料清单
姓名公司目录编号评论
SEER数据库美国国立卫生研究院国家癌症研究所
X-tile 软件耶鲁医学院
R-studio阳性淋巴结数,淋巴结转移率,log odds of positive lymph nodes,机器学习,竞争风险模型,结直肠印戒细胞癌,癌症特异性生存率,SEER数据库,预后预测,pN分期

参考文献

  1. Siegel, R. L., Giaquinto, A. N., Jemal, A. Cancer statistics, 2024. CA Cancer J Clin. 74 (1), 12-49 (2024).
  2. Korphaisarn, K., et al. Signet ring cell colorectal cancer: Genomic insights into a rare subpopulation of colorectal adenocarcinoma. Br J Cancer. 121 (6), 505-510 (2019).
  3. Willauer, A. N., et al. Clinical and molecular characterization of early-onset colorectal cancer. Cancer. 125 (12), 2002-2010 (2019).
  4. Watanabe, A., et al. A case of primary colonic signet ring cell carcinoma in a young man which preoperatively mimicked Phlebosclerotic colitis. Acta Med Okayama. 73 (4), 361-365 (2019).
  5. Kim, H., Kim, B. H., Lee, D., Shin, E. Genomic alterations in signet ring and mucinous patterned colorectal carcinoma. Pathol Res Pract. 215 (10), 152566(2019).
  6. Deng, X., et al. Neoadjuvant radiotherapy versus surgery alone for stage II/III mid-low rectal cancer with or without high-risk factors: A prospective multicenter stratified randomized trial. Ann Surg. 272 (6), 1060-1069 (2020).
  7. Buk Cardoso, L., et al. Machine learning for predicting survival of colorectal cancer patients. Sci Rep. 13 (1), 8874(2023).
  8. Monterrubio-Gómez, K., Constantine-Cooke, N., Vallejos, C. A. A review on statistical and machine learning competing risks methods. Biom J. 66 (2), e2300060(2024).
  9. Kim, H. J., Choi, G. S. Clinical implications of lymph node metastasis in colorectal cancer: Current status and future perspectives. Ann Coloproctol. 35 (3), 109-117 (2019).
  10. Xu, T., et al. Log odds of positive lymph nodes is an excellent prognostic factor for patients with rectal cancer after neoadjuvant chemoradiotherapy. Ann Transl Med. 9 (8), 637(2021).
  11. Chen, Y. R., et al. Prognostic performance of different lymph node classification systems in young gastric cancer. J Gastrointest Oncol. 12 (4), 285-1300 (2021).
  12. Bouvier, A. M., et al. How many nodes must be examined to accurately stage gastric carcinomas? Results from a population based study. Cancer. 94 (11), 2862-2866 (2002).
  13. Coburn, N. G., Swallow, C. J., Kiss, A., Law, C. Significant regional variation in adequacy of lymph node assessment and survival in gastric cancer. Cancer. 107 (9), 2143-2151 (2006).
  14. Li Destri, G., Di Carlo, I., Scilletta, R., Scilletta, B., Puleo, S. Colorectal cancer and lymph nodes: the obsession with the number 12. World J Gastroenterol. 20 (8), 1951-1960 (2014).
  15. Dinaux, A. M., et al. Outcomes of persistent lymph node involvement after neoadjuvant therapy for stage III rectal cancer. Surgery. 163 (4), 784-788 (2018).
  16. Sun, Y., Zhang, Y., Huang, Z., Chi, P. Prognostic implication of negative lymph node count in ypN+ rectal cancer after neoadjuvant chemoradiotherapy and construction of a prediction nomogram. J Gastrointest Surg. 23 (5), 1006-1014 (2019).
  17. Xu, Z., Jing, J., Ma, G. Development and validation of prognostic nomogram based on log odds of positive lymph nodes for patients with gastric signet ring cell carcinoma. Chin J Cancer Res. 32 (6), 778-793 (2020).
  18. Scarinci, A., et al. The impact of log odds of positive lymph nodes (LODDS) in colon and rectal cancer patient stratification: a single-center analysis of 323 patients. Updates Surg. 70 (1), 23-31 (2018).
  19. Nitsche, U., et al. Prognosis of mucinous and signet-ring cell colorectal cancer in a population-based cohort. J Cancer Res Clin Oncol. 142 (11), 2357-2366 (2016).
  20. Kang, H., O'Connell, J. B., Maggard, M. A., Sack, J., Ko, C. Y. A 10-year outcomes evaluation of mucinous and signet-ring cell carcinoma of the colon and rectum. Dis Colon Rectum. 48 (6), 1161-1168 (2005).
  21. Sung, C. O., et al. Clinical significance of signet ring cells in colorectal mucinous adenocarcinoma. Mod Pathol. 21 (12), 1533-1541 (2008).
  22. Alvi, M. A., et al. Molecular profiling of signet ring cell colorectal cancer provides a strong rationale for genomic targeted and immune checkpoint inhibitor therapies. Br J Cancer. 117 (2), 203-209 (2017).
  23. Brownlee, S., et al. Evidence for overuse of medical services around the world. Lancet. 390 (10090), 156-168 (2017).

访问受限。请登录或开始试用以查看此内容。

重印与许可

标签

机器学习模型癌症生存预测LODDS分类淋巴结比率随机森林XGBoost模型神经网络模型竞争风险模型