Identification and functional annotation of cuproptosis–ferroptosis co-expressed genes
Four training datasets comprising 421 UC samples and 97 healthy controls were integrated. A total of 551 differentially expressed genes, including 362 upregulated and 189 downregulated genes, were identified. Correlation analysis yielded 444 ferroptosis–cuproptosis-correlated genes, and intersection with the differentially expressed genes produced 32 cuproptosis–ferroptosis co-expressed differentially expressed genes (CF-DEGs).
The 32 CF-DEGs were enriched in responses to wounding, copper ions, fatty-acid transport, lipopolysaccharide, inflammatory bowel disease, NF-κB signaling, TNF signaling, and ferroptosis (Supplementary Table 4).
Prioritization of core biomarkers through integrated machine learning and WGCNA
Weighted gene co-expression network analysis (WGCNA) identified MEpurple, MEbrown, and MEblack as the modules most strongly associated with UC. Together, the selected modules contained 1,426 genes (Supplementary Table 5).
Least absolute shrinkage and selection operator (LASSO), support vector machine–recursive feature elimination (SVM-RFE), and random forest selected 21, 32, and 19 features, respectively. Thirteen genes were shared by all three models. Integration of the machine-learning consensus, the CytoHubba-ranked protein–protein interaction core, and the selected WGCNA modules yielded five UC-associated candidate biomarkers: LCN2, IDO1, CXCL2, NOS2, and CD274.
Diagnostic evaluation and pathway enrichment of the biomarker signature
All five candidate biomarkers were positively correlated and upregulated in UC in the training cohort (Figure 2A, B). Each marker achieved an area under the receiver operating characteristic curve > 0.80 in the training cohort and > 0.75 in the independent GSE47908 validation cohort (Figure 2C–E).
The five-gene nomogram showed favorable calibration, with a Hosmer–Lemeshow P > 0.05 and mean absolute error < 0.1, and discrimination in the analyzed cohort, with an area under the curve and concordance index of 0.945 and a 95% confidence interval for the concordance index of 0.922–0.968 (Figure 3).
Gene set enrichment analysis linked the five candidates to JAK–STAT, interferon–RIPK1/3, and Toll-like receptor–NF-κB signaling (Figure 4).

Figure 2: Candidate-biomarker expression and diagnostic performance. (A) Correlation heatmap. (B, C) Biomarker expression and receiver operating characteristic curves in the training cohort. (D, E) Biomarker expression and receiver operating characteristic curves in the GSE47908 validation cohort. ROC: receiver operating characteristic; AUC: area under the curve; UC: ulcerative colitis. Please click here to view a larger version of this figure.

Figure 3: Development and assessment of the UC nomogram. (A) Five-gene nomogram. (B) Calibration plot. (C) Receiver operating characteristic curve. ROC: receiver operating characteristic; UC: ulcerative colitis; AUC: area under the curve; C-index: concordance index. Please click here to view a larger version of this figure.

Figure 4: Gene set enrichment analysis of the five candidate biomarkers. (A–E) Gene set enrichment analysis results for LCN2, IDO1, CXCL2, NOS2, and CD274, respectively. GSEA: gene set enrichment analysis. Please click here to view a larger version of this figure.
Immune microenvironment landscape and regulatory network analysis
UC samples showed increased proportions of neutrophils, activated memory CD4⁺ T cells, M1 macrophages, and activated mast cells, with reciprocal decreases in M2 macrophages, resting mast cells, and resting dendritic cells. These patterns were reproduced in the validation cohort (Figure 5A–E).
The five candidates correlated positively with neutrophils, activated memory CD4⁺ T cells, and M1 macrophages and inversely with resting mast cells and M2 macrophages (Figure 5F–J). The gene miRNA network contained 98 nodes and 146 edges. hsa-miR-34a-5p and hsa-miR-16-5p showed the highest biomarker connectivity, whereas AR and RELA were the most connected transcription factors (Figure 6).

Figure 5: Immune-infiltration analysis. (A) Immune-cell composition. (B) Between-group differences in immune-cell proportions. (C) Immune-cell correlation heatmap. (D, E) Immune-cell infiltration in the training and validation cohorts. (F–J) Correlations between candidate biomarkers and immune-cell populations. UC: ulcerative colitis. Please click here to view a larger version of this figure.

Figure 6: Predicted miRNA and transcription-factor networks. (A) Gene–miRNA network. Circles indicate candidate biomarkers, and squares indicate miRNAs. (B) Transcription factor–gene network. Diamonds indicate candidate biomarkers, and inverted triangles indicate transcription factors. miRNA: microRNA; TF: transcription factor. Please click here to view a larger version of this figure.
Spatiotemporal expression dynamics at single-cell resolution
After quality control, 31,712 cells formed 22 clusters that were annotated as nine major cell populations (Figure 7A–C). Epithelial and plasma-cell proportions were higher in UC. LCN2 and NOS2 were enriched in epithelial cells, whereas IDO1, CXCL2, and CD274 were enriched in myeloid cells (Figure 7D–G).
Epithelial-cell subclustering identified 11 subsets, with expansion of inflammatory colonocytes and enrichment of LCN2 and NOS2 in this subset (Figure 8). Myeloid-cell subclustering identified seven subsets, with increased monocytes, reduced macrophages, and enrichment of IDO1, CXCL2, and CD274 in monocytes (Figure 9).
Inflammatory colonocytes accumulated late in the epithelial trajectory, with increasing LCN2 and NOS2 expression (Figure 10). Monocytes exhibited a distinct UC-associated trajectory, with dynamic expression of IDO1, CXCL2, and CD274 (Figure 11).
Inflammatory colonocytes showed the strongest outgoing signaling and prominent communication with monocytes. APP–CD74 was a leading ligand–receptor pair between these subsets (Figure 12).

Figure 7: Candidate-biomarker expression across cell populations. (A) Cell clusters. (B) Annotation markers. (C) Nine annotated cell populations. (D, E) Cell distributions and proportions in healthy and UC samples. (F, G) Candidate-biomarker expression visualized by UMAP and bubble plot. UC: ulcerative colitis; UMAP: uniform manifold approximation and projection. Please click here to view a larger version of this figure.

Figure 8: Epithelial-cell subclustering and candidate expression. (A) Initial epithelial-cell clusters. (B) Annotation markers. (C) Annotated epithelial-cell subsets. (D) Subset proportions in healthy and UC samples. (E) LCN2 and NOS2 expression across epithelial-cell subsets. t-SNE: t-distributed stochastic neighbor embedding; UC: ulcerative colitis. Please click here to view a larger version of this figure.

Figure 9: Myeloid-cell subclustering and candidate expression. (A) Initial myeloid-cell clusters. (B) Annotation markers. (C) Annotated myeloid-cell subsets. (D) Subset proportions in healthy and UC samples. (E) IDO1, CXCL2, and CD274 expression across myeloid-cell subsets. t-SNE: t-distributed stochastic neighbor embedding; UC: ulcerative colitis. Please click here to view a larger version of this figure.

Figure 10: Epithelial-cell pseudotime analysis. (A) Pseudotime trajectory and state assignments. (B) Distribution of healthy and UC epithelial cells along the trajectory. (C) LCN2 and NOS2 expression dynamics along pseudotime. UC: ulcerative colitis. Please click here to view a larger version of this figure.

Figure 11: Myeloid-cell pseudotime analysis. (A) Pseudotime trajectory and state assignments. (B) Distribution of healthy and UC myeloid cells along the trajectory. (C) IDO1, CXCL2, and CD274 expression dynamics along pseudotime. UC: ulcerative colitis. Please click here to view a larger version of this figure.

Figure 12: Epithelial–myeloid communication. (A) Interaction number and strength. (B) Outgoing and incoming signaling strengths. (C) Interactions involving inflammatory colonocytes. (D) Communication-strength heatmap. (E) Ligand–receptor pairs involving inflammatory colonocytes. Please click here to view a larger version of this figure.
In vitro experimental validation of ferroptosis and cuproptosis interventions
Lipopolysaccharide (LPS) exposure reduced Caco-2 cell viability and increased IL-6 and IL-1β expression relative to the vehicle control (Figure 13A–C and Supplementary Figure 5).
LPS increased Fe2⁺, malondialdehyde, and the messenger RNA abundance of LCN2, IDO1, CXCL2, and NOS2, whereas ferrostatin-1 reversed each LPS-associated change (Figure 13D–I). For the LPS versus LPS + ferrostatin-1 comparisons, exact two-sided P values ranged from 9.45 × 10⁻5 to 0.0027 after averaging technical replicates within each biological replicate.
LPS reduced viability and increased IL-6, IL-1β, FDX1/DLAT, CD274, and copper-sensitive fluorescence, whereas tetrathiomolybdate reversed these changes (Figure 14 A–E). For the LPS versus LPS + tetrathiomolybdate comparisons, exact two-sided P values ranged from < 1 × 10⁻15 to 0.0008.
Together, the analyses identified five UC-associated diagnostic candidates, localized their expression to epithelial and myeloid populations, and showed that LCN2, IDO1, CXCL2, and NOS2 responded to ferroptosis inhibition, whereas CD274 responded to copper chelation in Caco-2 cells.

Figure 13: Ferrostatin-1 attenuates LPS-associated inflammatory and ferroptosis-related changes in Caco-2 cells. (A) Cell viability. (B, C) IL-6 and IL-1β messenger RNA expression. (D, E) Intracellular malondialdehyde and Fe2⁺ levels. (F–I) LCN2, IDO1, CXCL2, and NOS2 messenger RNA expression. Technical triplicates were averaged within each independent biological replicate; error bars indicate standard deviation. Available biological replicate counts were Control n = 4 in A and C, Control n = 2 in F, and n = 6 for all other panel/group combinations. Exact two-sided P values and assumption checks are reported in Supplementary Table 5. Fer-1: ferrostatin-1; LPS: lipopolysaccharide; MDA: malondialdehyde. Please click here to view a larger version of this figure.

Figure 14: Tetrathiomolybdate attenuates LPS-associated inflammatory and cuproptosis-related changes in Caco-2 cells. (A) Cell viability. (B, C) IL-6 and IL-1β messenger RNA expression. (D) FDX1 and DLAT messenger RNA expression. (E) CD274 messenger RNA expression. (F) Intracellular Cu2⁺ fluorescence; scale bar = 50 µm. Technical triplicates were averaged within each of six independent biological replicates per group; error bars indicate standard deviation. Exact two-sided P values and assumption checks are reported in Supplementary Table 5. TTM: tetrathiomolybdate; LPS: lipopolysaccharide. Please click here to view a larger version of this figure.
Raw and processed data, analysis scripts, and cell-experiment source spreadsheets are publicly available at https://doi.org/10.5281/zenodo.21202720.
Supplementary Figure 1: Data preprocessing and differential-expression diagnostics. (A–F) Sample distributions before (A–C) and after (D–F) batch correction. (G) Volcano plot of differentially expressed genes. (H) Heatmap of normalized expression patterns. DEGs: differentially expressed genes; GEO: Gene Expression Omnibus.Please click here to download this file.
Supplementary Figure 2: Enrichment analysis of CF-DEGs. (A) Gene Ontology enrichment. (B) Kyoto Encyclopedia of Genes and Genomes enrichment. CF-DEGs: cuproptosis–ferroptosis co-expressed differentially expressed genes; GO: Gene Ontology; KEGG: Kyoto Encyclopedia of Genes and Genomes.Please click here to download this file.
Supplementary Figure 3: WGCNA screening diagnostics. (A) Soft-threshold selection. (B) Module-eigengene clustering. (C) Gene-module dendrogram. (D) Module–trait associations. (E–G) Gene-significance/module-membership relationships for the purple, brown, and black modules. WGCNA: weighted gene co-expression network analysis; UC: ulcerative colitis.Please click here to download this file.
Supplementary Figure 4: Multistage biomarker screening. (A, B) LASSO feature selection. (C, D) SVM-RFE feature selection. (E, F) Random-forest feature ranking. (G) Consensus among the three models. (H, I) PPI network and CytoHubba core nodes. (J) Integration of machine learning, WGCNA, and PPI results. (K) Chromosomal locations of the five candidates. LASSO: least absolute shrinkage and selection operator; SVM-RFE: support vector machine–recursive feature elimination; RF: random forest; PPI: protein–protein interaction; WGCNA: weighted gene co-expression network analysis.Please click here to download this file.
Supplementary Figure 5: LPS model validation in Caco-2 cells. (A) Cell viability. (B, C) IL-6 and IL-1β messenger RNA expression. Technical triplicates were averaged within six independent biological replicates per group; error bars indicate standard deviation. LPS: lipopolysaccharide.Please click here to download this file.
Supplementary Table 1: GEO datasets used for discovery and validation. Dataset accession numbers, platforms, sample composition, and cohort assignments for the GEO datasets included in the discovery and validation analyses. GEO: Gene Expression Omnibus.Please click here to download this file.
Supplementary Table 2: Reverse transcription quantitative PCR primer sequences. Primer sequences used for reverse transcription quantitative PCR analysis of the genes evaluated in this study.Please click here to download this file.
Supplementary Table 3: Statistical details for Figures 13, 14, and Supplementary Figure 1. Biological-replicate counts, assumption checks, statistical analysis methods, and exact P values for the indicated experimental comparisons.Please click here to download this file.
Supplementary Table 4: Cuproptosis–ferroptosis co-expressed differentially expressed genes. List of CF-DEGs identified by intersecting ferroptosis–cuproptosis-correlated genes with differentially expressed genes.Please click here to download this file.
Supplementary Table 5: Genes in the WGCNA modules selected for biomarker screening. List of genes contained in the WGCNA modules selected for subsequent biomarker screening. WGCNA: weighted gene co-expression network analysis.Please click here to download this file.