This protocol integrates public transcriptomic datasets and colonic epithelial cell validation to identify S100P as a shared biomarker associated with inflammatory bowel disease, colorectal cancer, and pancreatic adenocarcinoma.
Subskrypcja JoVE jest wymagana do oglądania tego materiału. Zaloguj się lub rozpocznij bezpłatny okres próbny.
Artykuł badawczy
* These authors contributed equally
This protocol integrates public transcriptomic datasets and colonic epithelial cell validation to identify S100P as a shared biomarker associated with inflammatory bowel disease, colorectal cancer, and pancreatic adenocarcinoma.
Inflammatory bowel disease (IBD) is associated with an increased risk of colorectal cancer (CRC) and pancreatic adenocarcinoma (PAAD), yet the molecular features shared among these diseases remain incompletely understood. This study aimed to identify common genes and biological pathways associated with IBD, CRC, and PAAD through integrated transcriptomic analysis and experimental validation. Gene expression datasets for IBD, CRC, and PAAD were obtained from The Cancer Genome Atlas and Gene Expression Omnibus databases. Weighted gene co-expression network analysis and differential expression analysis were performed to identify disease-associated and shared genes. Gene Ontology and Kyoto Encyclopedia of Genes and Genomes (analyses were used to explore enriched biological functions and pathways. Immune cell infiltration was evaluated using Cell-type Identification by Estimating Relative Subsets of RNA Transcripts. Receiver operating characteristic analysis was performed to assess the diagnostic performance of common genes. Single-cell RNA sequencing analysis was conducted to examine the cellular distribution of S100P. In addition, the effects of S100P downregulation were evaluated in lipopolysaccharide (LPS)-stimulated colonic epithelial cells. A total of 162 disease-associated genes and four common genes were identified. Functional enrichment analyses indicated significant enrichment of immune- and inflammation-related pathways, including the interleukin-17 signaling pathway. Immune infiltration analysis revealed similar trends in several immune cell populations across IBD, CRC, and PAAD. Single-cell analysis showed elevated S100P expression in epithelial cells from all three diseases. Downregulation of S100P restored the proliferative capacity of LPS-stimulated colonic epithelial cells and reduced inflammatory cytokine expression. Integrated transcriptomic analysis identified S100P as a biomarker associated with IBD, CRC, and PAAD and highlighted shared immune-related features across these diseases.
Inflammatory bowel disease (IBD) represents a spectrum of immune-related disorders that affect the gastrointestinal tract, including ulcerative colitis and Crohn’s disease1. The etiology of IBD is highly complex, involving mucosal immune abnormalities, dysbiosis, and genetic susceptibility2. IBD is a global health concern with increasing incidence rates and substantial economic burdens, and its growing prevalence has attracted considerable attention1. Importantly, patients with IBD have a significantly increased risk of developing colorectal cancer (CRC)3 and pancreatic adenocarcinoma (PAAD)4. Although this association has been well documented, the genetic crosstalk among IBD, CRC, and PAAD remains incompletely understood.
Previous studies strongly suggest that IBD, CRC, and PAAD share common pathogenic processes; however, specific and sensitive diagnostic biomarkers remain lacking, and the shared pathogenic mechanisms have not been fully elucidated. Fortunately, with the rapid advancement and widespread availability of high-throughput sequencing technologies, numerous transcriptomic datasets from patients with IBD, CRC, and PAAD have become publicly available, enabling systematic investigation of the molecular interconnections among these diseases2,5,6.
This study employed high-throughput sequencing data from patients with IBD, CRC, and PAAD to identify disease-associated and shared genes using weighted gene co-expression network analysis (WGCNA) and differential expression analysis. These genes were further investigated to identify potential shared signaling pathways among IBD, CRC, and PAAD. Additionally, we evaluated the diagnostic value of these common genes. Single-cell RNA sequencing (scRNA-seq) analysis demonstrated that the key gene S100P was predominantly expressed in epithelial cells. Finally, we investigated the biological role of S100P in IBD.
In conclusion, this study aimed to identify common diagnostic biomarkers and biological pathways associated with IBD, CRC, and PAAD, thereby providing valuable clinical insights into the shared prevention and treatment of these diseases.
Dostęp ograniczony. Zaloguj się lub rozpocznij wersję próbną, aby wyświetlić tę treść.
This study used publicly available, de-identified datasets from The Cancer Genome Atlas (TCGA) and the Gene Expression Omnibus (GEO), as well as established commercial cell lines. No newly recruited human participants, identifiable patient information, or patient-derived samples were involved. All analyses were performed in accordance with relevant institutional guidelines and the terms of use of the public databases. Therefore, additional institutional ethical approval and informed consent were not required for this study.
Data Source
RNA-seq data for the IBD cohorts (GSE179285; platform: GPL6480 and GSE24287; platform: GPL6480), CRC cohorts (TCGA-CRC; platform: Illumina HiSeq 2000 and GSE87211; platform: GPL13497), and PAAD cohorts (GSE128735; platform: GPL20301 and GSE62452; platform: GPL6244) were downloaded from TCGA and the GEO. All datasets were accessed on December 5, 2025.
For each dataset, samples were strictly divided into two subgroups, with disease lesion/tumor tissues serving as the case group and corresponding non-lesion normal tissues serving as the control group. Specifically, the IBD cohort contained 297 intestinal mucosal samples from patients with IBD and 56 normal intestinal mucosal samples from healthy individuals; the CRC cohort included 841 primary colorectal tumor tissues and 211 matched adjacent normal colorectal epithelial tissues; and the PAAD cohort consisted of 114 PAAD tumor tissues and 106 normal pancreatic parenchymal tissues.
All datasets within the same disease category were integrated uniformly. The limma package’s normalizeBetweenArrays function was applied to perform cross-sample quantile normalization, effectively eliminating inter-platform batch effects and standardizing gene expression values across different datasets for subsequent differential expression analysis.
Screening of IBD-, CRC-, and PAAD-Related Genes and Common Genes
First, differentially expressed genes (DEGs) were screened from the IBD, CRC, and PAAD cohorts using the limma package, and the original P values were corrected using the Benjamini-Hochberg false discovery rate (FDR) method. In the IBD, CRC, and PAAD cohorts, the screening criteria were set at |logFC| > 0.4 and P < 0.05. Additionally, WGCNA was performed on all genes, with a minimum module gene threshold of 100 (soft-threshold power = 0.90; network type = signed). Consequently, common DEGs and module genes were identified across the three cohorts. Genes consistently identified by both methods were defined as common genes, whereas the remaining genes were categorized as related genes.
PPI and Functional Enrichment Analysis
These analyses were performed on the disease-associated genes. Protein–protein interaction (PPI) analysis was conducted using the STRING database (interaction score > 0.40). Functional enrichment analysis included Gene Ontology (GO) and Kyoto Encyclopedia of Genes and Genomes (KEGG) analyses, which were performed using the clusterProfiler, enrichplot, and org.Hs.eg.db packages (P < 0.05 and FDR-adjusted q value [Benjamini–Hochberg method] < 0.05).
Immune Microenvironment Profiling
CIBERSORT is a reliable algorithm for estimating immune cell infiltration levels from gene expression data using the default LM22 signature matrix7. In this study, the CIBERSORT algorithm was used to estimate the extent of immune cell infiltration in samples from the IBD, CRC, and PAAD cohorts to explore the shared characteristics of the immune microenvironment among the three diseases. The analysis was performed with 1,000 permutations to calculate P values for each sample, and quantile normalization (QN = TRUE) was applied to the mixture expression file. Only samples with a CIBERSORT P value < 0.05 were retained for subsequent analyses, ensuring the reliability of the deconvolution results.
Assessment of the Diagnostic Value of Common Genes
The diagnostic value of the common genes in the IBD, CRC, and PAAD cohorts was evaluated using receiver operating characteristic (ROC) analysis with the pROC package in R. The optimal trade-off between sensitivity and specificity was visualized using ROC curves.
qRT-PCR, Cell Transfection, and Colony Formation Assay
qRT-PCR and cell transfection were conducted according to previous studies8,9,10. Transient transfection was performed using jetPRIME transfection reagent (Polyplus, China) according to the manufacturer's instructions. Cells were incubated with the transfection mixture for 6 h, after which the medium was replaced with complete DMEM. Subsequent experiments were performed 48 h after transfection.
Briefly, total cellular RNA was extracted using TRIzol reagent. RNA was reverse-transcribed into cDNA using PrimeScript RT Master Mix. Quantitative PCR was performed using TB Green qPCR. β-actin was used as the internal reference gene for expression normalization. Biological experiments were performed in triplicate. The primer sequences and the siS100P sequence can be found in a previous study 11.
NCM460, FHC, HCT116, SW116, PANC1, and BXPC2 cells were obtained as listed in the Table of Materials. All cell lines underwent cell identification and mycoplasma testing.
During the experiments, all cells were passaged for 3–5 generations. All cells were cultured in complete DMEM containing 10% fetal bovine serum and 1% penicillin–streptomycin.
The colony formation assay was performed as described in a previous study12. Briefly, 1,000 cells were seeded into each well of a 6-well plate and cultured for 10 days before the experiment was terminated. Cells were fixed with 4% paraformaldehyde, stained with 0.1% crystal violet, and colonies were counted using ImageJ.
Analysis of Common Genes Based on scRNA-seq Data
scRNA-seq data from the IBD dataset (GSE214695), CRC dataset (GSE166555), and PAAD dataset (GSE154778) were preprocessed as described in previous studies8,13. Raw count matrices were collapsed by mean expression for duplicate gene symbols using limma::avereps. Initial filtering retained genes detected in at least three cells and cells containing at least 50 unique transcripts. Cells with a mitochondrial transcript fraction >5% or fewer than 50 detected genes were removed. Log normalization was performed with a scale factor of 10,000, followed by variance-stabilizing transformation to identify the top 1,500 highly variable genes, which were Z-score standardized prior to principal component analysis (PCA). Cluster-defining marker genes were filtered using log2(fold change) > 0.5, a detection fraction ≥0.25 in target clusters, and an adjusted P value <0.05.
Briefly, data preprocessing was conducted using the Seurat package, and cell type annotation was performed using the SingleR package (version 2.6.0). Cell clustering was performed in Seurat using k-nearest neighbor graph construction and t-SNE embedding based on PCA dimensions 1–20. The distribution and expression levels of the common genes across different cell types were then examined.
Construction of the IBD Model
According to previous studies14, lipopolysaccharide (LPS) was used to induce inflammation in normal human colonic epithelial cells (FHC and NCM460), thereby generating an inflammation-mimicking IBD model. Biological experiments were performed in triplicate. Cells were routinely cultured in a humidified incubator at 37°C with 5% CO2. When cell confluence reached approximately 50%–70%, the culture medium was replaced with fresh complete medium, and the cells were treated with 10 ng/mL LPS for 12 h. An equal volume of sterile phosphate-buffered saline (PBS) was used as the vehicle control. The culture volume was 2 mL per well in 6-well plates. After treatment, the medium was removed, the cells were washed twice with pre-cooled sterile PBS, and the cells were collected for subsequent analyses.
Statistical Analysis
All bioinformatics analyses were performed using R software (version 4.1.2). Comparisons between two groups were performed using Student's t-test, whereas comparisons among multiple groups were performed using one-way analysis of variance (ANOVA). Correlation analysis was performed using the Spearman method. All cellular experiments were repeated at least three times, and the data are presented as the mean ± standard deviation (SD). A P value or FDR < 0.05 was considered statistically significant. NS, not significant; P < 0.05 (*), P < 0.01 (**), and P < 0.001 (***).
Dostęp ograniczony. Zaloguj się lub rozpocznij wersję próbną, aby wyświetlić tę treść.
DEG w IBD, CRC i PAAD
W pierwszej kolejności zintegrowano kohorty IBD (GSE179285 i GSE24287), kohorty CRC (TCGA-CRC i GSE87211) oraz kohorty PAAD (GSE128735 i GSE62452), a następnie oceniono proces integracji za pomocą PCA oraz wykresów gęstości ekspresji genów. Wyniki wykazały, że efekty serii między różnymi zbiorami danych zostały skutecznie wyeliminowane po integracji (Rysunek 1A–F).
Dostęp ograniczony. Zaloguj się lub rozpocznij wersję próbną, aby wyświetlić tę treść.
Previous studies have highlighted the association of IBD, an immune-related gastrointestinal disorder, with several diseases, including neurodegenerative diseases17, multiple sclerosis18, amyotrophic lateral sclerosis19, endometriosis20, rheumatoid arthritis2, Hodgkin’s lymphoma21, CRC22, and PAAD23. IBD is generally considere...
Dostęp ograniczony. Zaloguj się lub rozpocznij wersję próbną, aby wyświetlić tę treść.
Conflict of Interest:
The authors declare no competing interests.
The authors acknowledge financial support from the National Key R&D Plan of China (Grant No. 2023YFB3210400).
Dostęp ograniczony. Zaloguj się lub rozpocznij wersję próbną, aby wyświetlić tę treść.
| Nazwa | Firma | Numer katalogowy | Komentarze |
|---|---|---|---|
| BXPC2 cell line | Cell Bank of the Chinese Academy of Sciences | N/A | Human pancreatic adenocarcinoma cell line |
| CIBERSORT | N/A | N/A | Immune cell infiltration analysis |
| clusterProfiler package | Bioconductor | v4.8.0 | Gene Ontology (GO) and Kyoto Encyclopedia of Genes and Genomes (KEGG) enrichment analyses |
| Dulbecco's Modified Eagle Medium (DMEM) | Thermo Fisher Scientific | 11965092 | Cell culture medium |
| Fetal bovine serum (FBS) | Thermo Fisher Scientific | A5256701 | Cell culture supplement |
| FHC cell line | Cell Bank of the Chinese Academy of Sciences | N/A | Human normal colonic epithelial cell line |
| GEO dataset (GSE128735) | Gene Expression Omnibus (GEO) | N/A | Public gene expression dataset for pancreatic adenocarcinoma |
| GEO dataset (GSE154778) | Gene Expression Omnibus (GEO) | N/A | Public single-cell RNA sequencing dataset for pancreatic adenocarcinoma |
| GEO dataset (GSE166555) | Gene Expression Omnibus (GEO) | N/A | Public single-cell RNA sequencing dataset for colorectal cancer |
| GEO dataset (GSE179285) | Gene Expression Omnibus (GEO) | N/A | Public gene expression dataset for inflammatory bowel disease |
| GEO dataset (GSE214695) | Gene Expression Omnibus (GEO) | N/A | Public single-cell RNA sequencing dataset for inflammatory bowel disease |
| GEO dataset (GSE24287) | Gene Expression Omnibus (GEO) | N/A | Public gene expression dataset for inflammatory bowel disease |
| GEO dataset (GSE62452) | Gene Expression Omnibus (GEO) | N/A | Public gene expression dataset for pancreatic adenocarcinoma |
| GEO dataset (GSE87211) | Gene Expression Omnibus (GEO) | N/A | Public gene expression dataset for colorectal cancer |
| ggplot2 package | CRAN | v3.4.2 | Data visualization |
| HCT116 cell line | Cell Bank of the Chinese Academy of Sciences | N/A | Human colorectal cancer cell line |
| ImageJ | National Institutes of Health (NIH) | v1.8.0 | Colony counting |
| jetPRIME Transfection Reagent | Polyplus | 101000046 | Cell transfection |
| Lipopolysaccharide (LPS) | Beyotime Co., Ltd | S1735 | Induction of an inflammation-mimicking IBD cell model |
| limma package | Bioconductor | v3.54.0 | Differential expression analysis |
| NCM460 cell line | Cell Bank of the Chinese Academy of Sciences | N/A | Human normal colonic epithelial cell line |
| org.Hs.eg.db package | Bioconductor | v3.23.1 | Gene annotation for enrichment analysis |
| PANC1 cell line | Cell Bank of the Chinese Academy of Sciences | N/A | Human pancreatic adenocarcinoma cell line |
| Penicillin–streptomycin | Thermo Fisher Scientific | 15140-122 | Antibiotic supplement for cell culture |
| pheatmap package | CRAN | v1.0.12 | Heatmap visualization |
| pROC package | CRAN | v1.19.0.1 | Receiver operating characteristic (ROC) analysis |
| PrimeScript RT Master Mix | Takara Bio | RR036A | Reverse transcription of RNA into cDNA |
| Primer sets for qRT-PCR | Tsingke Biotech Co., Ltd | N/A | Primer sequences reported in Reference 11 |
| R software | R Foundation for Statistical Computing | v4.1.2 | Statistical and bioinformatics analyses |
| Seurat package | CRAN | v4.0 | Single-cell RNA sequencing data preprocessing and analysis |
| siS100P | Tsingke Biotech Co., Ltd | N/A | Small interfering RNA targeting S100P |
| SingleR package | Bioconductor | v2.6.0 | Cell type annotation for single-cell RNA sequencing |
| STRING database | STRING Consortium | N/A | Protein-protein interaction analysis |
| SW1116 cell line | Cell Bank of the Chinese Academy of Sciences | N/A | Human colorectal cancer cell line |
| TB Green qPCR Mix | Takara Bio | RR430B | Quantitative real-time PCR |
| TCGA-CRC dataset | The Cancer Genome Atlas (TCGA) | N/A | Public colorectal cancer transcriptomic dataset |
| TRIzol reagent | Invitrogen | 15596-026 | Total RNA extraction |
| WGCNA package | CRAN | v1.73 | Weighted gene co-expression network analysis |
Dostęp ograniczony. Zaloguj się lub rozpocznij wersję próbną, aby wyświetlić tę treść.