Artykuł badawczy

S100P as a Shared Biomarker in Inflammatory Bowel Disease, Colorectal Cancer, and Pancreatic Adenocarcinoma: An Integrated Transcriptomic Analysis

80 wyświetleń

DOI:

10.3791/71735

11 sierpnia 2026

* These authors contributed equally

W tym artykule

Podsumowanie

This protocol integrates public transcriptomic datasets and colonic epithelial cell validation to identify S100P as a shared biomarker associated with inflammatory bowel disease, colorectal cancer, and pancreatic adenocarcinoma.

Streszczenie

Inflammatory bowel disease (IBD) is associated with an increased risk of colorectal cancer (CRC) and pancreatic adenocarcinoma (PAAD), yet the molecular features shared among these diseases remain incompletely understood. This study aimed to identify common genes and biological pathways associated with IBD, CRC, and PAAD through integrated transcriptomic analysis and experimental validation. Gene expression datasets for IBD, CRC, and PAAD were obtained from The Cancer Genome Atlas and Gene Expression Omnibus databases. Weighted gene co-expression network analysis and differential expression analysis were performed to identify disease-associated and shared genes. Gene Ontology and Kyoto Encyclopedia of Genes and Genomes (analyses were used to explore enriched biological functions and pathways. Immune cell infiltration was evaluated using Cell-type Identification by Estimating Relative Subsets of RNA Transcripts. Receiver operating characteristic analysis was performed to assess the diagnostic performance of common genes. Single-cell RNA sequencing analysis was conducted to examine the cellular distribution of S100P. In addition, the effects of S100P downregulation were evaluated in lipopolysaccharide (LPS)-stimulated colonic epithelial cells. A total of 162 disease-associated genes and four common genes were identified. Functional enrichment analyses indicated significant enrichment of immune- and inflammation-related pathways, including the interleukin-17 signaling pathway. Immune infiltration analysis revealed similar trends in several immune cell populations across IBD, CRC, and PAAD. Single-cell analysis showed elevated S100P expression in epithelial cells from all three diseases. Downregulation of S100P restored the proliferative capacity of LPS-stimulated colonic epithelial cells and reduced inflammatory cytokine expression. Integrated transcriptomic analysis identified S100P as a biomarker associated with IBD, CRC, and PAAD and highlighted shared immune-related features across these diseases.

Wprowadzenie

Inflammatory bowel disease (IBD) represents a spectrum of immune-related disorders that affect the gastrointestinal tract, including ulcerative colitis and Crohn’s disease1. The etiology of IBD is highly complex, involving mucosal immune abnormalities, dysbiosis, and genetic susceptibility2. IBD is a global health concern with increasing incidence rates and substantial economic burdens, and its growing prevalence has attracted considerable attention1. Importantly, patients with IBD have a significantly increased risk of developing colorectal cancer (CRC)3 and pancreatic adenocarcinoma (PAAD)4. Although this association has been well documented, the genetic crosstalk among IBD, CRC, and PAAD remains incompletely understood.

Previous studies strongly suggest that IBD, CRC, and PAAD share common pathogenic processes; however, specific and sensitive diagnostic biomarkers remain lacking, and the shared pathogenic mechanisms have not been fully elucidated. Fortunately, with the rapid advancement and widespread availability of high-throughput sequencing technologies, numerous transcriptomic datasets from patients with IBD, CRC, and PAAD have become publicly available, enabling systematic investigation of the molecular interconnections among these diseases2,5,6.

This study employed high-throughput sequencing data from patients with IBD, CRC, and PAAD to identify disease-associated and shared genes using weighted gene co-expression network analysis (WGCNA) and differential expression analysis. These genes were further investigated to identify potential shared signaling pathways among IBD, CRC, and PAAD. Additionally, we evaluated the diagnostic value of these common genes. Single-cell RNA sequencing (scRNA-seq) analysis demonstrated that the key gene S100P was predominantly expressed in epithelial cells. Finally, we investigated the biological role of S100P in IBD.

In conclusion, this study aimed to identify common diagnostic biomarkers and biological pathways associated with IBD, CRC, and PAAD, thereby providing valuable clinical insights into the shared prevention and treatment of these diseases.

Protokół

This study used publicly available, de-identified datasets from The Cancer Genome Atlas (TCGA) and the Gene Expression Omnibus (GEO), as well as established commercial cell lines. No newly recruited human participants, identifiable patient information, or patient-derived samples were involved. All analyses were performed in accordance with relevant institutional guidelines and the terms of use of the public databases. Therefore, additional institutional ethical approval and informed consent were not required for this study.

Data Source
RNA-seq data for the IBD cohorts (GSE179285; platform: GPL6480 and GSE24287; platform: GPL6480), CRC cohorts (TCGA-CRC; platform: Illumina HiSeq 2000 and GSE87211; platform: GPL13497), and PAAD cohorts (GSE128735; platform: GPL20301 and GSE62452; platform: GPL6244) were downloaded from TCGA and the GEO. All datasets were accessed on December 5, 2025.

For each dataset, samples were strictly divided into two subgroups, with disease lesion/tumor tissues serving as the case group and corresponding non-lesion normal tissues serving as the control group. Specifically, the IBD cohort contained 297 intestinal mucosal samples from patients with IBD and 56 normal intestinal mucosal samples from healthy individuals; the CRC cohort included 841 primary colorectal tumor tissues and 211 matched adjacent normal colorectal epithelial tissues; and the PAAD cohort consisted of 114 PAAD tumor tissues and 106 normal pancreatic parenchymal tissues.

All datasets within the same disease category were integrated uniformly. The limma package’s normalizeBetweenArrays function was applied to perform cross-sample quantile normalization, effectively eliminating inter-platform batch effects and standardizing gene expression values across different datasets for subsequent differential expression analysis.

Screening of IBD-, CRC-, and PAAD-Related Genes and Common Genes
First, differentially expressed genes (DEGs) were screened from the IBD, CRC, and PAAD cohorts using the limma package, and the original P values were corrected using the Benjamini-Hochberg false discovery rate (FDR) method. In the IBD, CRC, and PAAD cohorts, the screening criteria were set at |logFC| > 0.4 and P < 0.05. Additionally, WGCNA was performed on all genes, with a minimum module gene threshold of 100 (soft-threshold power = 0.90; network type = signed). Consequently, common DEGs and module genes were identified across the three cohorts. Genes consistently identified by both methods were defined as common genes, whereas the remaining genes were categorized as related genes.

PPI and Functional Enrichment Analysis
These analyses were performed on the disease-associated genes. Protein–protein interaction (PPI) analysis was conducted using the STRING database (interaction score > 0.40). Functional enrichment analysis included Gene Ontology (GO) and Kyoto Encyclopedia of Genes and Genomes (KEGG) analyses, which were performed using the clusterProfiler, enrichplot, and org.Hs.eg.db packages (P < 0.05 and FDR-adjusted q value [Benjamini–Hochberg method] < 0.05).

Immune Microenvironment Profiling
CIBERSORT is a reliable algorithm for estimating immune cell infiltration levels from gene expression data using the default LM22 signature matrix7. In this study, the CIBERSORT algorithm was used to estimate the extent of immune cell infiltration in samples from the IBD, CRC, and PAAD cohorts to explore the shared characteristics of the immune microenvironment among the three diseases. The analysis was performed with 1,000 permutations to calculate P values for each sample, and quantile normalization (QN = TRUE) was applied to the mixture expression file. Only samples with a CIBERSORT P value < 0.05 were retained for subsequent analyses, ensuring the reliability of the deconvolution results.

Assessment of the Diagnostic Value of Common Genes
The diagnostic value of the common genes in the IBD, CRC, and PAAD cohorts was evaluated using receiver operating characteristic (ROC) analysis with the pROC package in R. The optimal trade-off between sensitivity and specificity was visualized using ROC curves.

qRT-PCR, Cell Transfection, and Colony Formation Assay
qRT-PCR and cell transfection were conducted according to previous studies8,9,10. Transient transfection was performed using jetPRIME transfection reagent (Polyplus, China) according to the manufacturer's instructions. Cells were incubated with the transfection mixture for 6 h, after which the medium was replaced with complete DMEM. Subsequent experiments were performed 48 h after transfection.

Briefly, total cellular RNA was extracted using TRIzol reagent. RNA was reverse-transcribed into cDNA using PrimeScript RT Master Mix. Quantitative PCR was performed using TB Green qPCR. β-actin was used as the internal reference gene for expression normalization. Biological experiments were performed in triplicate. The primer sequences and the siS100P sequence can be found in a previous study 11.

NCM460, FHC, HCT116, SW116, PANC1, and BXPC2 cells were obtained as listed in the Table of Materials. All cell lines underwent cell identification and mycoplasma testing.

During the experiments, all cells were passaged for 3–5 generations. All cells were cultured in complete DMEM containing 10% fetal bovine serum and 1% penicillin–streptomycin.

The colony formation assay was performed as described in a previous study12. Briefly, 1,000 cells were seeded into each well of a 6-well plate and cultured for 10 days before the experiment was terminated. Cells were fixed with 4% paraformaldehyde, stained with 0.1% crystal violet, and colonies were counted using ImageJ.

Analysis of Common Genes Based on scRNA-seq Data
scRNA-seq data from the IBD dataset (GSE214695), CRC dataset (GSE166555), and PAAD dataset (GSE154778) were preprocessed as described in previous studies8,13. Raw count matrices were collapsed by mean expression for duplicate gene symbols using limma::avereps. Initial filtering retained genes detected in at least three cells and cells containing at least 50 unique transcripts. Cells with a mitochondrial transcript fraction >5% or fewer than 50 detected genes were removed. Log normalization was performed with a scale factor of 10,000, followed by variance-stabilizing transformation to identify the top 1,500 highly variable genes, which were Z-score standardized prior to principal component analysis (PCA). Cluster-defining marker genes were filtered using log2(fold change) > 0.5, a detection fraction ≥0.25 in target clusters, and an adjusted P value <0.05.

Briefly, data preprocessing was conducted using the Seurat package, and cell type annotation was performed using the SingleR package (version 2.6.0). Cell clustering was performed in Seurat using k-nearest neighbor graph construction and t-SNE embedding based on PCA dimensions 1–20. The distribution and expression levels of the common genes across different cell types were then examined.

Construction of the IBD Model
According to previous studies14, lipopolysaccharide (LPS) was used to induce inflammation in normal human colonic epithelial cells (FHC and NCM460), thereby generating an inflammation-mimicking IBD model. Biological experiments were performed in triplicate. Cells were routinely cultured in a humidified incubator at 37°C with 5% CO2. When cell confluence reached approximately 50%–70%, the culture medium was replaced with fresh complete medium, and the cells were treated with 10 ng/mL LPS for 12 h. An equal volume of sterile phosphate-buffered saline (PBS) was used as the vehicle control. The culture volume was 2 mL per well in 6-well plates. After treatment, the medium was removed, the cells were washed twice with pre-cooled sterile PBS, and the cells were collected for subsequent analyses.

Statistical Analysis
All bioinformatics analyses were performed using R software (version 4.1.2). Comparisons between two groups were performed using Student's t-test, whereas comparisons among multiple groups were performed using one-way analysis of variance (ANOVA). Correlation analysis was performed using the Spearman method. All cellular experiments were repeated at least three times, and the data are presented as the mean ± standard deviation (SD). A P value or FDR < 0.05 was considered statistically significant. NS, not significant; P < 0.05 (*), P < 0.01 (**), and P < 0.001 (***).

Wyniki

DEG w IBD, CRC i PAAD
W pierwszej kolejności zintegrowano kohorty IBD (GSE179285 i GSE24287), kohorty CRC (TCGA-CRC i GSE87211) oraz kohorty PAAD (GSE128735 i GSE62452), a następnie oceniono proces integracji za pomocą PCA oraz wykresów gęstości ekspresji genów. Wyniki wykazały, że efekty serii między różnymi zbiorami danych zostały skutecznie wyeliminowane po integracji (Rysunek 1A–F).

Wykresy PCA korekcji efektu serii i wykresy ekspresji; porównanie przed i po analizie; normalizacja ekspresji genów.
Rycina 1. Normalizacja kohort choroby zapalnej jelit (IBD), raka jelita grubego (CRC) oraz gruczolakoraka trzustki (PAAD). (A) Wykresy analizy głównych składowych (PCA) przed i po normalizacji kohort IBD (GSE179285 i GSE24287). (B) Wykresy rozkładu ekspresji genów przed i po normalizacji kohort IBD. (C) Wykresy PCA przed i po normalizacji kohort CRC (TCGA-CRC i GSE87211). (D) Wykresy rozkładu ekspresji genów przed i po normalizacji kohort CRC. (E) Wykresy PCA przed i po normalizacji kohort PAAD (GSE62452 i GSE128735). (F) Wykresy rozkładu ekspresji genów przed i po normalizacji kohort PAAD. Kliknij tutaj, aby wyświetlić powiększoną wersję tej ryciny.

Po integracji zbiorów danych i korekcji efektu serii przeprowadzono analizę różnicowej ekspresji w kohortach IBD, CRC i PAAD. W obrębie każdej kohorty porównano tkanki z uszkodzeniami chorobowymi lub tkanki nowotworowe z dopasowanymi tkankami prawidłowymi w celu zidentyfikowania DEG. W kohorcie IBD zidentyfikowano 183 DEG, w tym 64 geny o obniżonej i 119 genów o podwyższonej ekspresji (Rycyna 2A). W kohorcie CRC zidentyfikowano 5 064 DEG, w tym 2 477 genów o obniżonej i 2 587 genów o podwyższonej ekspresji (Rycina 2B). W kohorcie PAAD zidentyfikowano 2 293 DEG, w tym 901 genów o obniżonej i 1 392 genów o podwyższonej ekspresji (Rycina 2C). Ostatecznie, spośród kohort IBD, CRC i PAAD zidentyfikowano 40 wspólnych DEG (Rycina 2D).

Mapy ciepła ekspresji genów, wykresy wulkaniczne i diagram Venna dla analizy różnicowej w sekwencjonowaniu RNA.
Rycina 2. Analiza ekspresji różnicowej w kohortach IBD, CRC i PAAD. (A) Mapa ciepła i wykres wulkaniczny genów o różnej ekspresji (DEGs) w kohorcie IBD. (B) Mapa ciepła i wykres wulkaniczny DEGs w kohorcie CRC. (C) Mapa ciepła i wykres wulkaniczny DEGs w kohorcie PAAD. (D) Diagram Venna przedstawiający wspólne DEGs w kohortach IBD, CRC i PAAD. Kliknij tutaj, aby wyświetlić powiększoną wersję tej ryciny.

WGCNA w IBD, CRC i PAAD
Analizę WGCNA przeprowadzono dla kohort IBD, CRC i PAAD. W kohorcie IBD zidentyfikowano trzy moduły ściśle powiązane z cechami klinicznymi, przy czym moduły MEbrown i MEturquoise wykazały najsilniejsze korelacje (Rycina 3A). Podobnie w kohorcie CRC zidentyfikowano osiem modułów ściśle powiązanych z cechami klinicznymi, z których najsilniejsze korelacje wykazały moduły MEbrown i MEturquoise (Rycina 3B). W kohorcie PAAD zidentyfikowano jeden moduł ściśle powiązany z cechami klinicznymi, a najsilniejsze korelacje zaobserwowano dla modułów MEblack i MEbrown (Rycina 3C).

Diagramy relacji między modułami genów a cechami oraz diagram Venna dla wyników analiz IBD, CRC i PAAD.
Rycina 3. Ważona analiza sieci współekspresji genów (WGCNA) kohort IBD, CRC i PAAD. (A) Mapa ciepła przedstawiająca korelacje między modułami współekspresji genów a cechami klinicznymi w kohorcie IBD. (B) Mapa ciepła przedstawiająca korelacje między modułami współekspresji genów a cechami klinicznymi w kohorcie CRC. (C) Mapa ciepła przedstawiająca korelacje między modułami współekspresji genów a cechami klinicznymi w kohorcie PAAD. (D) Diagram Venna przedstawiający wspólne geny zidentyfikowane w kluczowych modułach współekspresji w kohortach IBD, CRC i PAAD. Kliknij tutaj, aby wyświetlić powiększoną wersję tej ryciny.

Na podstawie genów z modułów istotnie skorelowanych z cechami klinicznymi zidentyfikowano 434 geny powiązane z modułami, które potencjalnie biorą udział w IBD, CRC oraz PAAD (Ryc. 3D).

Analiza wzbogacenia funkcjonalnego wspólnych genów w IBD, CRC i PAAD
Poprzez połączenie wyników analizy różnicowej ekspresji i WGCNA zidentyfikowano 40 wspólnych DEG oraz 122 wspólne geny powiązane z modułami. Aby zbadać wspólne charakterystyki molekularne IBD, CRC i PAAD, geny te zintegrowano w celu dalszej analizy, co pozwoliło na wyłonienie 158 genów związanych z tymi chorobami.

W pierwszej kolejności przeanalizowano 158 genów przy użyciu bazy danych STRING w celu skonstruowania sieci PPI (Rycina 4A). Następnie przeprowadzono analizy wzbogacenia GO i KEGG. Analiza GO wykazała istotne wzbogacenie w procesach biologicznych, w tym w procesach metabolicznych hormonów (Rycina 4B,C), natomiast analiza KEGG wykazała wzbogacenie w szlakach, w tym w szlaku sygnalizacyjnym interleukiny-17 (IL-17) oraz szlaku sygnalizacyjnym receptora aktywowanego przez proliferatory peroksysomów (PPAR) (Rycina 4D,E).

Analiza sieci, wykres interakcji genów. Zawiera węzły połączone ścieżkami, wizualizacje danych.
Rycina 4. Analiza wzbogacenia funkcjonalnego genów związanych z chorobami zidentyfikowanych w IBD, CRC i PAAD. (A) Sieć interakcji białko-białko (PPI) genów związanych z chorobami. (B) Analiza wzbogacenia Gene Ontology (GO) przedstawiona w formie wykresu kropkowego dla wzbogaconych terminów procesów biologicznych, komponentów komórkowych i funkcji molekularnych. (C) Sieć gen-koncepcja GO pokazująca relacje między wzbogaconymi terminami GO a powiązanymi genami. (D) Analiza wzbogacenia ścieżek Kyoto Encyclopedia of Genes and Genomes (KEGG) przedstawiona w formie wykresu kropkowego. (E) Sieć gen-ścieżka KEGG pokazująca relacje między wzbogaconymi ścieżkami a powiązanymi genami. Kliknij tutaj, aby wyświetlić większą wersję tej ryciny.

Analiza korelacji wspólnych genów z infiltracją komórek odpornościowych w IBD, CRC i PAAD
Do oszacowania poziomów infiltracji komórek odpornościowych w kohortach IBD, CRC i PAAD wykorzystano algorytm CIBERSORT. W kohorcie IBD zaobserwowano różnice w poziomach infiltracji komórek odpornościowych pomiędzy próbkami prawidłowymi a próbkami z IBD (Rysunek 5A) oraz oceniono korelacje pomiędzy wspólnymi genami a populacjami komórek odpornościowych (Rysunek 5B). Podobnie w kohorcie CRC zaobserwowano różnice w poziomach infiltracji komórek odpornościowych pomiędzy próbkami prawidłowymi a próbkami z CRC (Rysunek 5C) oraz oceniono korelacje pomiędzy wspólnymi genami a populacjami komórek odpornościowych (Rysunek 5D). W kohorcie PAAD również zaobserwowano różnice w poziomach infiltracji komórek odpornościowych pomiędzy próbkami prawidłowymi a próbkami z PAAD (Rysunek 5E) oraz oceniono korelacje pomiędzy wspólnymi genami a populacjami komórek odpornościowych (Rysunek 5F).

Wykresy skrzypcowe i diagramy korelacji ekspresji białek w grupach normalnych, IBD, CRC i grupach leczonych.
Rysunek 5. Analiza infiltracji komórek odpornościowych w kohortach IBD, CRC i PAAD. (A) Wykresy skrzypcowe przedstawiające szacowane proporcje populacji komórek odpornościowych w kohorcie IBD. (B) Analiza korelacji między wspólnymi genami a populacjami komórek odpornościowych w kohorcie IBD, obejmująca mapę ciepła korelacji komórek odpornościowych oraz sieć powiązań gen-komórka odpornościowa. (C) Wykresy skrzypcowe przedstawiające szacowane proporcje populacji komórek odpornościowych w kohorcie CRC. (D) Analiza korelacji między wspólnymi genami a populacjami komórek odpornościowych w kohorcie CRC, obejmująca mapę ciepła korelacji komórek odpornościowych oraz sieć powiązań gen-komórka odpornościowa. (E) Wykresy skrzypcowe przedstawiające szacowane proporcje populacji komórek odpornościowych w kohorcie PAAD. (F) Analiza korelacji między wspólnymi genami a populacjami komórek odpornościowych w kohorcie PAAD, obejmująca mapę ciepła korelacji komórek odpornościowych oraz sieć powiązań gen-komórka odpornościowa. Kliknij tutaj, aby wyświetlić powiększoną wersję tego rysunku.

Ocena potencjalnej wartości wspólnych genów w IBD, CRC i PAAD
Wzorce ekspresji wspólnych genów poddano dalszej ocenie w kohortach IBD, CRC i PAAD. Gen S100P był konsekwentnie nadekspresyjny we wszystkich trzech kohortach (Rysunek 6A,D,G). Ponadto, za pomocą analizy ROC oceniono wydajność diagnostyczną wspólnych genów.

Wykresy pudełkowe analizy ekspresji genów, krzywe ROC; czułość, swoistość w wynikach badania biomarkerów nowotworowych.
Rycina 6. Wydajność diagnostyczna wspólnych genów w kohortach IBD, CRC i PAAD. (A) Poziomy ekspresji FXYD3, S100P, PLA2G2A i MUC1 w kohorcie IBD. (B) Krzywe charakterystyki operacyjnej odbiornika (ROC) przedstawiające wydajność diagnostyczną poszczególnych wspólnych genów w kohorcie IBD. (C) Krzywa ROC przedstawiająca wydajność diagnostyczną połączonego modelu diagnostycznego w kohorcie IBD. (D) Poziomy ekspresji FXYD3, S100P, PLA2G2A i MUC1 w kohorcie CRC. (E) Krzywe ROC przedstawiające wydajność diagnostyczną poszczególnych wspólnych genów w kohorcie CRC. (F) Krzywa ROC przedstawiająca wydajność diagnostyczną połączonego modelu diagnostycznego w kohorcie CRC. (G) Poziomy ekspresji FXYD3, S100P, PLA2G2A i MUC1 w kohorcie PAAD. (H) Krzywe ROC przedstawiające wydajność diagnostyczną poszczególnych wspólnych genów w kohorcie PAAD. (I) Krzywa ROC przedstawiająca wydajność diagnostyczną połączonego modelu diagnostycznego w kohorcie PAAD. Wartości pola pod krzywą (AUC) oraz odpowiadające im 95% przedziały ufności są podane w odpowiednich miejscach. Kliknij tutaj, aby wyświetlić powiększoną wersję tej ryciny.

W kohorcie IBD wartości pola pod krzywą (AUC) wynosiły 0,626 dla FXYD3, 0,597 dla S100P, 0,670 dla PLA2G2A oraz 0,697 dla MUC1, natomiast łączny model diagnostyczny wykazał AUC na poziomie 0,815 (Rysunek 6B,C). Podobnie, w kohorcie CRC wartości AUC wynosiły 0,839 dla FXYD3, 0,738 dla S100P, 0,716 dla PLA2G2A oraz 0,672 dla MUC1, podczas gdy łączny model diagnostyczny dał AUC wynoszące 0,925 (Rysunek 6E,F). W kohorcie PAAD wartości AUC wynosiły 0,852 dla FXYD3, 0,896 dla S100P, 0,637 dla PLA2G2A oraz 0,733 dla MUC1, natomiast łączny model diagnostyczny uzyskał AUC na poziomie 0,901 (Rysunek 6H,I).

Analiza scRNA-seq w oparciu o wspólne geny
Po wstępnym przetwarzaniu danych scRNA-seq z zestawu danych IBD zidentyfikowano 17 klastrów komórkowych i 8 typów komórek. Następnie oceniono rozkład wspólnych genów w różnych populacjach komórek, stwierdzając, że wspólne geny są ekspresjonowane głównie w komórkach nabłonkowych (Rysunek 7A). Podobnie, wstępne przetwarzanie zestawu danych CRC pozwoliło na zidentyfikowanie 20 klastrów komórkowych i 8 typów komórek, przy czym wspólne geny były również ekspresjonowane głównie w komórkach nabłonkowych (Rysunek 7B). Na koniec, wstępne przetwarzanie zestawu danych PAAD pozwoliło zidentyfikować 19 klastrów komórkowych i 7 typów komórek, a wspólne geny były analogicznie ekspresjonowane głównie w komórkach nabłonkowych (Rysunek 7C).

Wykresy UMAP przedstawiające klastrowanie typów komórek; panele A, B, C ilustrują ekspresję genów w IBD, CRC i PAAD.
Rycina 7Analiza sekwencjonowania RNA pojedynczych komórek (scRNA-seq) wspólnych genów w tkankach z IBD, CRC i PAAD. (A) Wizualizacja klastrów komórek w zbiorze danych IBD (GSE214695) z wykorzystaniem algorytmu UMAP (uniform manifold approximation and projection), odpowiadające im adnotacje typów komórek oraz wykresy cech (feature plots) przedstawiające ekspresję FXYD3, S100P, PLA2G2A i MUC1 w poszczególnych populacjach komórek. (BWizualizacja UMAP klastrów komórkowych w zbiorze danych CRC (GSE166555), odpowiadające im adnotacje typów komórek oraz wykresy cech (feature plots) przedstawiające ekspresję FXYD3, S100P, PLA2G2A i MUC1 w poszczególnych populacjach komórek.C) Wizualizacja UMAP klastrów komórek w zbiorze danych PAAD (GSE154778), odpowiadające im adnotacje typów komórek oraz wykresy cech przedstawiające ekspresję FXYD3, S100P, PLA2G2A i MUC1 w poszczególnych populacjach komórek. Skale kolorystyczne wskazują względne poziomy ekspresji genów. Kliknij tutaj, aby wyświetlić powiększoną wersję tej ryciny.

Funkcja biologiczna S100P w IBD
Biorąc pod uwagę stałą nadekspresję S100P w kohortach IBD, CRC i PAAD, a także wcześniejsze doniesienia opisujące jego rolę w CRC i PAAD15,16, w niniejszym badaniu dalej analizowano funkcję biologiczną S100P w IBD.

Po pierwsze, ekspresja S100P była znacząco zwiększona w modelu IBD indukowanym LPS, a wydajność wyciszenia siS100P została potwierdzona (Rysunek 8A,B). Ponadto traktowanie LPS zmniejszyło proliferację komórek, podczas gdy zahamowanie ekspresji S100P częściowo przywróciło proliferację komórek (Rysunek 8C). Co więcej, wyciszenie S100P znacząco zmniejszyło ekspresję IL-1β, IL-6 i TNF-α w liniach komórkowych FHC i NCM460 w modelu IBD (Rysunek 8D,E). Na koniec, obniżenie poziomu S100P zmniejszyło ekspresję IL17RA w komórkach nabłonkowych jelita grubego, komórkach CRC oraz komórkach PAAD (Rysunek 8F).
 

Wykresy słupkowe, testy kolonogenne, wykresy stężeń cytokin dla wpływu siRNA na komórki, odpowiedź zapalna.
Rycina 8. Obniżenie poziomu S100P osłabia indukowane lipopolisacharydem (LPS) reakcje zapalne w komórkach nabłonka okrężnicy. (A) Względna ekspresja mRNA S100P w komórkach FHC po stymulacji LPS i wyciszeniu S100P. (B) Względna ekspresja mRNA S100P w komórkach NCM460 po stymulacji LPS i wyciszeniu S100P. (C) Reprezentatywne obrazy tworzenia kolonii i ilościowe oznaczenie proliferacji komórek w komórkach FHC i NCM460 po stymulacji LPS i wyciszeniu S100P. (D) Pomiar poziomu IL-1β, IL-6 i TNF-α metodą ELISA w komórkach FHC po stymulacji LPS i wyciszeniu S100P. (E) Pomiar poziomu IL-1β, IL-6 i TNF-α metodą ELISA w komórkach NCM460 po stymulacji LPS i wyciszeniu S100P. (F) Względna ekspresja mRNA IL-17RA po wyciszeniu S100P w komórkach FHC, NCM460, HCT116, SW1116, PANC-1 i BxPC-3. Kliknij tutaj, aby wyświetlić powiększoną wersję tej ryciny.

   

Dostępność danych
Zbiory danych analizowane w niniejszym badaniu są publicznie dostępne w TCGA oraz w repozytorium GEO pod numerami dostępu TCGA-CRC, GSE179285, GSE24287, GSE87211, GSE128735, GSE62452, GSE214695, GSE166555 oraz GSE154778. W trakcie tego badania nie wygenerowano nowych zbiorów danych z sekwencjonowania.

Dyskusja

Previous studies have highlighted the association of IBD, an immune-related gastrointestinal disorder, with several diseases, including neurodegenerative diseases17, multiple sclerosis18, amyotrophic lateral sclerosis19, endometriosis20, rheumatoid arthritis2, Hodgkin’s lymphoma21, CRC22, and PAAD23. IBD is generally considered to increase the risk of developing both CRC24 and PAAD23. Although epidemiological associations among these diseases have been reported, their shared molecular characteristics remain incompletely understood. In this study, we integrated bulk transcriptomic datasets, scRNA-seq data, and cellular functional experiments to investigate shared molecular features among IBD, CRC, and PAAD. Our findings provide additional insight into potential molecular and immune-related characteristics shared across these disease settings.

In this study, we identified 158 disease-associated genes through differential expression analysis and WGCNA, comprising 40 common DEGs and 122 common module-associated genes. Functional enrichment analysis demonstrated significant enrichment of pathways including the IL-17 signaling pathway. Previous studies have reported important roles for IL-17 signaling in the progression of IBD, CRC, and PAAD25,26,27. These observations suggest that IL-17-related processes may represent a common biological feature across these diseases. However, the present study did not directly investigate the mechanistic relationship between S100P and IL-17 signaling, and additional functional studies are needed to clarify this association. Four common genes (FXYD3, S100P, PLA2G2A, and MUC1) were identified across the analyzed datasets. Previous studies have demonstrated that FXYD3 regulates the growth of PAAD cells28, S100P has been implicated in CRC and PAAD progression15,16, and MUC1 participates in the progression of IBD, CRC, and PAAD29,30,31. In contrast, PLA2G2A has been less extensively investigated across all three diseases. ROC analysis demonstrated that each of the four genes showed measurable diagnostic performance, whereas the combined diagnostic model achieved higher diagnostic performance than the individual genes. Our findings therefore extend previous observations by identifying these genes as shared molecular features across IBD, CRC, and PAAD.

Immune cell infiltration analysis demonstrated similar infiltration patterns for naïve B cells, resting natural killer (NK) cells, and M0 macrophages across the IBD, CRC, and PAAD cohorts. Furthermore, scRNA-seq analysis showed that FXYD3, S100P, PLA2G2A, and MUC1 were predominantly expressed in epithelial cells. These findings provide additional information regarding the cellular distribution of the identified genes and their potential association with immune-related characteristics. However, the present analyses do not establish direct interactions between epithelial cells and the immune microenvironment, and further mechanistic studies are required.

Finally, we investigated the biological role of S100P in an LPS-induced inflammation-mimicking IBD model. S100P expression was significantly increased following LPS stimulation, and siRNA-mediated knockdown effectively reduced its expression. Under the experimental conditions examined, S100P knockdown partially restored cell proliferation and reduced the expression of the inflammatory cytokines IL-1β, IL-6, and TNF-α. In addition, S100P downregulation was associated with reduced IL17RA expression. These findings support further investigation of S100P in inflammation-related cellular responses.

This study has several limitations. First, only loss-of-function experiments were performed, and overexpression or rescue experiments were not conducted; therefore, a direct causal role of S100P cannot be established. Second, S100P expression was evaluated primarily at the mRNA level, and protein-level validation was not performed. Third, the functional experiments were limited to an inflammation-mimicking cell model, whereas CRC and PAAD models were not experimentally investigated. Finally, although multiomics analyses strengthened the identification of shared molecular features, additional mechanistic studies and in vivo validation are required to further clarify the biological roles of S100P and the other shared genes.

In conclusion, this study identified shared molecular features, biological pathways, and immune-related characteristics across IBD, CRC, and PAAD. Four common genes (FXYD3, S100P, PLA2G2A, and MUC1) demonstrated diagnostic potential across the analyzed datasets, and S100P was further investigated in an inflammation-mimicking IBD model. These findings provide a foundation for future studies investigating the shared molecular mechanisms linking IBD, CRC, and PAAD.

Oświadczenia

Conflict of Interest:
The authors declare no competing interests.

Podziękowania

The authors acknowledge financial support from the National Key R&D Plan of China (Grant No. 2023YFB3210400).

Materiały

Lista materiałów użytych w tym artykule
NazwaFirmaNumer katalogowyKomentarze
BXPC2 cell lineCell Bank of the Chinese Academy of SciencesN/AHuman pancreatic adenocarcinoma cell line
CIBERSORTN/AN/AImmune cell infiltration analysis
clusterProfiler packageBioconductorv4.8.0Gene Ontology (GO) and Kyoto Encyclopedia of Genes and Genomes (KEGG) enrichment analyses
Dulbecco's Modified Eagle Medium (DMEM)Thermo Fisher Scientific11965092Cell culture medium
Fetal bovine serum (FBS)Thermo Fisher ScientificA5256701Cell culture supplement
FHC cell lineCell Bank of the Chinese Academy of SciencesN/AHuman normal colonic epithelial cell line
GEO dataset (GSE128735)Gene Expression Omnibus (GEO)N/APublic gene expression dataset for pancreatic adenocarcinoma
GEO dataset (GSE154778)Gene Expression Omnibus (GEO)N/APublic single-cell RNA sequencing dataset for pancreatic adenocarcinoma
GEO dataset (GSE166555)Gene Expression Omnibus (GEO)N/APublic single-cell RNA sequencing dataset for colorectal cancer
GEO dataset (GSE179285)Gene Expression Omnibus (GEO)N/APublic gene expression dataset for inflammatory bowel disease
GEO dataset (GSE214695)Gene Expression Omnibus (GEO)N/APublic single-cell RNA sequencing dataset for inflammatory bowel disease
GEO dataset (GSE24287)Gene Expression Omnibus (GEO)N/APublic gene expression dataset for inflammatory bowel disease
GEO dataset (GSE62452)Gene Expression Omnibus (GEO)N/APublic gene expression dataset for pancreatic adenocarcinoma
GEO dataset (GSE87211)Gene Expression Omnibus (GEO)N/APublic gene expression dataset for colorectal cancer
ggplot2 packageCRANv3.4.2Data visualization
HCT116 cell lineCell Bank of the Chinese Academy of SciencesN/AHuman colorectal cancer cell line
ImageJNational Institutes of Health (NIH)v1.8.0Colony counting
jetPRIME Transfection ReagentPolyplus101000046Cell transfection
Lipopolysaccharide (LPS)Beyotime Co., LtdS1735Induction of an inflammation-mimicking IBD cell model
limma packageBioconductorv3.54.0Differential expression analysis
NCM460 cell lineCell Bank of the Chinese Academy of SciencesN/AHuman normal colonic epithelial cell line
org.Hs.eg.db packageBioconductorv3.23.1Gene annotation for enrichment analysis
PANC1 cell lineCell Bank of the Chinese Academy of SciencesN/AHuman pancreatic adenocarcinoma cell line
Penicillin–streptomycinThermo Fisher Scientific15140-122Antibiotic supplement for cell culture
pheatmap packageCRANv1.0.12Heatmap visualization
pROC packageCRANv1.19.0.1Receiver operating characteristic (ROC) analysis
PrimeScript RT Master MixTakara BioRR036AReverse transcription of RNA into cDNA
Primer sets for qRT-PCRTsingke Biotech Co., LtdN/APrimer sequences reported in Reference 11
R softwareR Foundation for Statistical Computingv4.1.2Statistical and bioinformatics analyses
Seurat packageCRANv4.0Single-cell RNA sequencing data preprocessing and analysis
siS100PTsingke Biotech Co., LtdN/ASmall interfering RNA targeting S100P
SingleR packageBioconductorv2.6.0Cell type annotation for single-cell RNA sequencing
STRING databaseSTRING ConsortiumN/AProtein-protein interaction analysis
SW1116 cell lineCell Bank of the Chinese Academy of SciencesN/AHuman colorectal cancer cell line
TB Green qPCR MixTakara BioRR430BQuantitative real-time PCR
TCGA-CRC datasetThe Cancer Genome Atlas (TCGA)N/APublic colorectal cancer transcriptomic dataset
TRIzol reagentInvitrogen15596-026Total RNA extraction
WGCNA packageCRANv1.73Weighted gene co-expression network analysis

Bibliografia

  1. Xu R, Du W, Yang Q, Du A. ITGB2 related to immune cell infiltration as a potential therapeutic target of inflammatory bowel disease using bioinformatics and functional research. Journal of cellular and molecular medicine. 2024;28(15):e18501.
  2. Sun HW, Zhang X, Shen CC. The shared circulating diagnostic biomarkers and molecular mechanisms of systemic lupus erythematosus and inflammatory bowel disease. Frontiers in immunology. 2024;15:1354348.
  3. Faye AS, Holmer AK, Axelrad JE. Cancer in inflammatory bowel disease. Gastroenterology clinics of North America. 2022;51(3):649-66.
  4. Yu J, et al. Risk of hepato-pancreato-biliary cancer is increased by primary sclerosing cholangitis in patients with inflammatory bowel disease: A population-based cohort study. United European gastroenterology journal. 2022;10(2):212-24.
  5. Xia B, et al. Identification of potential shared gene signatures between gastric cancer and type 2 diabetes: A data-driven analysis. Frontiers in medicine. 2024;11:1382004.
  6. Luo Y, et al. Exploring the molecular mechanism of comorbidity of type 2 diabetes mellitus and colorectal cancer: Insights from bulk omics and single-cell sequencing validation. Biomolecules. 2024;14(6).
  7. Newman AM, et al. Robust enumeration of cell subsets from tissue expression profiles. Nature methods. 2015;12(5):453-7.
  8. Sun W, et al. Construction and validation of a novel senescence-related risk score can help predict the prognosis and tumor microenvironment of gastric cancer patients and determine that STK40 can affect the ROS accumulation and proliferation ability of gastric cancer cells. Frontiers in immunology. 2023;14:1259231.
  9. Hong J, et al. A zinc metabolism-related gene signature for predicting prognosis and characteristics of breast cancer. Front Immunol. 2023;14:1276280.
  10. Man KF, et al. CREB1-BCL2 drives mitochondrial resilience in RAS GAP-dependent breast cancer chemoresistance. Oncogene. 2025;44(16):1093-105.
  11. Zhou H, et al. S100P promotes trophoblast syncytialization during early placenta development by regulating YAP1. Frontiers in endocrinology. 2022;13:860261.
  12. Zhou C, et al. Novel exosome-associated LncRNA model predicts colorectal cancer prognosis and drug response. Hereditas. 2025;162(1):79.
  13. Hong J, et al. Integrative PANoptosis-focused omics analysis uncovers GSDMC as a candidate biomarker in breast cancer. Frontiers in Cell and Developmental Biology. 2026;Volume 14 - 2026.
  14. Qiu C, et al. Hsa_circ_0004662 accelerates the progression of ulcerative colitis via the microRNA-532/HMGB3 signalling axis. Journal of cellular and molecular medicine. 2025;29(6):e70430.
  15. Schmid F, et al. Calcium-binding protein S100P is a new target gene of MACC1, drives colorectal cancer metastasis and serves as a prognostic biomarker. British journal of cancer. 2022;127(4):675-85.
  16. Arumugam T, Simeone DM, Van Golen K, Logsdon CD. S100P promotes pancreatic cancer growth, survival, and invasion. Clinical cancer research: an official journal of the American Association for Cancer Research. 2005;11(15):5356-64.
  17. Zong J, et al. The two-directional prospective association between inflammatory bowel disease and neurodegenerative disorders: A systematic review and meta-analysis based on longitudinal studies. Frontiers in immunology. 2024;15:1325908.
  18. Yaqubi K, et al. Inflammatory bowel disease is associated with an increase in the incidence of multiple sclerosis: A retrospective cohort study of 24,934 patients. European journal of medical research. 2024;29(1):186.
  19. Li CY, et al. Genome-wide genetic links between amyotrophic lateral sclerosis and autoimmune diseases. BMC medicine. 2021;19(1):27.
  20. Shigesi N, et al. The association between endometriosis and autoimmune diseases: A systematic review and meta-analysis. Human reproduction update. 2019;25(4):486-503.
  21. Sonnenberg A, Duong HT, McCarty DJ, El-Serag HB. Concurrence of inflammatory bowel disease with multiple sclerosis or Hodgkin lymphoma. European journal of gastroenterology & hepatology. 2023;35(12):1349-53.
  22. Chacon-Millan P, et al. A combination of microarray-based profiling and biocomputational analysis identified miR331-3p and hsa-let-7d-5p as potential biomarkers of ulcerative colitis progression to colorectal cancer. International journal of molecular sciences. 2024;25(11).
  23. Everhov Å H, et al. Inflammatory bowel disease and pancreatic cancer: A Scandinavian register-based cohort study 1969-2017. Alimentary pharmacology & therapeutics. 2020;52(1):143-54.
  24. Contran N, et al. Colorectal cancer and inflammatory bowel diseases share common salivary proteomic pathways. Scientific reports. 2024;14(1):17711.
  25. Sun W, et al. Osthole pretreatment alleviates TNBS-induced colitis in mice via both cAMP/PKA-dependent and independent pathways. Acta pharmacologica Sinica. 2017;38(8):1120-8.
  26. Li SY, et al. Diosgenin exerts anti-tumor effects through inactivation of cAMP/PKA/CREB signaling pathway in colorectal cancer. European journal of pharmacology. 2021;908:174370.
  27. Castro-Pando S, et al. Pancreatic epithelial IL17/IL17RA signaling drives B7-H4 expression to promote tumorigenesis. Cancer immunology research. 2024;12(9):1170-83.
  28. Kayed H, et al. FXYD3 is overexpressed in pancreatic ductal adenocarcinoma and influences pancreatic cancer cell growth. International journal of cancer. 2006;118(1):43-54.
  29. Wang H, et al. Mesenchymal stem cells ameliorate DSS-induced experimental colitis by modulating the gut microbiota and MUC-1 pathway. Journal of inflammation research. 2023;16:2023-39.
  30. Li W, et al. MUC1-C drives stemness in progression of colitis to colorectal cancer. JCI insight. 2020;5(12).
  31. Murthy D, et al. The MUC1-HIF-1α signaling axis regulates pancreatic cancer pathogenesis through polyamine metabolism remodeling. Proceedings of the National Academy of Sciences of the United States of America. 2024;121(14):e2315509121.
  32. Sun W, et al. TRIM47 regulates energy metabolism via glycolytic reprogramming to drive hepatocellular carcinoma progression and represents an efficient therapeutic target. Adv Sci (Weinh). 2026;13(17):e16996.
  33. Sun W, et al. METTL4 enhances GLI1 translation through m(6)Am modification to promote tumor progression as a therapeutic target for hepatocellular carcinoma. J Adv Res. 2026.
  34. Yuan Y, et al. RNA nanotherapeutics for hepatocellular carcinoma treatment. Theranostics. 2025;15(3):965-92.

Przedruki i uprawnienia

Tagi

Badania nad rakiemnumer 234numer 234nieswoiste zapalenia jelitrak jelita grubegogruczolakorak trzustkidiagnostykaodporno ciowyS100P