The study was approved by the Ethical Committee of Hainan Medical University (approval number HMC1984.24) and was in accordance with the Declaration of Helsinki (as revised in 2013).
Study subjects
A study of 50 lung cancer patients diagnosed between 2017 and 2024 at Sanya Central Hospital in Hainan Province was conducted using the Chinese Medical Association Clinical Guidelines for the Diagnosis and Treatment of Lung Cancer (2024 Edition). CSCO staging (2024) aligned with the AJCC 8th Edition for cross-study comparability, as validated in Chinese cohorts12, with two independent oncologist reviews to ensure consistency. The patient selection process, tissue availability, RNA quality assessment, and final sample inclusion for qRT-PCR analysis are summarized in Figure 1.
Study data
This study used bioinformatics to analyze YTHDC2 expression levels in NSCLC and their relationship with clinicopathological features, offering insights into potential mechanisms. Bioinformatics analyses were performed exclusively using publicly available RNA sequencing data from The Cancer Genome Atlas (TCGA), including lung adenocarcinoma (LUAD), lung squamous cell carcinoma (LUSC), and corresponding normal lung tissue samples downloaded through the GDC Data Portal (https://portal.gdc.cancer.gov/). The downloaded dataset included RNA sequencing expression data together with available clinical variables (patient identifier, sample identifier, age, sex, pathological stage, survival status, and overall survival time) for eligible TCGA-LUAD and TCGA-LUSC cases. Samples lacking gene expression or survival information were excluded from downstream survival and ROC analyses. The dataset used in this study is provided as Supplementary File 1. No clinical specimens collected at Sanya Central Hospital were used for the bioinformatics analyses. Detailed procedures for the GEPIA, Kaplan–Meier Plotter, and survivalROC analyses are provided in the Bioinformatics analysis section below. The complete bioinformatics workflow, including data acquisition, preprocessing, gene expression analysis, survival analysis, and ROC analysis, is summarized in Figure 2.
Inclusion and exclusion criteria
The inclusion and exclusion criteria of the patients are presented in Table 1. A power analysis using statistical power analysis software determined that at least 26 cases per group would provide 80% power to detect a moderate effect size (d = 0.8, α = 0.05, two-tailed). While initial enrollment aimed for 50 pairs in the lung cancer group, the final qRT-PCR analysis included 30 tumor samples and 19 normal adjacent tissue samples after exclusions for tissue or RNA quality. This reduction in sample size reflects real-world clinical constraints and highlights the importance of RNA integrity and tissue availability in translational studies. At n = 19 for comparisons, the minimum detectable effect size is d = 1.0 (80% power, α = 0.05). Thus, the experiment was sufficiently powered to detect large, but not small-to-moderate, differences in YTHDC2 expression.
Tissue specimens and quantitative real-time PCR (qRT-PCR)
The pathological diagnoses were determined in a double-blind manner by two different pathologists. Tumor purity was estimated by pathologists (>70% malignant cells) and confirmed via ESTIMATE (TCGA). Adjacent normal tissues were macrodissected to minimize stromal contamination. The histological types of lung cancer included lung adenocarcinoma and squamous cell carcinoma, with 26 cases of lung adenocarcinoma and 4 cases of squamous cell carcinoma. The clinical staging of the 50 lung cancer patients was conducted according to the staging criteria outlined in “Chinese Medical Association Clinical Guidelines for Lung Cancer (2024 Edition)”. Pathologically confirmed NSCLC specimens (n = 50) represented typical clinical distributions (see Table 2). Of the 50 patients initially enrolled, 30 tumor tissue samples and 19 matched adjacent normal tissue samples met the RNA quality criteria. Because paired statistical analysis requires matched specimens from the same patient, the comparison of tumor versus adjacent normal tissue expression was performed using the 19 available matched pairs. These clinical specimens were used exclusively for experimental validation by qRT-PCR and were analyzed independently of the public TCGA datasets used for bioinformatics analysis. These selected specimens were processed immediately after pathological confirmation and handled under RNase-free conditions before RNA extraction. This subset reflects the cases in which both adequate tissue quantity and high-quality total RNA (RNA integrity number, RIN >7.0) could be obtained. Total RNA concentration and purity were measured before reverse transcription, and only samples with adequate RNA quality (RIN >7.0) were included for downstream analysis. Equal amounts of total RNA were reverse-transcribed into complementary DNA (cDNA) according to the manufacturer's protocol prior to quantitative PCR. This process ensured the validity of the gene expression data analyzed. The qRT-PCR experiments were carried out on a real-time PCR thermocycler using a probe-based quantitative PCR assay. All reactions were performed in triplicate, together with no-template controls to ensure analytical reproducibility. PCR amplification was performed under the following cycling conditions: an initial enzyme activation/denaturation step at 95°C for 10 min, followed by 40 cycles of denaturation at 95°C for 15 s and annealing/extension at 60°C for 60 s. Fluorescence signals were acquired at the end of each amplification cycle. All reagents and consumables were obtained from commercial suppliers (see Table of Materials). The primer and probe sequences used in qRT-PCR are provided in Table 3.
Bioinformatics analysis
GEPIA database analysis of YTHDC2 gene expression
The GEPIA database was used to analyze YTHDC2 expression in NSCLC. The GEPIA web server (http://gepia.cancer-pku.cn/) was accessed through a web browser. The Expression DIY module was selected, the gene symbol "YTHDC2" was entered, the LUAD and LUSC datasets were selected, the default normalization parameters were retained, and differential expression box plots were generated directly through the GEPIA interface. Statistical significance was defined as p < 0.05.
Kaplan-Meier plotter database for survival analysis of lung cancer patients
The study analyzed the relationship between YTHDC2 expression and prognosis in lung cancer patients using the Kaplan-Meier Plotter database. The lung cancer dataset was selected, the gene symbol "YTHDC2" was entered, the auto-selected best cutoff option was applied, and Kaplan–Meier curves for overall survival and post-progression survival were generated using the default analysis settings. Patients were automatically stratified into high- and low-expression groups using the optimal cutoff determined by the Kaplan–Meier Plotter platform, and hazard ratios with corresponding 95% confidence intervals were generated using the platform's default settings.
Running R packages in R software for ROC curve plotting
RNA sequencing expression data and the corresponding clinical metadata for eligible TCGA-LUAD and TCGA-LUSC cases were downloaded from the GDC Data Portal. The downloaded datasets were merged by patient identifier and imported into R for downstream analyses. The survivalROC package was used to generate time-dependent ROC curves at 1-, 3-, and 5-year prediction time points, and the corresponding area under the curve (AUC) values were calculated to evaluate the prognostic performance of YTHDC2 expression. Only TCGA-LUAD and TCGA-LUSC patients with available RNA-seq expression and survival information were included in the survival ROC analysis. The local clinical cohort was not used for survival prediction because long-term follow-up data were unavailable.
Tissue YTHDC2 expression by qRT-PCR
To analyze gene expression levels, qRT-PCR was performed. Equal volumes of cDNA were added to each reaction according to the manufacturer's recommended reaction conditions. Amplification was performed using a probe-based quantitative PCR assay, and fluorescence data were collected automatically at the end of each amplification cycle. Briefly, the RNA extraction process involved several steps, including sample preparation, deparaffinization, removal of residual liquid, proteinase K digestion, incubation, centrifugation, DNase treatment, DNase I addition, and ethanol precipitation. The sample was then bound to a silica-based RNA purification column and centrifuged at 8,000 × g for 30 s. The column was then washed with wash buffer 1, wash buffer 2, and wash buffer 2 diluted with ethanol, and dried at 13,000 × g for 2 min. The RNA was then eluted by adding 70 µL of RNase-Free Water to the center of the column membrane, followed by centrifugation at 13,000 × g for 1 min. Each primer pair demonstrated a single amplification product, which was confirmed by melting-curve analysis before relative gene expression was calculated. The primer efficiency (90–110%) was validated using standard curves prior to sample analysis. Melting curve analyses confirmed the presence of single amplicons and the absence of primer dimers. Relative YTHDC2 expression was calculated using the 2-ΔCt method, in which Ct values were normalized to the endogenous reference gene GAPDH. Because expression values were presented as normalized expression levels rather than fold changes relative to a calibrator sample, the results are reported as 2−ΔCt values. GAPDH was selected as the housekeeping gene because its expression showed minimal variability (CV < 5%) compared with the tested alternatives (ACTB, CV = 12%; 18S rRNA, CV = 18%), consistent with reference gene selection criteria in m6A studies.
Statistical analysis
Statistical analyses were performed using statistical software, including the creation of graphics and charts. To analyze quantitative and categorical data on the YTHDC2 gene expression in NSCLC patients, R software was used for bioinformatics analyses and ROC curve generation. ROC curves and AUC were used to evaluate the diagnostic performance of YTHDC2 expression in predicting survival. For regression-based analyses, effect estimates (odds ratios) with corresponding 95% confidence intervals were reported where applicable. The Wilcoxon signed-rank test was used to compare YTHDC2 expression between paired tumor and adjacent normal tissue samples, while Pearson correlation analysis was used to assess the association between YTHDC2 expression and clinicopathological features. Correlation coefficients (r) and corresponding p-values were reported. Because the qRT-PCR data consisted of paired tumor and adjacent normal tissue samples from the same patients and the gene expression data were not normally distributed, the Wilcoxon signed-rank test was used to compare YTHDC2 expression levels between paired tissues. This test does not assume normality of the data and is commonly used for skewed biological data. All statistical tests were two-sided, and p <0.05 was considered statistically significant. Continuous variables were assessed for normality before analysis. Continuous variables are presented as mean ± standard deviation or median (interquartile range), as appropriate.