$$\rightleftharpoonup{xx}$$
$$\longleftharp{xx}$$,
$$\longrightharp{xx}$$,
This study was approved by the institutional ethics committee and conducted in accordance with the Declaration of Helsinki and local regulations governing the use of human biospecimens. This study was approved by the ethics committee of Boai Hospital of Zhongshan (KY-2020-012-124). Written informed consent was obtained from all participants before any study-specific procedure.
Study cohorts and samples
The discovery cohort comprised paired tumor-draining lymph node (TDLN) and metastatic lymph node (TMLN) tissues collected intraoperatively from 6 patients with HER2-positive breast cancer. For each patient, the matched TDLN was defined as a draining lymph node without histologic evidence of metastasis and served as the within-patient control comparator, whereas the TMLN was pathology-confirmed metastatic tissue. Fresh tissues were briefly rinsed in cold phosphate-buffered saline (PBS), blotted dry, snap-frozen in liquid nitrogen within 30 min of excision, and stored at -80 °C until extraction. Inclusion criteria were pathologically confirmed invasive breast carcinoma, HER2-positive status per ASCO/CAP, availability of paired nodes, and no neoadjuvant therapy. Exclusion criteria were inadequate tissue or RNA integrity number (RIN) < 7.0. One initially screened sample failed QC due to RIN < 7.0 and was excluded, so all discovery analyses consistently use n = 6. The independent validation cohort comprised 120 archival FFPE lymph-node specimens with clinical follow-up and was used for qRT-PCR and outcome analyses. For prognostic modeling, patients were compared by POSTN expression groups (High vs. Low, with a cut-point determined by X-tile), with the Low group serving as the reference comparator. Cases with missing follow-up or incomplete covariates were excluded from prognostic modelling.
Sample-size rationale (Discovery and validation)
The discovery stage used a paired TDLN-TMLN design to maximize within-patient contrast and reduce inter-individual variance while controlling multiplicity at FDR = 0.05. For the validation cohort (n = 120), POSTN expression was dichotomized at the cut-point determined by X-tile (High/Low = 40/80). Given 48 disease-free survival (DFS) events and α = 0.05, a Schoenfeld approximation indicates ≥ 80% power to detect clinically relevant hazard ratios of approximately HR ≥ 2.3 under equal allocation; this is consistent with the observed effect size (HR = 2.31, 95% CI 1.41-3.77).
Definition of HER2 positivity
HER2 positivity followed contemporary ASCO/CAP criteria13: immunohistochemistry (IHC) 3+ defined as uniform intense membrane staining in >10% of tumor cells, or in situ hybridization (ISH)-amplified defined as a HER2/CEP17 ratio ≥2.0 with an average HER2 copy number ≥4.0 signals per cell. IHC 2+ results underwent reflex ISH with blinded recounting of ≥20 invasive tumor cells to confirm amplification status.
Tissue processing and RNA isolation
All procedures were performed on ice unless specified, using RNase-free consumables. For each sample with ≤ 100 mg tissue, homogenization was performed in 1 mL of acid guanidinium thiocyanate-phenol-chloroform reagent (AGPC), followed by addition of 200 µL of chloroform with 15 s vigorous mixing and a 2-3 min room-temperature incubation. Phase separation was achieved by centrifugation at 12,000 x g for 15 min at 4 °C. The aqueous layer was then transferred to a fresh tube, and RNA was precipitated with 500 µL of isopropanol after a 10 min incubation at room temperature. Pelleting was completed at 12,000 x g for 10 min at 4 °C. The pellet was washed with 1 mL of 75% ethanol and centrifuged at 7,500 x g for 5 min at 4 °C, air-dried for 5-10 min, and dissolved in RNase-free water. An on-column DNase I digestion was used when genomic DNA carryover was suspected. Visual checkpoints included clear phase separation after chloroform extraction and an intact translucent pellet after isopropanol precipitation. Troubleshooting included repeating the ethanol wash for low A260/230 ratios and extending precipitation or ensuring cooling during pelleting for low yields.
RNA quality control
Quantification used spectrophotometry to monitor A260/280 and A260/230 with targets around 1.8-2.1, complemented by fluorometric measurement for accuracy. Integrity was assessed on a microfluidic electrophoresis system and required RIN ≥ 7.0. Distinct 18S/28S rRNA peaks and the absence of a genomic DNA smear were used as acceptance criteria; samples failing thresholds were re-extracted or excluded.
Library preparation
Stranded mRNA libraries were constructed with poly(A) selection for intact RNA, while rRNA depletion was permitted for partially degraded inputs. The typical input was ≥ 1 µg total RNA per library. Fragmentation was performed near 94 °C for 8 min; first-strand cDNA synthesis at 50 °C for 50 min; second-strand synthesis at 16 °C for 60 min; adapter ligation at 20 °C for 15 min; and PCR amplification used 10-12 cycles adjusted to avoid over-amplification. Cleanups employed bead ratios near 0.8x-1.0x, and the expected library size distribution was 300 bp, including adapters. Library quality was verified by microfluidic electrophoresis; adapter-dimer contamination prompted a stricter cleanup, and excessively broad size distributions were corrected by modestly shortening fragmentation.
Sequencing
Indexed libraries were sequenced in paired-end 150 bp (PE150) mode, targeting 30 million read pairs per library. Run-level quality required Q30 ≥ 90% and stable cluster densities with minimal lane bias. Libraries were randomized across lanes to mitigate batch effects, and run records documented lane assignments and any sentinel controls used to monitor cross-contamination.
Computational analysis and differential expression
Analyses were conducted in R (v4.3.2) on Linux. Raw read quality was assessed with FastQC (v0.11.9). Adapter and quality trimming were performed with fastp (v0.23.4) using automatic adapter detection, sliding-window trimming (window size 4 bp; mean Phred Q ≥ 20), a minimum read length of 50 bp, and poly-G trimming when relevant chemistries were detected. Reads were aligned to the GRCh37/hg19 reference genome using HISAT2 (v2.2.1) with library-appropriate strandness parameters and known splice-site indices. Gene-level counts were generated with featureCounts (Subread v2.0.3) against GENCODE v19 using paired-end counting, chimeric-read handling, multi-mapper-aware settings, and the correct strandness flag. Low-count genes were filtered by requiring counts ≥ 10 in at least three samples. Differential expression (DE) analysis was performed using DESeq2 (v1.40.2) under a paired design (design = ~ pair + condition) to compare TMLN versus TDLN, with default size-factor normalization and log2 fold-change shrinkage via apeglm. Potential outliers were assessed using Cook's distance. Transcriptome-wide multiplicity was controlled using the Benjamini-Hochberg false discovery rate (FDR). Significance was defined as FDR < 0.05 and an absolute log2 fold-change threshold (|log2FC|) ≥ 1. DESeq2's negative-binomial framework provides appropriate mean-variance modelling for count data and stabilizes fold-change estimates in small-to-moderate cohorts. As sensitivity analyses, we re-ran DE testing using edgeR and limma-voom with the same filtering criteria and FDR thresholds, which yielded concordant top-ranked signals, supporting the robustness of the main analysis choices.
Functional enrichment (GO and KEGG)
Functional annotation was performed in R using clusterProfiler (v4.8.3). Up- and down-regulated DEGs were analyzed separately. Gene identifiers were mapped to Entrez Gene IDs (organism: Homo sapiens) prior to enrichment. Gene Ontology (GO) enrichment was conducted using enrichGO (OrgDb: org.Hs.eg.db; ont = BP/CC/MF; pAdjustMethod = "BH"; pvalueCutoff = 0.05; qvalueCutoff = 0.05), and KEGG pathway enrichment was conducted using enrichKEGG (organism = "hsa"; pAdjustMethod = "BH"; pvalueCutoff = 0.05; qvalueCutoff = 0.05). The background universe was defined as all expressed genes retained after low-count filtering in the DE analysis. Enriched terms were additionally filtered to retain gene sets with 10-500 annotated genes (minGSSize = 10; maxGSSize = 500). Visualizations were generated using ggplot2 (v3.5.1) and ComplexHeatmap (v2.16.1), including dot plots of gene ratio and -log10(adjusted P) values; the term "immune-response pathways" refers to curated GO modules covering antigen processing and presentation, interferon signaling, and lymphocyte activation.
External database interrogation
Candidate genes were interrogated in external resources to provide orthogonal context. Tumor-versus-normal mRNA expression was queried using GEPIA2 (v2.0) with TCGA breast invasive carcinoma (TCGA-BRCA) tumor samples compared against normal breast tissues (GTEx/TCGA normals, as available). GEPIA2-reported expression values (log2[TPM+1]) and its default statistical testing framework for group comparisons were used to generate boxplots and P values for each queried gene. Protein-level localization was assessed using the Human Protein Atlas (HPA) by reviewing immunohistochemistry images and annotations from the Tissue Atlas and Pathology Atlas for breast tissue/cancer and recording the reported staining intensity and cellular compartment (e.g., stromal vs epithelial) when available.
qRT-PCR validation
Independent lymph-node tissues were processed as above to extract total RNA, and cDNA was synthesized from 1 µg RNA in 20 µL reverse-transcription reactions. Quantitative real-time PCR (qRT-PCR) was performed using SYBR-based chemistry, with each reaction containing 1x master mix and 0.2-0.4 µM of each primer. Primer specificity was confirmed by single-peak melt-curve profiles, and amplification efficiency was evaluated using dilution-series standard curves, with acceptable efficiencies defined as 90%-110%. Each sample was assayed in technical triplicate; replicates were required to meet a cycle-threshold (Ct) SD ≤ 0.3. Gene expression was normalized to GAPDH (or another validated reference gene, where applicable), and relative expression was calculated using the 2^-ΔΔCt method. Normality was assessed using Shapiro-Wilk tests; between-group comparisons used two-sided t-tests for approximately normal data or Mann-Whitney U tests otherwise.
Prognostic analysis in the validation cohort
Disease-free survival (DFS) was the primary endpoint and was calculated from the date of surgery to the first documented recurrence event or last follow-up (censored). For the primary analysis, POSTN expression was dichotomized at the cut-point determined by X-tile (High vs Low = 40/80); sensitivity analyses used alternative cut-points such as tertiles. Kaplan-Meier curves and log-rank tests were used for univariable comparisons. Multivariable Cox proportional-hazards models were fit to estimate hazard ratios (HRs) and 95% confidence intervals, adjusting a priori for age, tumor size, nodal status, histologic grade, and receptor status. The proportional-hazards assumption was assessed using Schoenfeld residuals, collinearity was examined using variance-inflation factors, and influential observations were evaluated using dfbetas. Complete case analyses were performed after excluding cases with missing follow-up or incomplete covariates, as described above.
General statistics and reporting
Unless stated otherwise, data are summarized as mean ± SD or median (IQR), two-sided tests are used, and p < 0.05 is considered statistically significant. For transcriptome-wide analyses and enrichment, multiplicity is controlled with BH-FDR, and effect sizes are accompanied by 95% confidence intervals whenever feasible.
Safety and waste disposal
Phenol-guanidinium reagents and organic solvents were handled in a certified fume hood with laboratory coats, nitrile gloves, and splash protection. Organic and halogenated waste were segregated into labeled containers and disposed of in accordance with institutional policies. Biological waste was autoclaved or chemically disinfected prior to disposal, and surfaces and tools were decontaminated with RNase-inactivating solutions. Reagent-specific hazards such as corrosivity and toxicity were documented in laboratory SOPs and observed during all procedures.