$$\rightleftharpoonup{xx}$$
$$\longleftharp{xx}$$,
$$\longrightharp{xx}$$,
Clinical Cohort and Sequencing Validation
The successful execution of the upstream RNA extraction and library preparation protocol (Figure 1) was confirmed by sequencing yield and quality metrics. In this representative dataset, bone marrow samples from five newly diagnosed AML patients and four relapsed AML patients yielded an average of approximately 6.0 GB of raw data per sample. Quality control assessment (Table 2) confirmed that base quality and read depth met the thresholds required for downstream bioinformatic analysis9. Low RNA integrity (for example, RIN < 6.0), low mapping rates, or high transcript degradation bias would represent suboptimal input quality and would compromise the reliability of downstream differential expression analysis.
Global Transcriptomic Variance and PCA
To assess global transcriptome variance and inspect clinical grouping, PCA was performed on the normalized expression data. In this representative dataset, the newly diagnosed and relapsed groups showed separation in two-dimensional space (Figure 2A)20, with PC1 and PC2 accounting for 23.82% and 18.75% of the total variance, respectively. The Venn diagrams in Figure 2B,C provide an additional descriptive summary of genes detected across samples within the newly diagnosed and relapsed groups, supporting sample-level reproducibility checks before downstream differential expression analysis. Because the cohort was small and unpaired, PCA separation was interpreted as an illustrative workflow output rather than definitive evidence of disease-state-specific biology.
Differential Expression Gene (DEG) Analysis
Applying the established protocol thresholds (|log2FC| ≥ 1 and adjusted P-value ≤ 0.05) to the DESeq2 output identified 2,025 DEGs, comprising 772 upregulated and 1,253 downregulated genes in the relapsed group (Figure 3A). Candidate transcripts with high variation included FOXC1 (log2FC = 7.55, P = 4.92 x 10-5), HOXA11 (log2FC = 7.76), HOXA11-AS (log2FC = 7.23), and AXL (log2FC = 3.50), alongside downregulated RHOB, PTX3, and CXCL8. Existing literature links several of these genes to AML stemness, signaling, or therapy response13,29; however, the present workflow identifies them only as relapse-associated candidate transcripts. Any definitive mechanistic role in clinical resistance requires subsequent independent functional validation.
Functional and Pathway Enrichment (GO, KEGG, and GSEA)
The functional annotation protocol mapped the DEGs to broader biological systems. GO analysis identified enrichment of terms related to small GTPase-mediated signal transduction, metal ion transport, and chromatin assembly (Figure 4A–C). KEGG pathway mapping identified associations with ECM-receptor interactions and cytokine-cytokine receptor interactions (Figure 4D). GSEA showed enrichment of RNA biosynthetic processes in the relapsed group and enrichment of energy metabolism pathways in the newly diagnosed group (Figure 5A). These enrichment results provide a descriptive roadmap of altered gene sets and should be interpreted as hypothesis-generating associations rather than proven drivers of relapse.
Protein-Protein Interaction (PPI) Network Construction
The initial STRING network contained 56 nodes and 193 interactions. After removal of disconnected or orphan nodes, the displayed Cytoscape subnetwork contained 42 nodes and 136 interactions (Figure 5B). Network modular analysis prioritized TP53, CCL2, CXCL8, and IL6 as central mathematical hubs with the highest number of interactions. Because the PPI network relies on database-predicted interaction scores (e.g., ATF3 score: 0.982), hub identification should be interpreted as target prioritization for future empirical studies rather than direct evidence of p53-mediated apoptosis evasion or other resistance mechanisms.
The raw RNA sequencing data generated in this protocol have been deposited in the Figshare repository and are publicly accessible via the following DOI: https://doi.org/10.6084/m9.figshare.30655814. The processed data and associated analysis files are included in the article and its supplementary materials. Representative command-line parameters and analysis settings used to reproduce the computational workflow are provided as Supplementary File 1. All data supporting the findings of this study are available without restriction.
| Patient ID | Age (Years) | Sex | Molecular Mutations | Survival/Follow-up (Months) | Clinical Status |
| R_AML_1 | 70 | Male | FLT3-ITD (+) | 22 | Deceased |
| R_AML_2 | 29 | Female | NPM1 (+) | 11 | Alive |
| R_AML_3 | 40 | Male | CEBPA (+) | 17 | Alive |
| R_AML_4 | 55 | Female | Triple Negative* | 24 | Deceased |
Table 1: Demographic and Clinical Characteristics of Patients in the Relapsed AML (R_AML) Group. Table 1 summarizes the demographic and clinical features of the relapsed AML cohort used in the representative analysis, including patient-level clinical characteristics relevant to interpretation of the transcriptomic workflow.
| Sample | Library | Raw_reads | Raw_bases | Clean_reads | Clean_bases | Error_rate | Q20 | Q30 | GC_pct |
| AML_1 | FRAS25
0244891-1r | 48705066 | 7.31G | 47807532 | 7.17G | 0.01 | 99.35 | 97.48 | 47.48 |
| AML_2 | FRAS25
0244896-1r | 42969940 | 6.45G | 42237962 | 6.34G | 0.01 | 99.35 | 97.44 | 46.74 |
| AML_3 | FRAS2502
44906-1r | 48738386 | 7.31G | 47744462 | 7.16G | 0.01 | 99.36 | 97.48 | 47.28 |
| AML_4 | FRAS250
244915-1r | 48723650 | 7.31G | 47688240 | 7.15G | 0.01 | 99.29 | 97.26 | 47.45 |
| AML_5 | FRAS2502
44920-1r | 49508198 | 7.43G | 47740308 | 7.16G | 0.01 | 99.37 | 97.53 | 47.73 |
| R_AML_1 | FRAS2502
44892-1r | 47879408 | 7.18G | 46671584 | 7.0G | 0.01 | 99.39 | 97.49 | 47.63 |
| R_AML_2 | FRAS2502
70005-1r | 47657378 | 7.15G | 46957882 | 7.04G | 0.01 | 99.39 | 97.49 | 50.5 |
| R_AML_3 | FRAS250
405722-1r | 58754766 | 8.81G | 56867112 | 8.53G | 0.01 | 99.38 | 97.42 | 46.52 |
| R_AML_4 | FRAS2502
44902-1r | 48491122 | 7.27G | 47469334 | 7.12G | 0.01 | 99.23 | 97.21 | 46.43 |
Table 2: Summary of data quality. Table 2 reports sequencing quality metrics for each sample, including read yield, base quality, GC content, and mapping-related quality-control information used to determine whether samples were suitable for downstream analysis.

Figure 1: Workflow of the protocol. The workflow summarizes the major experimental and computational stages, including clinical sample collection, RNA quality control, library preparation and sequencing, read processing and alignment, transcript quantification, differential expression analysis, GO/KEGG enrichment, GSEA, and PPI network construction. Please click here to view a larger version of this figure.

Figure 2: Quantitative analysis of samples. (A) Principal component analysis (PCA) was performed to evaluate intergroup differences and within-group sample reproducibility. PCA was conducted using linear algebraic methods based on normalized gene expression values across all samples. (B, C) Venn diagrams showing genes detected across samples in the AML and R_AML groups, respectively. Sample-restricted regions indicate genes detected in individual samples, while overlapping areas represent genes commonly detected across two or more samples. Please click here to view a larger version of this figure.

Figure 3: Differential gene expression analysis. (A) Bar plot showing the number of differentially expressed genes (DEGs) between comparison groups, identified by DESeq2 with thresholds of adjusted P-value ≤ 0.05 and |log2FoldChange| ≥ 1. (B) Volcano plot of DEGs. The x-axis represents log2FoldChange values, and the y-axis represents -log10(P-value). Blue dashed lines indicate the threshold lines used for DEG selection. (C) Hierarchical clustering heatmap of DEGs. The x-axis denotes sample names, and the y-axis shows normalized expression values of the DEGs. Please click here to view a larger version of this figure.

Figure 4: Functional enrichment analysis of differentially expressed genes. (A) GO enrichment bar plot. The x-axis represents GO terms, and the y-axis shows enrichment significance, expressed as -log10(padj). Colors represent BP (Biological Process), CC (Cellular Component), and MF (Molecular Function). (B) GO enrichment bubble plot. The x-axis represents the ratio of DEGs annotated to each GO term relative to the total number of DEGs, and the y-axis indicates the GO terms. Bubble size corresponds to the number of annotated genes, and color gradients represent enrichment significance. (C) KEGG enrichment bar plot. The x-axis represents KEGG pathways, and the y-axis denotes enrichment significance. (D) KEGG enrichment bubble plot. Bubble size indicates the number of annotated genes, and color gradients reflect enrichment significance. Please click here to view a larger version of this figure.

Figure 5: GSEA enrichment and protein–protein interaction (PPI) network analysis. (A) Bar plot showing normalized enrichment scores (NES) for selected significant gene sets. Positive NES values indicate enrichment in the R_AML group, whereas negative NES values indicate enrichment in the newly diagnosed AML group. (B) Protein–protein interaction (PPI) network. Each node represents a protein, and each edge denotes an interaction between connected proteins. Please click here to view a larger version of this figure.