Public transcriptomic data preprocessing and differentially expressed gene identification
Integration and ComBat correction of GSE22356, GSE33463, and GSE48149 reduced systematic differences among the datasets. Boxplots showed that sample expression distributions became more consistent after correction. Principal component analysis indicated that samples clustered mainly according to dataset source before correction but became more intermixed after correction, indicating effective batch-effect reduction (Figure 1).
Using thresholds of an absolute log2 fold change >0.585 and an adjusted P value < 0.05, 78 differentially expressed genes were identified, including 44 upregulated and 34 downregulated genes (Figure 2A). The heatmap showed that interferon-related genes, including XAF1, MX1, IFI44L, EPSTI1, PARP9, IFIH1, CXCL10, GBP1, STAT1, SAMHD1, TNFSF10, and TLR7, were generally upregulated in PH-related samples. In contrast, erythroid-related genes, including HBG1, HBD, ALAS2, CA1, and SLC4A1, tended to be downregulated (Figure 2B).

Figure 1: Batch-effect correction. (A) Boxplots of the merged expression matrix before and after ComBat correction. (B) Principal component analysis plots showing sample distributions before and after batch-effect correction. Please click here to view a larger version of this figure.

Figure 2: Differentially expressed genes. (A) Volcano plot showing upregulated and downregulated genes in PH-related samples. (B) Heatmap showing differentially expressed genes between the control and disease groups. Please click here to view a larger version of this figure.
Weighted gene co-expression network construction and functional enrichment
Sample clustering showed stable overall clustering with no obvious outliers (Figure 3A). The scale-free topology fit index approached 0.8 at a power of 9, and β = 9 was selected for network construction (Figure 3B). Gene clustering and dynamic module identification yielded multiple co-expression modules (Figure 3C). The blue module showed the strongest association with PH status (r = 0.55, P = 1 × 10−19), whereas the turquoise and grey modules also showed correlations with PH (Figure 3D).
Consensus genes obtained by intersecting the key weighted gene co-expression network analysis module genes with the differentially expressed genes were enriched in antiviral immune defense, NF-κB and JAK-STAT regulation, inflammatory-factor responses, cytokine and chemokine receptor binding, and transcriptional regulation (Figure 4A). Kyoto Encyclopedia of Genes and Genomes enrichment analysis identified cytokine-cytokine receptor interaction, chemokine signaling, NOD-like receptor signaling, Toll-like receptor signaling, tumor necrosis factor signaling, and interleukin-17 signaling (Figure 4B), supporting immune-inflammatory dysregulation as a molecular basis of pulmonary vascular remodeling.

Figure 3: Weighted gene co-expression network analysis. (A) Sample clustering tree and trait heatmap. (B) Soft-threshold selection plot. (C) Gene dendrogram and module colors. (D) Module-trait relationship heatmap. Please click here to view a larger version of this figure.

Figure 4: Functional enrichment analysis. (A) Gene Ontology enrichment results for the consensus genes. (B) Kyoto Encyclopedia of Genes and Genomes pathway enrichment results for the consensus genes. Please click here to view a larger version of this figure.
Protein-protein interaction network and hub-gene screening
The STRING protein-protein interaction network constructed from the consensus genes revealed an interconnected immune-inflammatory network (Figure 5A). Degree-based CytoHubba ranking showed that FN1, CD44, JUN, TGFB1, CXCL8, and BCL2 had high connectivity (Figure 5B). These hub genes may participate in inflammatory signaling, cell adhesion, extracellular matrix remodeling, and pulmonary vascular structural remodeling.

Figure 5: Protein-protein interaction network and hub genes. (A) STRING protein-protein interaction network of the consensus genes. (B) Degree-ranked hub genes identified using CytoHubba. Please click here to view a larger version of this figure.
Machine-learning feature selection and diagnostic performance
Least absolute shrinkage and selection operator regression identified six candidate genes with non-zero coefficients after cross-validation (Figure 6A, B). Support vector machine-recursive feature elimination retained eight genes and yielded a cross-validation accuracy of 0.883 and an error of 0.117 (Figure 7A). The random forest out-of-bag error stabilized when the number of trees was ≥100, and IFIH1, JUN, and TLR7 ranked among the top genes according to the importance score (Figure 7B).
The intersection of the least absolute shrinkage and selection operator, support vector machine-recursive feature elimination, and random forest results identified five core genes: CXCL10, JUN, IFIH1, MX1, and TLR7 (Figure 8A). Single-gene receiver operating characteristic analysis showed moderate-to-good diagnostic discrimination, with area under the curve values of 0.842 for IFIH1, 0.833 for JUN, 0.827 for TLR7, 0.814 for CXCL10, and 0.759 for MX1 (Figure 8B).

Figure 6: Least absolute shrinkage and selection operator regression analysis. (A) Coefficient path generated by the least absolute shrinkage and selection operator regression. (B) Cross-validation error plot. Please click here to view a larger version of this figure.

Figure 7: Support vector machine-recursive feature elimination and random forest analyses. (A) Support vector machine-recursive feature elimination feature-selection plot. (B) Random forest model and gene-importance ranking. Please click here to view a larger version of this figure.

Figure 8: Machine-learning summary. (A) Venn diagram showing the intersection of the three feature-selection algorithms. (B) Receiver operating characteristic curves for the five core genes. Please click here to view a larger version of this figure.
External validation in GSE117261
GSE117261 was used as an independent lung-tissue validation cohort and was not included in differential-expression screening, weighted gene co-expression network construction, or machine-learning feature selection (Table 3). The completed validation results showed heterogeneous replication across the five genes (Table 4). CXCL10 was increased (log₂ fold change = 0.677; P = 0.0410; FDR = 0.144) and yielded an AUC of 0.639 (95% CI, 0.491–0.786), with a cutoff of 6.245, sensitivity of 0.724, and specificity of 0.600. JUN was increased (log₂ fold change = 0.463; P = 0.00248; FDR = 0.0194) and yielded an AUC of 0.714 (95% CI, 0.593–0.835), with a cutoff of 8.708, sensitivity of 0.707, and specificity of 0.720.
IFIH1 (log2 fold change = 0.107; P = 0.371; FDR = 0.591; AUC = 0.543, 95% CI, 0.398–0.687), MX1 (log2 fold change = 0.109; P = 0.488; FDR = 0.690; AUC = 0.475, 95% CI, 0.337–0.614), and TLR7 (log2 fold change = -0.050; P = 0.543; FDR = 0.733; AUC = 0.546, 95% CI, 0.414–0.679) did not meet the prespecified external-support criteria. TLR7 also showed a direction opposite to the qRT-PCR result. The exploratory five-gene model fitted and evaluated within GSE117261 yielded an apparent AUC of 0.740, whereas repeated nested cross-validation yielded an AUC of 0.656. Thus, JUN received the strongest independent support, CXCL10 showed limited directionally consistent evidence, and replication of IFIH1, MX1, and TLR7 was weak or discordant.
| Item | Description |
| Dataset accession | GSE117261 |
| Data source | Gene Expression Omnibus (GEO) |
| Sample type | Human lung-tissue transcriptomic microarray data |
| Sample size | 58 PAH samples and 25 failed-donor control samples |
| Platform | GPL6244 / Affymetrix Human Gene 1.0 ST Array |
| Validation objective | Expression differences, single-gene ROC analysis, and exploratory five-gene combined ROC modeling for CXCL10, JUN, IFIH1, MX1, and TLR7 |
| Role in this study | Independent external-validation dataset; not included in the original training, WGCNA, or feature-selection analyses |
Table 3: Basic information for the GSE117261 independent external-validation dataset. The table summarizes the dataset source, sample type, sample size, platform, validation objectives, and role of GSE117261 in the study.
| Gene/model | PAH n | Control n | log2 fold change | P value | FDR | AUC | AUC 95% CI | Youden cutoff | Sensitivity | Specificity |
| CXCL10 | 58 | 25 | 0.677 | 0.041 | 0.144 | 0.639 | 0.491–0.786 | 6.245 | 0.724 | 0.6 |
| JUN | 58 | 25 | 0.463 | 0.00248 | 0.019 | 0.714 | 0.593–0.835 | 8.708 | 0.707 | 0.72 |
| IFIH1 | 58 | 25 | 0.107 | 0.371 | 0.591 | 0.543 | 0.398–0.687 | 5.3 | 0.759 | 0.44 |
| MX1 | 58 | 25 | 0.109 | 0.488 | 0.69 | 0.475 | 0.337–0.614 | 6.833 | 0.724 | 0.4 |
| TLR7 | 58 | 25 | -0.05 | 0.543 | 0.733 | 0.546 | 0.414–0.679 | 4.065 | 0.138 | 1 |
| Five-gene model (apparent/in-sample) | 58 | 25 | / | / | / | 0.74 | 0.621–0.859 | 0.613 | 0.845 | 0.6 |
| Five-gene model (repeated nested CV) | 58 | 25 | / | / | / | 0.656 | 0.517–0.795 | 0.655 | 0.845 | 0.52 |
Table 4: Actual expression-difference and ROC validation results for GSE117261. The table reports the PAH and control sample sizes, log₂ fold changes, P values, false-discovery-rate-adjusted values, AUCs, 95% confidence intervals, Youden-index cutoffs, sensitivities, specificities, and interpretations for the five genes and exploratory combined models.
Quantitative reverse transcription PCR validation
Quantitative reverse transcription PCR validation included 20 biologically independent PAH lung-tissue samples and 20 biologically independent control samples, with three technical replicates averaged for each biological sample. CXCL10, JUN, IFIH1, MX1, and TLR7 were significantly upregulated in PAH (Table 5; Figure 9A). The mean relative-expression values were approximately 3.470 for CXCL10, 2.560 for JUN, 2.760 for IFIH1, 2.650 for MX1, and 2.580 for TLR7. The corresponding P values/FDR values were 1.43 × 10⁻7/7.15 × 10⁻7, 4.17 × 10⁻5/4.17 × 10⁻5, 1.10 × 10⁻5/1.38 × 10⁻5, 1.58 × 10⁻6/2.63 × 10⁻6, and 1.37 × 10⁻6/2.63 × 10⁻6, respectively.
Single-gene ROC analysis based on qRT-PCR expression showed AUCs of 0.988 for CXCL10 (95% CI, 0.961–1.000), 0.880 for JUN (95% CI, 0.758–1.000), 0.908 for IFIH1 (95% CI, 0.820–0.995), 0.945 for MX1 (95% CI, 0.874–1.000), and 0.948 for TLR7 (95% CI, 0.886–1.000) (Figure 9B; Table 5). The five-gene logistic regression model achieved an apparent AUC of 1.000 (DeLong 95% CI: 1.000–1.000), with sensitivity and specificity of 1.000 (Figure 9C). Expression-direction consistency between quantitative reverse transcription PCR validation and GSE117261 was visualized using a heatmap (Figure 9D). Because the same 40 samples were used for both fitting and evaluation, this reflected the apparent in-sample performance. In 100 repeats of stratified five-fold cross-validation using an L2-regularized logistic-regression model, the pooled out-of-fold AUC also remained 1.000 (95% CI, 1.000–1.000), and every repeat yielded an AUC of 1.000. Despite this internal stability, the cohort was small and independent prospective validation remains necessary. Directional comparison with GSE117261 showed concordant increases for CXCL10, JUN, IFIH1, and MX1, but discordant direction for TLR7 (Figure 9D).

Figure 9: Quantitative reverse transcription PCR validation. (A) Boxplots showing the relative expression of CXCL10, JUN, IFIH1, MX1, and TLR7 in 20 biologically independent PAH lung-tissue samples and 20 biologically independent control samples. Each biological sample was measured in three technical replicates, and the mean Ct value was used for analysis. For each boxplot, the center line represents the median, the box represents the interquartile range, the whiskers extend to 1.5 times the interquartile range, and individual points beyond the whiskers represent outliers. (B) Single-gene ROC curves based on qRT-PCR expression values; AUCs and DeLong 95% confidence intervals are shown. (C) ROC curves for the five-gene logistic-regression model, showing both apparent/in-sample performance and pooled out-of-fold performance from 100 repeated stratified five-fold cross-validation analyses. (D) Heatmap showing expression-direction consistency between qRT-PCR and GSE117261; CXCL10, JUN, IFIH1, and MX1 were concordantly increased, whereas TLR7 was discordant. Please click here to view a larger version of this figure.
| Gene/model | PAH n | Control n | Control 2^-ΔΔCt, mean ± SD | PAH 2^-ΔΔCt, mean ± SD | Direction | P value | FDR | AUC | AUC 95% CI | Youden cutoff | Sensi-
tivity | Specifi-
city | Validation type |
| CXCL10 | 20 | 20 | 1.099 ± 0.502 | 3.470 ± 1.043 | Upregulated | 1.43E-07 | 7.15E-07 | 0.987 | 0.961–1.000 | 2.028 | 1 | 0.95 | Single-gene qRT-PCR analysis |
| JUN | 20 | 20 | 1.158 ± 0.738 | 2.560 ± 1.609 | Upregulated | 4.17E-05 | 4.17E-05 | 0.88 | 0.758–1.000 | 1.423 | 0.85 | 0.9 | Single-gene qRT-PCR analysis |
| IFIH1 | 20 | 20 | 1.091 ± 0.440 | 2.760 ± 1.398 | Upregulated | 1.10E-05 | 1.38E-05 | 0.908 | 0.820–0.995 | 1.579 | 0.8 | 0.85 | Single-gene qRT-PCR analysis |
| MX1 | 20 | 20 | 1.132 ± 0.582 | 2.650 ± 1.085 | Upregulated | 1.58E-06 | 2.63E-06 | 0.945 | 0.874–1.000 | 1.779 | 0.9 | 0.9 | Single-gene qRT-PCR analysis |
| TLR7 | 20 | 20 | 1.080 ± 0.444 | 2.580 ± 1.284 | Upregulated | 1.37E-06 | 2.63E-06 | 0.947 | 0.886–1.000 | 1.764 | 0.85 | 0.9 | Single-gene qRT-PCR analysis |
| Five-gene model (apparent/in-sample) | 20 | 20 | Not appli-
cable | Not appli-
cable | Not applicable | / | / | 1 | 1.000–1.000 | 0.998 | 1 | 1 | Same 40 biological samples used for model fitting and evaluation |
| Five-gene model (100× repeated 5-fold CV) | 20 | 20 | Not appli-
cable | Not appli-
cable | Not applicable | / | / | 1 | 1.000–1.000 | 0.716 | 1 | 1 | Internal cross-validation using L2-regularized logistic regression |
Table 5: Complete qRT-PCR expression and ROC results for CXCL10, JUN, IFIH1, MX1, and TLR7, including the five-gene combined-model analyses. The table reports the PAH and control sample sizes, relative-expression values, expression directions, P values, false-discovery-rate-adjusted values, AUCs, 95% confidence intervals, Youden-index cutoffs, sensitivities, specificities, and validation types for the individual genes and combined models.
Single-cell transcriptomic validation in GSE210248
GSE210248 provided cell-level mechanistic support by showing that PAH pulmonary arterial remodeling was accompanied by altered communication between immune cells and vascular structural cells. This observation was consistent with the bulk transcriptomic enrichment of inflammatory responses, chemokine signaling, Toll-like receptor signaling, and tumor necrosis factor signaling.
Single-cell evidence suggested that the PAH pulmonary artery signaling network shifted toward structural cells, including smooth muscle cells and fibroblasts. Smooth muscle cells exhibited multiple states, including oxygen-sensing/pericyte-like, contractile, synthetic, and fibroblast-like. These results support a disease model in which immune-inflammatory activation and vascular structural-cell remodeling jointly drive PH/PAH progression.
Candidate compound screening and molecular docking
Connectivity Map screening identified BRD-K91900765 as the highest-ranked candidate compound among the top 10 hits, with a Logit score of 10.13 and a predicted probability of 0.085 (Figure 10). Compound curation showed that BRD-K91900765 corresponds to VX-745/neflamapimod, a selective p38α/MAPK14 inhibitor with a PubChem compound identifier of 3038525 and a molecular mass of 436.27 g/mol (Table 6).
Exploratory docking of BRD-K91900765/VX-745 with the five biomarker-associated proteins yielded top Vina scores of -7.5 kcal/mol for CXCL10, -7.6 kcal/mol for JUN, -7.5 kcal/mol for IFIH1, -8.7 kcal/mol for MX1, and -8.1 kcal/mol for TLR7 (Tables 7–11; Figure 11A–E). These results indicated predicted structural compatibility only and did not establish the five proteins as direct pharmacological targets. Preliminary absorption, distribution, metabolism, excretion, and toxicity predictions suggested that the compound had several drug-like properties, although the relatively high calculated cLogP requires further evaluation (Table 12). Docking against the established VX-745 target, MAPK14/p38α (PDB ID: 1OUK), was included as a positive reference analysis. The top MAPK14 cavity, C1, yielded a Vina score of -7.9 kcal/mol, a cavity volume of 3560 Å3, a docking-box center of (2, 22, 34), and dimensions of (22, 31, 31) (Table 13; Figure 11F).

Figure 10: Connectivity Map candidate compound ranking. Ranking of the candidate compounds identified through Connectivity Map screening. BRD-K91900765 was the highest-ranked compound, with a Logit score of 10.13 and a predicted probability of 0.085. Please click here to view a larger version of this figure.

Figure 11: Three-dimensional molecular docking diagrams for BRD-K91900765/VX-745. (A) Exploratory docking with CXCL10. (B) Exploratory docking with JUN. (C) Exploratory docking with IFIH1. (D) Exploratory docking with MX1. (E) Exploratory docking with TLR7. (F) Positive-reference docking with the established pharmacological target MAPK14/p38α (PDB ID: 1OUK). Panels A–E indicate predicted structural compatibility and do not establish direct pharmacological targeting. Please click here to view a larger version of this figure.
| Item | Description |
| CMap/Broad ID | BRD-K91900765 (common batch format: BRD-K91900765-001-xx-x) |
| Common name/aliases | VX-745; neflamapimod; VRT-031745; VD-31745 |
| Chemical name | 5-(2,6-dichlorophenyl)-2-(2,4-difluorophenyl)sulfanylpyrimido[1,6-b]pyridazin-6-one |
| PubChem CID | 3038525 |
| CAS number | 209410-46-8 |
| Molecular formula / relative molecular mass | C19H9Cl2F2N3OS; 436.27 g/mol |
| Canonical SMILES | C1=CC(=C(C(=C1)Cl)C2=C3C=CC(=NN3C=NC2=O)SC4=C(C=C(C=C4)F)F)Cl |
| InChIKey | VEPKQEUBKLEPRA-UHFFFAOYSA-N |
| Established major pharmacological target | MAPK14/p38α; p38β inhibition has also been reported with lower selectivity than p38α |
Table 6: Chemical and pharmacological information for BRD-K91900765/VX-745. The table summarizes the compound identifiers, aliases, chemical name, molecular formula, molecular mass, structural descriptors, and established pharmacological target of BRD-K91900765/VX-745.
| CurPocket ID | Vina score (kcal/mol) | Cavity volume (ų) | Center (x, y, z) | Docking size (x, y, z) |
| C1 | -7.5 | 7568 | 49, 15, 4 | 34, 33, 35 |
| C4 | -7 | 160 | 46, 0, 4 | 22, 22, 22 |
| C3 | -6 | 208 | 34, 15, 9 | 22, 22, 22 |
| C2 | -5.8 | 465 | 45, -6, 19 | 22, 22, 22 |
| C5 | -5.1 | 150 | 63, 0, 20 | 22, 22, 22 |
Table 7: Predicted docking pockets for BRD-K91900765/VX-745 with CXCL10 (PDB ID: 1LV9). The table reports the ranked cavity identifiers, Vina scores, cavity volumes, docking-box centers, and docking-box dimensions generated by CB-Dock2.
| CurPocket ID | Vina score (kcal/mol) | Cavity volume (ų) | Center (x, y, z) | Docking size (x, y, z) |
| C3 | -7.6 | 1593 | 26, 35, 70 | 22, 22, 22 |
| C2 | -7.5 | 1927 | 25, 21, 60 | 35, 22, 22 |
| C1 | -7.4 | 5420 | 32, 37, 48 | 35, 22, 31 |
| C5 | -6 | 464 | 26, 5, 48 | 22, 22, 22 |
| C4 | -5.3 | 683 | 59, 32, 39 | 22, 22, 22 |
Table 8: Predicted docking pockets for BRD-K91900765/VX-745 with JUN (PDB ID: 1JUN). The table reports the ranked cavity identifiers, Vina scores, cavity volumes, docking-box centers, and docking-box dimensions generated by CB-Dock2.
| CurPocket ID | Vina score (kcal/mol) | Cavity volume (ų) | Center (x, y, z) | Docking size (x, y, z) |
| C1 | -7.5 | 583 | 15, 4, 17 | 22, 22, 22 |
| C3 | -7.2 | 146 | 19, 21, 24 | 22, 22, 22 |
| C4 | -6.6 | 141 | 31, 19, 15 | 22, 22, 22 |
| C2 | -6.2 | 273 | 38, 9, 24 | 22, 22, 22 |
| C5 | -6.2 | 125 | 27, -7, 10 | 22, 22, 22 |
Table 9: Predicted docking pockets for BRD-K91900765/VX-745 with IFIH1 (PDB ID: 3B6E). The table reports the ranked cavity identifiers, Vina scores, cavity volumes, docking-box centers, and docking-box dimensions generated by CB-Dock2.
| CurPocket ID | Vina score (kcal/mol) | Cavity volume (ų) | Center (x, y, z) | Docking size (x, y, z) |
| C1 | -8.7 | 523 | -18, -14, -5 | 22, 22, 22 |
| C4 | -7.3 | 262 | -13, -6, -9 | 22, 22, 22 |
| C3 | -7 | 306 | 12, 12, -7 | 22, 22, 22 |
| C2 | -6.9 | 408 | -3, 1, -12 | 22, 22, 22 |
| C5 | -6.2 | 179 | 24, 25, 13 | 22, 22, 22 |
Table 10: Predicted docking pockets for BRD-K91900765/VX-745 with MX1 (PDB ID: 5GTM). The table reports the ranked cavity identifiers, Vina scores, cavity volumes, docking-box centers, and docking-box dimensions generated by CB-Dock2.
| CurPocket ID | Vina score (kcal/mol) | Cavity volume (ų) | Center (x, y, z) | Docking size (x, y, z) |
| C2 | -8.1 | 7347 | 112, 137, 148 | 33, 28, 35 |
| C3 | -8.1 | 2604 | 124, 124, 176 | 30, 22, 22 |
| C1 | -8 | 7676 | 137, 112, 148 | 34, 29, 35 |
| C5 | -6.9 | 1976 | 117, 157, 88 | 22, 22, 22 |
| C4 | -6.6 | 2040 | 132, 92, 91 | 22, 22, 22 |
Table 11: Predicted docking pockets for BRD-K91900765/VX-745 with TLR7 (PDB ID: 7CYN). The table reports the ranked cavity identifiers, Vina scores, cavity volumes, docking-box centers, and docking-box dimensions generated by CB-Dock2.
| Category | Parameter | Result | Interpretation |
| Physicochemical property | Molecular weight | 436.27 g/mol | Below 500 Da, meeting the Lipinski molecular-weight threshold |
| Physicochemical property | cLogP | Approximately 5.49 | Slightly higher than 5, suggesting high lipophilicity and the need to consider solubility and nonspecific binding |
| Physicochemical property | TPSA | Approximately 47.26 Ų | Low polar surface area, consistent with potentially favorable membrane permeability |
| Drug-likeness | HBA/HBD | May-00 | Meets the Lipinski thresholds for hydrogen-bond acceptors and donors |
| Drug-likeness | Rotatable bonds | 3 | Low conformational flexibility, favorable for stable binding conformations |
| Structural alerts | PAINS/Brenk alerts | Not detected | No common pan-assay interference or reactive structural alerts detected |
| Toxicity prediction | Ames mutagenicity | Predicted non-Ames toxic | Suggests a low predicted mutagenic risk; experimental validation is still required |
| Toxicity prediction | Carcinogenicity | Predicted non-carcinogenic | Suggests a relatively low predicted long-term carcinogenic risk; experimental validation is still required |
| Pharmacokinetic note | Oral availability/brain penetrance | Literature and databases indicate an orally available, brain-penetrant small molecule | Consistent with its development background as a p38α inhibitor; reevaluation is still needed for PH indications |
Table 12: Preliminary physicochemical, drug-likeness, ADMET, and toxicity predictions for BRD-K91900765/VX-745. The table summarizes predicted physicochemical properties, drug-likeness measures, structural alerts, toxicity endpoints, and pharmacokinetic characteristics. These computational predictions are preliminary and do not replace experimental pharmacokinetic or toxicological validation.
| CurPocket ID | Vina score (kcal/mol) | Cavity volume (ų) | Center (x, y, z) | Docking size (x, y, z) |
| C1 | -7.9 | 3560 | 2, 22, 34 | 22, 31, 31 |
| C5 | -7.5 | 254 | -9, 28, 60 | 22, 22, 22 |
| C3 | -7.4 | 351 | 12, 7, 37 | 22, 22, 22 |
| C2 | -6.3 | 827 | 18, 5, 28 | 22, 22, 22 |
| C4 | -6.3 | 334 | -16, 17, 38 | 22, 22, 22 |
Table 13: Predicted docking pockets for BRD-K91900765/VX-745 with its established pharmacological target MAPK14/p38α (PDB ID: 1OUK), included as the positive-reference analysis. The table reports the ranked cavity identifiers, Vina scores, cavity volumes, docking-box centers, and docking-box dimensions generated using the same docking workflow applied to the five biomarker-associated proteins.
Collectively, the discovery analyses and qRT-PCR results support CXCL10, JUN, IFIH1, MX1, and TLR7 as candidate PH/PAH biomarkers associated with immune-inflammatory dysregulation and pulmonary vascular remodeling, although independent GSE117261 replication was strongest for JUN and variable for the other genes. BRD-K91900765/VX-745 is a computationally prioritized drug-repositioning candidate with a plausible mechanism of MAPK14/p38α inhibition; target-binding, cellular, pharmacokinetic, toxicity, and animal-model validation are required before therapeutic interpretation.
DATA AVAILABILITY:
All public transcriptomic datasets used in this study are available from the Gene Expression Omnibus database under accession numbers GSE22356, GSE33463, GSE48149, GSE117261, and GSE210248. All coding files, processed datasets, de-identified qRT-PCR raw and analyzed data, model outputs, and molecular docking input/output files have been consolidated into a structured Zenodo repository. The repository includes a README that describes each file, software and package versions, script execution order, and complete reproduction steps - https://zenodo.org/records/21682282