A subscription to JoVE is required to view this content. Sign in or start your free trial.

Research Article

Bioinformatics and Machine-Learning Identification of Pulmonary Hypertension Biomarkers and Candidate Therapeutic Compounds

82 views

⸱

DOI:

10.3791/73519

⸱

August 25th, 2026

In This Article

Summary

This article presents a reproducible bioinformatics workflow integrating public transcriptomic datasets, machine learning, external validation, quantitative reverse transcription PCR, Connectivity Map screening, and molecular docking to identify pulmonary hypertension biomarkers and candidate therapeutic compounds.

Abstract

This study aimed to identify pulmonary hypertension (PH)-associated molecular biomarkers and candidate small-molecule compounds using public transcriptomic data and independent validation resources. Three Gene Expression Omnibus datasets (GSE22356, GSE33463, and GSE48149) were integrated following normalization, probe annotation, and ComBat batch-effect correction. Differential expression analysis, weighted gene co-expression network analysis, functional enrichment analysis, protein-protein interaction network analysis, and three machine-learning algorithms were used to identify core feature genes. Diagnostic performance was evaluated using receiver operating characteristic curves. External validation included an independent lung-tissue cohort (GSE117261), a pulmonary artery single-cell RNA-sequencing dataset (GSE210248), and quantitative reverse transcription PCR validation in independent lung-tissue samples. Connectivity Map-based drug repositioning and molecular docking were used to screen candidate compounds. Seventy-eight differentially expressed genes were identified, and CXCL10, JUN, IFIH1, MX1, and TLR7 were selected as core feature genes. In the independent GSE117261 lung-tissue cohort, JUN showed the strongest external support, whereas replication of the other genes was variable. Quantitative reverse transcription PCR in 20 biologically independent pulmonary arterial hypertension samples and 20 control samples confirmed upregulation of all five genes. The apparent five-gene qRT-PCR model and 100 repeated stratified five-fold cross-validation analyses both yielded an area under the curve of 1.000, although the small cohort requires cautious interpretation and independent prospective validation. Single-cell analysis of GSE210248 supported altered communication between immune and structural cells and smooth muscle cell phenotypic switching. BRD-K91900765/VX-745 ranked highest in Connectivity Map screening. MAPK14/p38α, its established pharmacological target, was included as a positive-reference docking protein, whereas docking against the five biomarker-associated proteins was treated as exploratory. These findings support the five genes as candidate PH biomarkers and VX-745 as a computational drug-repositioning hypothesis requiring experimental validation.

Introduction

Pulmonary hypertension (PH) is a progressive cardiopulmonary syndrome characterized by persistently elevated pulmonary arterial pressure, increased pulmonary vascular resistance, and eventual right ventricular failure. Current hemodynamic criteria define PH as a mean pulmonary arterial pressure at rest of >20 mmHg, as measured by right heart catheterization1. Among the different clinical subtypes, pulmonary arterial hypertension (PAH) is one of the most severe forms and is characterized by progressive pulmonary vascular remodeling. Its pathological features include endothelial dysfunction, abnormal proliferation and migration of pulmonary arterial smooth muscle cells, adventitial fibroblast activation, extracellular matrix deposition, inflammatory cell infiltration, and narrowing or obliteration of the distal pulmonary arteries2. These changes indicate that PH/PAH is not only a disorder of vasoconstriction but also a complex vascular remodeling disease driven by coordinated molecular, cellular, and immune-inflammatory mechanisms.

Current PAH therapies mainly target the prostacyclin, endothelin, nitric oxide–soluble guanylate cyclase, and phosphodiesterase type 5 pathways3˒4. Although these treatments improve symptoms, exercise capacity, and hemodynamic parameters, their effects remain largely vasodilatory and hemodynamic. Their ability to reverse established pulmonary vascular remodeling is limited, and many patients continue to experience disease progression despite combination therapy. Therefore, identifying novel molecular biomarkers and therapeutic candidates that reflect the remodeling process represents an important unmet need. In particular, immune-inflammatory activation, interferon-related signaling, Toll-like receptor pathways, chemokine-mediated immune recruitment, and smooth muscle cell phenotypic switching have emerged as potential contributors to PH/PAH progression5˒6.

High-throughput transcriptomic datasets provide valuable resources for identifying disease-associated molecular signatures in PH/PAH. However, studies based on a single dataset are often limited by small sample sizes, batch effects, platform heterogeneity, and insufficient validation. Differential expression analysis can identify genes with altered expression but may not fully capture disease-related co-expression modules or network-level interactions. Weighted gene co-expression network analysis (WGCNA) can identify gene modules associated with disease traits, whereas protein-protein interaction (PPI) network analysis can reveal highly connected genes within biological networks. Machine-learning methods can also prioritize genes with diagnostic or classification value. However, reliance on a single algorithm may introduce model-specific bias. Integrating differential expression analysis, WGCNA, PPI network analysis, and multiple machine-learning algorithms may therefore improve the robustness of biomarker discovery.

Another major challenge in transcriptomic biomarker studies is biological interpretation. Bulk-tissue signals may reflect changes in gene expression within resident vascular cells, immune cell infiltration, or altered proportions of multiple cell populations. Single-cell RNA sequencing provides an opportunity to place bulk-derived candidate genes within a cellular context. In PH/PAH, pulmonary vascular remodeling involves endothelial cells, smooth muscle cells, fibroblasts, monocytes/macrophages, lymphocytes, and other immune or structural cells. Disease progression is also associated with altered cell-cell communication and smooth muscle cell phenotypic switching. Thus, combining bulk transcriptomic screening with single-cell validation may help determine whether candidate biomarkers are associated with immune activation, vascular structural remodeling, or an imbalance in multicellular communication.

In addition to biomarker discovery, transcriptomic signatures can be used for computational drug repositioning. The Connectivity Map (CMap) links disease-associated gene-expression profiles with small molecules that may reverse or modulate those signatures7. When combined with compound curation and molecular docking, this strategy can generate experimentally testable therapeutic hypotheses. Although CMap prediction and molecular docking cannot establish drug efficacy, they can prioritize candidate compounds for future target-binding assays, cell-based experiments, and animal-model validation.

An integrated and reproducible workflow was developed to identify PH/PAH biomarkers and candidate therapeutic compounds. Three public Gene Expression Omnibus transcriptomic datasets were integrated following normalization and batch-effect correction. Differential expression analysis, WGCNA, functional enrichment analysis, PPI network analysis, and three machine-learning algorithms were used to screen robust feature genes. Receiver operating characteristic analysis, an independent lung-tissue validation cohort, pulmonary artery single-cell RNA-sequencing evidence, and quantitative reverse transcription PCR validation in independent samples were used to further evaluate the selected genes. Finally, CMap-based drug repositioning and molecular docking were applied to identify candidate compounds. The novelty of the study lies in its multilayered validation framework, which connects bulk transcriptomic discovery, machine-learning prioritization, independent validation, experimental quantitative reverse transcription PCR confirmation, single-cell mechanistic interpretation, and computational compound screening. The study hypothesis was that PH/PAH is driven by a coordinated immune-inflammatory and vascular remodeling program and that robust genes within this program may serve as candidate biomarkers and provide drug-repositioning opportunities.

Access restricted. Please log in or start a trial to view this content.

Protocol

The public Gene Expression Omnibus (GEO) datasets analyzed in this study contained de-identified transcriptomic data from previously published studies and did not require additional ethical approval. The Ethics Committee of Huaihua University approved the independent human lung-tissue quantitative reverse transcription PCR (qRT-PCR) validation study (approval no. 2024(A05112)). Written informed consent was obtained from all participants or their legally authorized representatives before sample collection. The approval and consent procedures applied to all 20 pulmonary arterial hypertension (PAH) and 20 control lung-tissue samples included in the qRT-PCR validation. The research tools used for this protocol are listed in the Table of Materials.

1. Collection and preprocessing of public transcriptomic datasets

Pulmonary hypertension (PH)-related microarray datasets GSE22356, GSE33463, and GSE48149 were obtained from the GEO database. PH/pulmonary arterial hypertension (PAH) and control samples were extracted according to the original phenotype annotations. Expression matrices and platform annotation files were downloaded using reproducible R scripts and the GEOquery package.

Probe annotation and gene-symbol mapping were performed consistently across the datasets. When multiple probes mapped to the same gene, the mean expression value was calculated. Quantile normalization was applied, and genes with low expression or low variance were removed. The datasets were merged, and batch effects were corrected using the ComBat algorithm in the sva package8. The correction was evaluated using boxplots and principal component analysis.

2. Identification of differentially expressed genes

The limma package was used to compare expression levels between PH and control samples in the batch-corrected expression matrix9. A linear model was fitted, and empirical Bayes statistics were applied. Differentially expressed genes were defined using an adjusted P value <0.05 and an absolute log2 fold change > 0.585. The results were visualized using volcano plots and heatmaps.

3. Weighted gene co-expression network construction

A weighted gene co-expression network was constructed using the WGCNA package10. Sample clustering was performed to detect outliers. The soft-thresholding power was selected based on the scale-free topology fit index. Gene modules were identified using the dynamic tree-cutting algorithm. Module eigengenes were correlated with the PH phenotype, and the disease-associated module with the strongest correlation was selected. Genes in the key module were intersected with the differentially expressed genes to obtain consensus genes.

4. Functional enrichment analysis

Gene Ontology biological process, cellular component, and molecular function categories were analyzed using clusterProfiler11. Kyoto Encyclopedia of Genes and Genomes pathway enrichment analysis was performed to identify signaling pathways12. A P value < 0.05 and a q value < 0.2 were used as the enrichment thresholds, and enriched terms were visualized using bubble plots11.

5. Protein-protein interaction network construction and hub-gene identification

The consensus gene list was submitted to the STRING database, with Homo sapiens selected as the species and an interaction confidence threshold > 0.413. The interaction file was imported into Cytoscape, and the CytoHubba plug-in was used to rank genes by node degree. Highly connected genes were defined as hub genes.

6. Selection of diagnostic feature genes using machine learning

Three independent feature-selection algorithms were applied. First, least absolute shrinkage and selection operator logistic regression was performed using the glmnet package and 10-fold cross-validation to identify genes with non-zero coefficients14. Second, recursive feature elimination with support vector machines was applied to remove redundant features and select the feature subset that achieved the highest cross-validation accuracy15. Third, a random forest model was constructed, and the features were ranked by mean decrease in Gini impurity16. The intersection of the gene sets derived from the three algorithms was used to define the final set of core feature genes. The pROC package was used to generate receiver operating characteristic curves and calculate area under the curve values17.

7. Validation of core genes using independent bulk and single-cell datasets

GSE117261 was used as an independent external lung-tissue validation cohort containing 58 PAH samples and 25 failed-donor control samples18. This dataset was not used in the discovery differential-expression analysis, weighted gene co-expression network construction, or machine-learning feature selection. The expression matrix was normalized and annotated, and differential expression was analyzed using limma v3.68.0. The Benjamini-Hochberg false-discovery-rate correction was applied across the entire annotated transcriptome. Single-gene receiver operating characteristic (ROC) curves were calculated using pROC v1.19.0.1, DeLong 95% confidence intervals, and Youden-index cutoffs. An exploratory five-gene logistic regression model was fitted within GSE117261, and its internal performance was additionally evaluated using repeated nested cross-validation.

GSE210248 (Table 1) was used as a single-cell pulmonary artery validation dataset containing samples from three patients with PAH and three healthy donors19. The data were processed using Seurat v5.5.1 for quality control, normalization, dimensionality reduction, clustering, and cell annotation20. Major cell populations, including endothelial cells, smooth muscle cells, fibroblasts, monocytes/macrophages, and T/natural killer cells, were identified. Cell-cell communication was analyzed using CellChat v2.1.2 and the CellChatDB.human ligand-receptor database21. A CellChat object was created from the normalized Seurat expression matrix and cell-type metadata. Overexpressed genes and ligand-receptor interactions were identified; communication probabilities were calculated; interactions involving cell groups with fewer than 10 cells were removed; and pathway-level communication networks were inferred and aggregated. This dataset was used only for external mechanistic validation and not for model training.

ItemDescription
DatasetGSE210248
Data type10x Genomics/droplet-based single-cell RNA sequencing; high-throughput transcriptomic profiling
Human samplesThree PAH pulmonary artery samples and three healthy donor pulmonary artery samples
Tissue sourceEx vivo pulmonary artery tissue, primarily reflecting the cellular ecology of the pulmonary vascular wall and vascular-remodeling process
Main analytical purposeCell-type localization, smooth-muscle-cell phenotypic switching, immune-structural cell communication, and mechanistic consistency validation of candidate genes

Table 1: Basic information for the GSE210248 single-cell validation dataset. The table summarizes the dataset accession, sequencing platform, tissue source, sample composition, and analytical purpose of the single-cell pulmonary artery validation analysis.

8. Validation of gene expression by qRT-PCR

The qRT-PCR validation included 20 biologically independent PAH lung-tissue samples from patients with PH/PAH and 20 biologically independent control lung-tissue samples. Total RNA was extracted using the Total RNA Extraction Kit. RNA concentration and purity were assessed using a spectrophotometer, and RNA integrity was evaluated by agarose gel electrophoresis. Only RNA samples with A260/280 values between 1.8 and 2.1 and no visible degradation were included.

Equal amounts of RNA were reverse-transcribed into complementary DNA using the Solarbio Universal RT-PCR Kit (AMV; catalog no. RP1200). Quantitative PCR for CXCL10, JUN, IFIH1, MX1, and TLR7 was performed using SYBR Green PCR Master Mix on a Real-Time PCR System. Each biological sample was analyzed in three technical replicates, together with no-template and no-reverse-transcription controls. The mean Ct value of the three technical replicates was used for subsequent analysis; technical replicates were not treated as independent observations. Primers spanning exon-exon junctions and producing 80–200 bp amplicons were used (Table 2). Primer specificity was verified using NCBI Primer-BLAST and melting-curve analysis22.

β-actin (ACTB) was used as the internal reference gene to normalize the expression levels of the target genes. Relative expression was calculated using the 2-ΔΔCt method23. Two-sided Mann-Whitney U tests were used for between-group comparisons based on the data distribution, and Benjamini-Hochberg false-discovery-rate correction was applied across the five genes. Single-gene ROC curves were generated with DeLong 95% confidence intervals, and optimal cutoffs were selected using the Youden index. The five-gene logistic regression model was initially fitted and evaluated on the same 40 biological samples; this estimate was therefore defined as the apparent in-sample performance. To assess potential overfitting, 100 stratified five-fold cross-validation repeats were performed using an L2-regularized logistic regression model, and the pooled out-of-fold ROC performance was calculated.

GeneRefSeq accessionForward primer (5′–3′)Reverse primer (5′–3′)Product size (bp)Tm (°C)Exon-spanning
CXCL10NM_001565.4GTCAAGCCAT
AATTGTTC
ATAGTGCCAG
GGTAGAGT
14146.1Yes
JUNNM_002228.4ACAAGTGGCA
GAGTCCCG
CGCCCAAGTT
CAACAACC
15254.5Yes
IFIH1NM_022168GCACAGAGCG
GTAGACCCT
GCCCTGAAGC
ACGAGATG
18254.7Yes
MX1NM_002462.5TTAGCCGTGG
TGATTTAGC
CAAGGTGGAG
CGATTCTG
15652.3Yes
TLR7NM_016562.4ATTGCCCTCGT
TGTTATA
TTCCTGGAGTT
TGTTGAT
17948.1Yes
ACTBNM_001101.3CTCACCATGGAT
GATGATATCGC
AGGAATCCTTCT
GACCCATGC
19456.2Yes

Table 2: Primer sequences used for quantitative reverse transcription PCR. The table lists the target genes, RefSeq accession numbers, forward and reverse primer sequences, product sizes, melting temperatures, and exon-spanning status of the primers used for qRT-PCR.

9. Candidate compound screening and molecular docking

The upregulated and downregulated core-gene signatures were submitted to the Connectivity Map database to identify small molecules predicted to reverse the PH-associated expression profile7. Candidates were ranked by Logit score and prediction probability.

The three-dimensional structure of BRD-K91900765/VX-745 was obtained from PubChem under CID 303852524. Compound-related pharmacological information was curated from public drug databases, and structural descriptors were calculated using DrugBank and SwissADME25,26. Protein structures were obtained from the RCSB Protein Data Bank using the following PDB identifiers: CXCL10, 1LV9; JUN, 1JUN; IFIH1, 3B6E; MX1, 5GTM; TLR7, 7CYN; and MAPK14/p38α, 1OUK27. Blind cavity detection and molecular docking were performed using CB-Dock2 v2.0 with the AutoDock Vina v1.2.0 scoring engine28,29. Protein and ligand files were uploaded to CB-Dock2, candidate cavities were automatically detected, and docking was performed within the cavity-specific boxes generated by the server. For each protein, the cavity identifier, Vina score, cavity volume, docking-box center, docking-box dimensions, and protein-ligand complex file were recorded. The pose with the most negative Vina score was selected as the top-ranked predicted conformation. MAPK14/p38α was included as the established pharmacological target and positive-reference docking protein for VX-745. Docking against CXCL10, JUN, IFIH1, MX1, and TLR7 was exploratory and was not interpreted as evidence of direct pharmacological targeting, binding, inhibition, or efficacy.

10. Statistical analysis and reproducibility control

All statistical analyses were performed in R unless otherwise specified. Two-sided P values < 0.05 were considered statistically significant. Multiple-testing correction was applied to the differential-expression, enrichment, external-validation, and qRT-PCR analyses as specified above. Cross-validation was used to assess the stability of the machine-learning and combined qRT-PCR models.

Access restricted. Please log in or start a trial to view this content.

Results

Public transcriptomic data preprocessing and differentially expressed gene identification

Integration and ComBat correction of GSE22356, GSE33463, and GSE48149 reduced systematic differences among the datasets. Boxplots showed that sample expression distributions became more consistent after correction. Principal component analysis indicated that samples clustered mainly according to dataset source before correction but became more intermixed after correction, indicating effec...

Access restricted. Please log in or start a trial to view this content.

Discussion

An integrated and reproducible workflow was developed to identify PH/PAH-associated molecular biomarkers and candidate therapeutic compounds by combining public transcriptomics, weighted gene co-expression network analysis, functional enrichment, protein-protein interaction network analysis, three machine-learning algorithms, external validation, single-cell transcriptomic interpretation, quantitative reverse transcription PCR confirmation, Connectivity Map screening, and molecular docking. CXCL10, JUN, IFIH1, MX1, and T...

Access restricted. Please log in or start a trial to view this content.

Disclosures

The authors declare no competing interests.

Acknowledgements

This study was supported by the Hunan Innovative Province Construction Project (No. 2022JJ30465).

Access restricted. Please log in or start a trial to view this content.

Materials

List of materials used in this article
NameCompanyCatalog NumberComments
2× SYBR Green PCR MastermixBeijing Solarbio Science & Technology Co., Ltd.Catalog no. SR1110Dye-based quantitative real-time PCR amplification and fluorescence detection
AI21.msvmRFE.R and e1071Custom R script with the CRAN e1071 packageAI21.msvmRFE.R; e1071 v1.7-17Support vector machine-recursive feature elimination
AgaroseBeijing Solarbio Science & Technology Co., Ltd.Catalog no. A8201; CAS 9012-36-6Assessment of total RNA integrity by agarose gel electrophoresis
Agarose gel electrophoresis apparatusBeijing Liuyi Biotechnology Co., Ltd.Model DYCZ-24DNElectrophoretic assessment of RNA integrity
AutoDock Vina scoring engineCenter for Computational Structural Biology, Scripps Researchv1.2.0; RRID: SCR_011958Protein-ligand pose scoring within the CB-Dock2 workflow
CB-Dock2Cao Laboratory, CB-Dock2 web serverv2.0; accessed July 2026Blind cavity detection and molecular docking of VX-745 with the selected protein structures
CellChatCellChat R packagev2.1.2Inference and visualization of cell-cell communication from the single-cell expression matrix
CellChatDB.humanDistributed with the CellChat R packageCellChatDB.human; Secreted Signaling subset; minimum cell threshold = 10Human ligand-receptor interaction database for CellChat
clusterProfilerBioconductor R packagev4.20.0; Bioconductor release 3.23Gene Ontology and Kyoto Encyclopedia of Genes and Genomes enrichment analyses
Connectivity Map (CMap/CLUE)Broad InstituteL1000/CLUE resource; RRID: SCR_016204; accessed July 2026Computational drug-repositioning analysis
Custom oligonucleotide primersBeijing Solarbio Science & Technology Co., Ltd.Custom synthesized; primer sequences provided in Table 2Amplification of ACTB, CXCL10, JUN, IFIH1, MX1, and TLR7
cytoHubbaCytoscape App Storev0.1Degree-based hub-gene ranking in the protein-protein interaction network
CytoscapeCytoscape Consortiumv3.10.4; RRID: SCR_003032Protein-protein interaction network visualization and analysis
DrugBankDrugBank Knowledgebasev6.0; RRID: SCR_002700Compound identity and pharmacological-information curation
Gel documentation systemBeijing Liuyi Biotechnology Co., Ltd.Model WO-9413BVisualization and recording of agarose-gel RNA integrity results
Gene Expression Omnibus (GEO)National Center for Biotechnology InformationGSE22356, GSE33463, GSE48149, GSE117261, and GSE210248; RRID: SCR_005012Retrieval of bulk and single-cell transcriptomic datasets
GEOqueryBioconductor R packagev2.80.0; Bioconductor release 3.23Programmatic downloading and import of GEO expression and phenotype data
glmnetCRAN R packagev5.0Least absolute shrinkage and selection operator logistic regression and regularized logistic modeling
limmaBioconductor R packagev3.68.0; Bioconductor release 3.23; RRID: SCR_010943Differential expression analysis and empirical Bayes statistics
NanoDrop spectrophotometerThermo Fisher ScientificNanoDrop ND-1000; software v3.8Measurement of RNA concentration and A260/280 and A260/230 purity ratios
NCBI Primer-BLASTNational Center for Biotechnology InformationWeb tool; RRID: SCR_003095; accessed July 2026Verification of primer specificity
pROCCRAN R packagev1.19.0.1; RRID: SCR_024286Receiver operating characteristic analysis, DeLong confidence intervals, and Youden-index cutoffs
Protein Data Bank (PDB)RCSB Protein Data BankCXCL10: 1LV9; JUN: 1JUN; IFIH1: 3B6E; MX1: 5GTM; TLR7: 7CYN; MAPK14/p38α: 1OUK; RRID: SCR_012820Retrieval of experimentally determined protein structures for molecular docking
PubChemNational Center for Biotechnology InformationPubChem CID 3038525; RRID: SCR_004284Retrieval of the three-dimensional structure and chemical identifiers of BRD-K91900765/VX-745
RR Foundation for Statistical Computingv4.6.1; RRID: SCR_001905Statistical computing, data processing, machine learning, and visualization
randomForestCRAN R packagev4.7-1.2Random forest feature selection and variable-importance ranking
Real-time PCR systemStratagene, now Agilent TechnologiesMx3000P Real-Time PCR SystemqRT-PCR amplification, fluorescence acquisition, melting-curve analysis, and Ct export
SeuratCRAN R package; Satija Laboratoryv5.5.1; RRID: SCR_016341Single-cell RNA-sequencing quality control, normalization, dimensionality reduction, clustering, and annotation
STRINGSTRING Consortiumv12.0; RRID: SCR_005223Protein-protein interaction network construction
sva (ComBat)Bioconductor R packagev3.60.0; Bioconductor release 3.23Correction of inter-dataset batch effects
SwissADMESwiss Institute of BioinformaticsWeb server; accessed July 2026Drug-likeness, physicochemical-property, and ADME prescreening
Total RNA Extraction KitBeijing Solarbio Science & Technology Co., Ltd.Catalog no. R1200Extraction and purification of total RNA from lung-tissue samples
Universal RT-PCR Kit (AMV)Beijing Solarbio Science & Technology Co., Ltd.Catalog no. RP1200Reverse transcription of total RNA into complementary DNA
WGCNACRAN R packagev1.74Weighted gene co-expression network construction and module-trait analysis

References

  1. Simonneau G, et al. Haemodynamic definitions and updated clinical classification of pulmonary hypertension. Eur Respir J. 2019;53(1):1801913. doi: 10.1183/13993003.01913-2018.
  2. Rabinovitch M, Guignabert C, Humbert M, Nicolls MR. Inflammation and immunity in the pathogenesis of pulmonary arterial hypertension. Circ Res. 2014;115(1):165–175.
  3. Humbert M, et al. 2022 ESC/ERS Guidelines for the diagnosis and treatment of pulmonary hypertension. Eur Heart J. 2022;43(38):3618–3731.
  4. Galiè N, et al. 2015 ESC/ERS Guidelines for the diagnosis and treatment of pulmonary hypertension. Eur Respir J. 2015;46(4):903–975.
  5. Soon E, et al. Elevated levels of inflammatory cytokines predict survival in idiopathic and familial pulmonary arterial hypertension. Circulation. 2010;122(9):920–927.
  6. George PM, et al. Evidence for the involvement of type I interferon in pulmonary arterial hypertension. Circ Res. 2014;114(4):677–688.
  7. Subramanian A, et al. A next-generation Connectivity Map: L1000 platform and the first 1,000,000 profiles. Cell. 2017;171(6):1437–1452.e17.
  8. Leek JT, et al. The sva package for removing batch effects and other unwanted variation in high-throughput experiments. Bioinformatics. 2012;28(6):882–883.
  9. Ritchie ME, et al. limma powers differential expression analyses for RNA-sequencing and microarray studies. Nucleic Acids Res. 2015;43(7):e47. doi: 10.1093/nar/gkv007.
  10. Langfelder P, Horvath S. WGCNA: an R package for weighted correlation network analysis. BMC Bioinformatics. 2008;9:559. doi: 10.1186/1471-2105-9-559.
  11. Yu G, Wang LG, Han Y, He QY. clusterProfiler: an R package for comparing biological themes among gene clusters. OMICS. 2012;16(5):284–287.
  12. Kanehisa M, Goto S. KEGG: Kyoto Encyclopedia of Genes and Genomes. Nucleic Acids Res. 2000;28(1):27–30.
  13. Szklarczyk D, et al. The STRING database in protein–protein association networks and functional enrichment analyses for any sequenced genome of interest. Nucleic Acids Res. 2023;51(D1):D638–D646.
  14. Friedman J, Hastie T, Tibshirani R. Regularization paths for generalized linear models via coordinate descent. J Stat Softw. 2010;33(1):1–22.
  15. Guyon I, Weston J, Barnhill S, Vapnik V. Gene selection for cancer classification using support vector machines. Mach Learn. 2002;46:389–422.
  16. Breiman L. Random forests. Mach Learn. 2001;45(1):5–32.
  17. Robin X, et al. pROC: an open-source package for R and S+ to analyze and compare ROC curves. BMC Bioinformatics. 2011;12:77. doi: 10.1186/s12859-011-0077-8.
  18. Stearman RS, et al. Systems analysis of the human pulmonary arterial hypertension lung transcriptome. Am J Respir Cell Mol Biol. 2019;60(6):637–649.
  19. Crnkovic S, et al. Single-cell transcriptomics reveals skewed cellular communication and phenotypic shift in pulmonary artery remodeling. JCI Insight. 2022;7(20):e153471. doi: 10.1172/jci.insight.153471.
  20. Hao Y, et al. Integrated analysis of multimodal single-cell data. Cell. 2021;184(13):3573–3587.e29.
  21. Jin S, et al. Inference and analysis of cell–cell communication using CellChat. Nat Commun. 2021;12(1):1088. doi: 10.1038/s41467-021-21246-9.
  22. Bustin SA, et al. MIQE 2.0: revision of the Minimum Information for Publication of Quantitative Real-Time PCR Experiments guidelines. Clin Chem. 2025;71(6):634–651.
  23. Livak KJ, Schmittgen TD. Analysis of relative gene expression data using real-time quantitative PCR and the 2−ΔΔCt method. Methods. 2001;25(4):402–408.
  24. Kim S, et al. PubChem 2023 update. Nucleic Acids Res. 2023;51(D1):D1373–D1380.
  25. Knox C, et al. DrugBank 6.0: the DrugBank Knowledgebase for 2024. Nucleic Acids Res. 2024;52(D1):D1265–D1275.
  26. Daina A, Michielin O, Zoete V. SwissADME: a free web tool to evaluate pharmacokinetics, drug-likeness, and medicinal chemistry friendliness of small molecules. Sci Rep. 2017;7:42717. doi: 10.1038/srep42717.
  27. Burley SK, et al. RCSB Protein Data Bank: powerful new tools for exploring 3D structures of biological macromolecules for basic and applied research and education. Nucleic Acids Res. 2021;49(D1):D437–D451. doi: 10.1093/nar/gkaa1038.
  28. Eberhardt J, Santos-Martins D, Tillack AF, Forli S. AutoDock Vina 1.2.0: new docking methods, expanded force field, and Python bindings. J Chem Inf Model. 2021;61(8):3891–3898. doi: 10.1021/acs.jcim.1c00203.
  29. Liu Y, et al. CB-Dock2: improved protein-ligand blind docking by integrating cavity detection, docking, and homologous template fitting. Nucleic Acids Res. 2022;50(W1):W159–W164. doi: 10.1093/nar/gkac394.
  30. Duffy JP, et al. The discovery of VX-745: a novel and selective p38α kinase inhibitor. ACS Med Chem Lett. 2011;2(10):758–763.
  31. Sheng Y, et al. Crocin inhibits neutrophil migration and activation to treat hypoxic pulmonary hypertension through targeting HCK. Phytomedicine. 2025;148:157334. doi: 10.1016/j.phymed.2025.157334.
  32. Cui H, et al. Leonurine ameliorates hypoxic pulmonary hypertension by inhibiting cross-talk between neutrophils and endothelial cells via SERPINB3 targeting. Int Immunopharmacol. 2026;176:116467. doi: 10.1016/j.intimp.2026.116467.

Access restricted. Please log in or start a trial to view this content.

Reprints and Permissions

Tags

Biomarker IdentificationTranscriptomic DataDifferential ExpressionGene Co-ExpressionProtein Interaction NetworkSingle-Cell RNA SequencingDrug RepositioningQuantitative PCR