Research Article

Transcriptomic Identification of Renin—Angiotensin System-Related Candidate Biomarkers and External Testing of a Hypertension Diagnostic Model

DOI:

10.3791/71252

June 22nd, 2026

* These authors contributed equally

In This Article

Summary

Loading...
$$\rightleftharpoonup{xx}$$ $$\longleftharp{xx}$$, $$\longrightharp{xx}$$,

Using public blood transcriptome datasets, this study identified candidate renin-angiotensin system (RAS)-related biomarkers for hypertension. Bioinformatics and machine learning yielded an eight-gene signature tested in an independent blood-based cohort, linking RAS transcriptomic alterations to immune and inflammatory signatures requiring further validation.

Abstract

Loading...
$$\rightleftharpoonup{xx}$$ $$\longleftharp{xx}$$, $$\longrightharp{xx}$$,

This work aimed to identify candidate renin-angiotensin system (RAS)-related blood transcriptomic biomarkers associated with hypertension, construct an externally tested candidate diagnostic model, and validate selected genes in an Ang II-induced endothelial cell model. Two public microarray datasets, GSE75360 and GSE74144, were analyzed. Differential expression analysis was performed using limma, followed by GO/KEGG enrichment, preranked GSEA, CIBERSORT immune infiltration analysis, protein-protein interaction and regulatory network construction, and machine learning-based feature selection using logistic regression and random forest. A logistic regression model based on the selected genes was developed in GSE75360 and externally tested in GSE74144. Experimental validation was performed in Ang II-induced HUVECs using qRT-PCR, western blotting, ELISA, and CST3/FURIN loss- and gain-of-function assays, followed by CCK-8, Transwell, inflammatory, oxidative stress, and endothelial function analyses. In GSE75360, 173 differentially expressed genes were identified, including 18 RAS-related differentially expressed genes. Eight candidate genes, LRP1, CTSD, MTHFR, AUTS2, FURIN, CST3, FCER1G, and TBXAS1, were selected by combined machine learning analyses. Enrichment and immune infiltration analyses indicated that these genes were mainly associated with immune and inflammatory signatures. The eight-gene model showed high discrimination in GSE75360 and retained moderate performance in GSE74144. In Ang II-treated HUVECs, most candidate genes were upregulated at the mRNA level, and CST3, FURIN, and TBXAS1 were further validated at the protein level. CST3 and FURIN modulation altered Ang II-induced endothelial viability, migration, expression of inflammatory/adhesion markers, ROS accumulation, and eNOS/NO-related functional readouts. This study identified candidate RAS-related biomarkers and immune-associated signatures in hypertension and developed an externally tested candidate diagnostic model. The in vitro findings support the functional relevance of CST3 and FURIN in Ang II-induced endothelial responses, warranting further clinical and mechanistic validation.

Introduction

Loading...
$$\rightleftharpoonup{xx}$$ $$\longleftharp{xx}$$, $$\longrightharp{xx}$$,

This study focuses on hypertension, a chronic cardiovascular disease that is prevalent globally1. Hypertension significantly increases the risk of serious health complications, including cardiovascular and cerebrovascular events as well as renal damage. This poses a grave threat to patients' quality of life and longevity, while creating a significant social and economic burden2,3. Currently, the diagnosis and treatment of hypertension mainly rely on blood pressure monitoring and antihypertensive medications. However, because the causes of hypertension are complex and heterogeneous, many patients do not respond adequately to current treatments, and molecularly informed stratification remains limited4,5. Accordingly, there is a need to identify candidate biomarkers and transcriptomic signatures associated with hypertension, while avoiding overinterpretation of blood-based findings as definitive tissue-level mechanisms.

In recent years, the renin-angiotensin system (RAS) has garnered significant attention in the pathogenesis and progression of hypertension6. Existing studies have shown that abnormal expression of RAS-related genes is closely associated with blood pressure regulation disorders and target organ damage, but the exact molecular networks and key regulatory factors are not fully understood7. Because hypertension is increasingly recognized as a disorder involving vascular, immune, and inflammatory dysregulation, peripheral blood mononuclear cells and white blood cells provide accessible surrogate tissues that may capture systemic RAS-associated transcriptomic alterations8. In this study, the RAS-related gene set was curated from GeneCards and PubMed searches and therefore included both canonical RAS pathway genes and genes reported in the literature to be functionally associated with RAS signaling.

This study employs two publicly available microarray datasets (GSE75360 and GSE74144), using bioinformatics and machine learning methodologies for transcriptomic analysis. Specific methods include limma differential analysis, GO/KEGG/GSEA enrichment analyses, CIBERSORT immune infiltration analysis, PPI network construction, and machine learning-based key gene selection and logistic regression diagnostic model construction and evaluation. The advantages of these methods lie in their ability to analyze large-scale transcriptomic data, externally test findings across datasets, and integrate complementary analytic approaches for candidate model construction. We hypothesized that blood-cell transcriptomic profiles could identify candidate RAS-related biomarkers linked to hypertension and associated immune signatures. Therefore, our aim was to identify candidate RAS-related genes, characterize their functional context, and construct an externally tested candidate diagnostic model, rather than to establish definitive molecular mechanisms or a clinically ready precision diagnostic tool.

Access restricted. Please log in or start a trial to view this content.

Protocol

Loading...
$$\rightleftharpoonup{xx}$$ $$\longleftharp{xx}$$, $$\longrightharp{xx}$$,

Overview of analysis workflow

The overall design of this study’s transcriptomic and machine learning-based analysis is illustrated in Figure 1, encompassing key steps: collection of Renin-Angiotensin System-related genes (RASRGs); screening of RAS-related differentially expressed genes (RASRDEGs) from hypertension datasets; functional enrichment analysis (GO/KEGG/GSEA); immune infiltration analysis (CIBERSORT); construction of protein-protein interaction (PPI) and regulatory networks; machine learning-based key gene selection (logistic regression, random forest [RF]); and evaluation of the hypertension diagnostic model. A complete list of software, databases, and online tools used in this study is provided in the Table of Materials.

Data download

Hypertension datasets GSE753608 and GSE74144 (Homo sapiens) were obtained via the R package GEOquery9 from the GEO database10. GSE75360 derived from peripheral blood mononuclear cells (platform: GPL10558) included 10 hypertension and 11 control samples; GSE74144 derived from white blood cells (platform: GPL13497) included 14 hypertension and 8 control samples (Table 1). Protein-coding RASRGs (1,264) were initially identified via GeneCards11 (keyword: "Renin-Angiotensin System") and PubMed (keyword: "Renin-Angiotensin System")12,13. The intersection of these RASRGs with genes in GSE75360/GSE74144 yielded 1,159 final RASRGs14. The two datasets were processed separately because they were generated on different microarray platforms. Probe annotation was performed according to the corresponding GPL platform annotation files, and normalized gene expression matrices were used for downstream analyses. Box plots were used to compare expression distributions before and after normalization.

Hypertension-related renin-angiotensin-related differentially expressed genes

Samples in the GSE75360 dataset were categorized into the hypertension group and control group. The limma software was employed to conduct differential gene expression analysis between the two groups14, with differentially expressed genes (DEGs) identified by the threshold of |logFC| > 0.45 and p-value < 0.05. The results of this differential analysis were visualized via volcano plots (generated using the R package ggplot2).

To obtain RASRDEGs, DEGs meeting the above threshold (|logFC| > 0.45, p-value < 0.05) were cross-referenced with RAS-related genes (RASRGs), and the intersection result was presented via a Venn diagram. Subsequently, the expression patterns of the identified RASRDEGs were visualized as a heatmap using the R package pheatmap, and the chromosomal localization of RASRDEGs was displayed via chromosome maps generated using the R package RCircos15.

Differentially expressed gene validation and ROC curve analysis

An intergroup plot was constructed to analyze RASRDEG expression differences between hypertension/control in GSE7536016, R package pROC was used to plot ROC curves and calculate AUC (0.5–0.7: low accuracy; 0.7–0.9: moderate; >0.9: high) for RASRDEG diagnostic efficacy.

Correlation analysis

Spearman’s correlation analysis was performed on RASRDEG expression in GSE75360; results were visualized via heatmap (R package ggplot2) (|r| < 0.3: no/weak correlation; 0.3–0.5: weak; 0.5–0.8: moderate; >0.8: strong).

Enrichment analysis of GO and KEGG

GO (Gene Ontology, 2024 release, http://geneontology.org/) is a widely used resource for large-scale functional enrichment, covering three domains: biological processes (BP), cellular components (CC), and molecular functions (MF)17. KEGG (Kyoto Encyclopedia of Genes and Genomes, Release 109.0, 2024, https://www.genome.jp/kegg/) stores data on genomes, biopathways, diseases, and drugs18.

RASRDEGs were subjected to GO annotation and KEGG pathway enrichment analysis using the R package clusterProfiler19. Enrichment test method: hypergeometric test; multiple test correction method: Benjamini-Hochberg (BH) method. Screening criterion: adjusted p-value < 0.05.

Gene set enrichment analysis (GSEA)

For cohort-level GSEA, all genes tested in the differential expression analysis of GSE75360 were ranked in descending order by logFC and used as the input gene list for clusterProfiler19. No DEG prefiltering was applied before GSEA. The c2 gene set collection from MSigDB20. Parameters: seed = 2022, 10–500 genes per set; screening criteria: corrected p < 0.05 (Benjamini-Hochberg, BH method), FDR < 0.2521.

Construction of hypertension diagnostic model

To identify key genes associated with hypertension, we employed two types of machine learning algorithms: logistic regression and random forests (RF). Logistic regression (binary dependent variable: hypertension/control) screened RASRDEGs with p < 0.05. Random Forest (RF, R package randomForest): parameters set.seed(520), ntree = 1000; MeanDecreaseGini (variable importance indicator) was extracted, and top15 RASRDEGs were selected. RASRDEGs were screened with a p value < 0.05 as the standard.

The RF (Random Forest) algorithm, an ensemble learning method under the Bagging category (integrating multiple decision trees), was applied via the R package randomForest22 (parameters: set.seed(520), ntree = 1000). MeanDecreaseGini (reflecting variable importance by average purity decrease during node splitting) of feature genes was extracted, and the top 15 RASRDEGs were selected. Finally, a Venn diagram of genes screened by logistic regression and RF was plotted to identify hypertension-related key genes.

Validation of hypertension diagnostic model

A logistic regression model was built based on key genes; linear predicted value (η) was calculated as:

Gene expression model equation, η = Intercept + Σ Coefficient(gene) * mRNA Expression(gene).

The R package pROC16 was used to plot ROC curves and evaluate the model’s efficacy in predicting hypertension risk. A nomogram was constructed via the R package rms23 to visualize the contribution of each key gene to the logistic regression model (reflecting the association between key genes and hypertension risk). Calibration curves were generated to assess the consistency between predicted and actual hypertension probabilities; decision curve analysis (DCA, R package ggDCA24) was performed to evaluate the model’s clinical utility (net benefit) in GSE75360 and GSE74144.

Single-gene GSEA

GSEA explores the role of genes associated with a specific gene in biological processes/pathways/diseases by analyzing its expression, aiding in understanding the gene’s functional role. For each focal gene in GSE75360, samples were split at the median into high- and low-expression groups. Differential expression analysis was then performed across all tested genes, and genome-wide logFC values were ranked from highest to lowest before GSEA with clusterProfiler19. No DEG prefiltering was applied before GSEA. Parameters: seed = 2020, 10–500 genes per set (c2 gene set collection from MSigDB21). Screening criteria: p < 0.05 (adj. p corrected via BH method).

Immune infiltration analysis (CIBERSORT)

The CIBERSORT algorithm25 (based on linear support vector regression) deconvoluted the transcriptome matrix to estimate immune cell composition in mixed samples (data with immune cell enrichment score > 0 were selected). The final immune cell infiltration matrix of GSE75360 was visualized via a proportion bar chart. Spearman’s correlation was used to analyze immune cell-immune cell and key gene-immune cell associations, with results presented as a correlation heatmap (R package pheatmap) and correlation bubble plot (R package ggplot2), respectively.

Protein-protein interaction (PPI) network

PPI networks are systems of interconnected proteins regulating biological processes via interactions. Using the STRING database26, a PPI network for key genes was constructed (minimum interaction score: 0.150, low confidence). Renin-angiotensin-related hub genes were selected by screening interacting genes. The GeneMANIA database27, which identifies functionally similar genes using genomic and proteomic datasets, was used to predict functionally similar genes of key RAS genes and to construct a protein interaction network.

Construction of regulatory network

mRNA-TF network: Transcription factors (TFs) regulate gene expression via post-transcriptional interaction with target genes. TFs targeting hub genes and their regulatory relationships were retrieved from the ChIPBase database28, and the mRNA-TF network was visualized using Cytoscape29.

mRNA-miRNA network: miRNAs modulate multiple target genes (single targets may be co-regulated by multiple miRNAs). StarBase v3.030 was used to identify miRNAs associated with RASRDEGs, and the mRNA-miRNA network was visualized via Cytoscape.

mRNA-drug network: Toxicogenomic databases31 were used to predict direct/indirect drug targets of hub genes. The mRNA-drug network (showing gene-drug interactions) was visualized with Cytoscape to complete network construction.

Ang II-induced HUVEC model

Human umbilical vein endothelial cells (HUVECs) were maintained at 37°C in a humidified incubator with 5% CO2. Cells were maintained in complete endothelial cell culture medium supplemented with fetal bovine serum and antibiotics according to the supplier’s instructions. To establish an in vitro hypertension-related endothelial injury model, HUVECs were treated with angiotensin II (Ang II; 100 nM) for 48 h. Vehicle-treated cells were used as the control group.

For gene intervention experiments, small interfering RNAs targeting CST3 or FURIN (si-CST3 and si-FURIN), corresponding negative control siRNA (si-NC), CST3 or FURIN overexpression plasmids (oe-CST3 and oe-FURIN), and the corresponding empty-vector control (oe-NC) were transfected into HUVECs using a commercial transfection reagent according to the manufacturer’s protocol. After transfection, cells were exposed to Ang II and then harvested for expression validation and functional assays. Knockdown and overexpression efficiencies were confirmed by qRT-PCR and western blotting.

qRT-PCR

Total RNA was isolated from HUVECs with a standard RNA extraction reagent, and complementary DNA was generated using a reverse transcription kit. SYBR Green chemistry was used for qRT-PCR. Expression levels of LRP1, CTSD, MTHFR, AUTS2, FURIN, CST3, FCER1G, TBXAS1, IL-6, TNF-α, VCAM1, ICAM1, and eNOS were normalized to GAPDH. and calculated by the 2−ΔΔCt method.

Western blotting

For western blot analysis, proteins were extracted with RIPA lysis buffer and quantified using a BCA assay. Equal protein amounts were resolved by SDS-PAGE and transferred to PVDF membranes. After blocking, membranes were incubated with primary antibodies against CST3, FURIN, TBXAS1, or GAPDH and then with suitable secondary antibodies. Bands were detected by chemiluminescence, and densitometry was normalized to GAPDH. Secreted CST3 in culture supernatants was quantified with an ELISA kit following the manufacturer’s protocol.

Cell viability

Cell viability was assessed using the Cell Counting Kit-8 (CCK-8) assay. Briefly, transfected and Ang II-treated HUVECs were seeded into 96-well plates, and absorbance at 450 nm was measured at 0, 24, 48, and 72 h after addition of the CCK-8 reagent. Cell migration was evaluated using Transwell chambers. After the indicated interventions, cells were seeded into the upper chambers, and migrated cells on the lower membrane surface were fixed, stained, and counted under a microscope in randomly selected fields.

Inflammatory test

To evaluate inflammatory activation, oxidative stress, and endothelial function, IL-6, TNF-α, VCAM1, ICAM1, and eNOS .mRNA levels were detected by qRT-PCR. Nitric oxide (NO) levels in the culture supernatant were measured using a commercial NO assay kit, and intracellular reactive oxygen species (ROS) levels were detected using DCF fluorescence according to the manufacturer’s instructions.

Statistical analysis

Transcriptomic processing and modeling were performed in R. Continuous variables were assessed for normality with the Shapiro-Wilk test. For two-group comparisons, independent-samples t-tests were used for normally distributed variables, whereas Wilcoxon rank-sum tests were used for non-normal variables. For three or more groups, one-way analysis of variance with appropriate post hoc testing was used when normality and homogeneity of variance assumptions were met; otherwise, the Kruskal-Wallis test was applied. CCK-8 time-course data were analyzed using two-way analysis of variance. Spearman correlation coefficients were calculated for association analyses. Unless otherwise stated, experimental results are shown as mean ± SD, and two-tailed p < 0.05 was considered significant.

Access restricted. Please log in or start a trial to view this content.

Results

Loading...
$$\rightleftharpoonup{xx}$$ $$\longleftharp{xx}$$, $$\longrightharp{xx}$$,

Cleaning of hypertension datasets

To ensure the reliability of subsequent analyses, datasets GSE75360 and GSE74144 were first subjected to probe annotation and data normalization using the R package limma. The distributions of gene expression values before and after normalization are shown in Supplemental File 1Supplemental Figure S1A-D, with orange representing hypertension samples and blue representing control samples.

Access restricted. Please log in or start a trial to view this content.

Discussion

Loading...
$$\rightleftharpoonup{xx}$$ $$\longleftharp{xx}$$, $$\longrightharp{xx}$$,

Hypertension is a complex cardiovascular disorder involving RAS dysregulation, immune-inflammatory activation, vascular injury, and endothelial dysfunction32. In the present study, we first used public blood transcriptomic datasets and machine learning approaches to identify RAS-related candidate biomarkers and to build an externally tested candidate diagnostic model. We then extended these bioinformatics findings through experimental validation in an Ang II-induced HUVEC model. This combined stra...

Access restricted. Please log in or start a trial to view this content.

Disclosures

Loading...
$$\rightleftharpoonup{xx}$$ $$\longleftharp{xx}$$, $$\longrightharp{xx}$$,

The authors have no conflicts of interest to declare.

Authors' contributions:

Conception and design of the work: Bai J, Wang Y;

Data collection: Yang X X, Du L, Shi J H, Li X X, Bai M;

Supervision: Bai J, Wang Y;

Analysis and interpretation of the data: Yang X X, Du L, Shi J H, Li X X, Bai M;

Statistical analysis: Bai J, Wang Y, Bai M;

Drafting the manuscript: Bai J, Wang Y;

Critical revision of the manuscript: all authors;

Approval of the final manuscript: all authors.

Acknowledgements

Loading...
$$\rightleftharpoonup{xx}$$ $$\longleftharp{xx}$$, $$\longrightharp{xx}$$,

Funding: Gansu Provincial Department of Education Higher Education Faculty Innovation Fund Project (No. 2026B-268)

Access restricted. Please log in or start a trial to view this content.

Materials

List of materials used in this article
NameCompanyCatalog NumberComments
96-well cell culture platesUsed for CCK-8 cell viability assays
Angiotensin II (Ang II)Used to induce endothelial injury in HUVECs at 100 nM for 48 h
Antibiotic solutionAdded to endothelial cell culture medium according to the supplier’s instructions
BCA protein assay kitUsed to determine protein concentration for western blotting
Cell Counting Kit-8 (CCK-8)Used to assess HUVEC viability/proliferation
Chemiluminescence detection reagentUsed to visualize western blot protein bands
ChIPBase v2.0Sun Lab / Research ResourceN/ADatabase used to retrieve transcription factor-target relationships for hub genes
CIBERSORTStanford University / Newman LabN/AImmune cell deconvolution algorithm used to estimate the abundance of 22 immune cell types
clusterProfiler (R package)BioconductorN/AR package used for GO, KEGG, and GSEA analyses
CO2 cell culture incubatorUsed to culture HUVECs at 37 °C with 5% CO2
Comparative Toxicogenomics Database (CTD)NC State University / Mount Desert Island Biological LaboratoryN/ADatabase used to predict gene-drug interactions
Complete endothelial cell culture mediumUsed for HUVEC culture
Computer workstation with internet accessAny standard research computing platformN/AUsed for data download, preprocessing, statistical analysis, and visualization
CST3 ELISA kitUsed to detect secreted CST3 in cell culture supernatant
CST3 overexpression plasmid (oe-CST3)Used for CST3 gain-of-function experiments in HUVECs
CytoscapeCytoscape ConsortiumN/ASoftware used to visualize mRNA-miRNA, mRNA-TF, and mRNA-drug regulatory networks
DCF fluorescent probe/ROS assay reagentUsed to detect intracellular ROS levels
Empty-vector control plasmid (oe-NC)Used as the overexpression control
Fetal bovine serumSupplement added to endothelial cell culture medium
FURIN overexpression plasmid (oe-FURIN)Used for FURIN gain-of-function experiments in HUVECs
GeneCardsWeizmann Institute of ScienceN/ADatabase used to curate RAS-related genes
GeneMANIAUniversity of TorontoN/AWeb-based tool used to identify functionally related genes
GEO databaseNCBIN/APublic database used to access GSE75360 and GSE74144 datasets
GEOquery (R package)BioconductorN/AR package used to download GEO datasets
ggDCA (R package)CRAN / GitHub source used by authorsN/AR package used for decision curve analysis
ggplot2 (R package)CRANN/AR package used for data visualization
Human umbilical vein endothelial cells (HUVECs)Cell model used for Ang II-induced endothelial injury experiments
limma (R package)BioconductorN/AR package used for normalization and differential expression analysis
MSigDBBroad InstituteN/AGene set database used for GSEA
Negative control siRNA (si-NC)Used as the knockdown control
Nitric oxide assay kitUsed to measure NO levels in culture supernatant
pheatmap (R package)CRANN/AR package used to generate heatmaps
Primary antibody against CST3Used for western blotting
Primary antibody against FURINUsed for western blotting
Primary antibody against GAPDHUsed as the western blot loading control
Primary antibody against TBXAS1Used for western blotting
pROC (R package)CRANN/AR package used to generate ROC curves and calculate AUC
PubMedU.S. National Library of MedicineN/ALiterature database used to supplement RAS-related gene curation
PVDF membraneUsed for western blot protein transfer
qRT-PCR primersUsed to detect candidate genes and inflammatory/endothelial markers
R softwareR Foundation for Statistical ComputingN/AStatistical computing environment used for all analyses
randomForest (R package)CRANN/AR package used for feature selection by random forest
RCircos (R package)CRAN / Bioconductor-associated resourceN/AR package used to visualize chromosomal localization of RASRDEGs
Reverse transcription kitUsed to synthesize complementary DNA
RIPA lysis bufferUsed for total protein extraction
rms (R package)CRANN/AR package used to construct the nomogram
RNA extraction reagentUsed to extract total RNA from HUVECs
Secondary antibodiesUsed for western blotting
siRNA targeting CST3 (si-CST3)Used for CST3 knockdown experiments
siRNA targeting FURIN (si-FURIN)Used for FURIN knockdown experiments
starBase v3.0Sun Lab / Research ResourceN/ADatabase used to identify miRNA-target interactions
STRINGSTRING ConsortiumN/ADatabase used to construct the protein-protein interaction network
SYBR Green qPCR reagentUsed for qRT-PCR detection
Transfection reagentUsed to transfect siRNAs and overexpression plasmids into HUVECs
Transwell chambersUsed for HUVEC migration assays

Reprints and Permissions

Request permission to reuse the text or figures of this JoVE article

Request Permission

Tags

Renin Angiotensin SystemHypertension BiomarkersTranscriptomic AnalysisDifferential ExpressionImmune InfiltrationMachine Learning ModelEndothelial Cell ModelGene ValidationProtein Interaction NetworkLogistic Regression

Related Articles