Research Article

Identification of Candidate Biomarkers Associated with Mitochondrial Dysfunction and SUMOylation in Heart Failure Based on Bioinformatics Approaches

DOI:

10.3791/72265

June 26th, 2026

In This Article

Summary

Loading...
$$\rightleftharpoonup{xx}$$ $$\longleftharp{xx}$$, $$\longrightharp{xx}$$,

Using bioinformatics, machine learning, and qPCR validation, this study identified five candidate biomarkers associated with SUMOylation and mitochondrial dysfunction in heart failure. These findings improve understanding of heart failure mechanisms and suggest potential directions for future diagnostic research.

Abstract

Loading...
$$\rightleftharpoonup{xx}$$ $$\longleftharp{xx}$$, $$\longrightharp{xx}$$,

Heart failure (HF) presents a persistent clinical challenge. While SUMOylation and mitochondrial function are vital for cardiomyocyte health, their combined influence on HF remains elusive. Two HF-related datasets were downloaded from GEO. The overlapping genes were obtained from all genes in the training set, SUMO-related genes, and mitochondrial-related genes. Three machine learning algorithms were applied to identify diagnostic key genes. Subsequently, diagnostic models were constructed and evaluated based on these genes. Besides, the immune microenvironment in HF versus healthy controls was assessed using CIBERSORT, MCP-counter, and ssGSEA. The differences in immune infiltration between HF and healthy controls were analyzed. Drug prediction and molecular docking were performed to identify potential drug candidates targeting these genes. Finally, qPCR was employed to validate gene expression levels in clinical samples. A total of 113 common genes with notable enrichment in mitochondrial regulation were identified. Five key genes, namely NFKB1, MYEF2, NSUN2, SQSTM1, and FKBP4, were identified by three machine learning algorithms. Functional enrichment analyses linked these genes to immune response, RNA processing, and cell cycle regulation. Moreover, immune infiltration profiling revealed that neutrophil infiltration contributes to dysregulated immune responses in HF. Molecular docking revealed that the small-molecule drug IMX-942 has a favorable binding affinity with SQSTM1 (-5.8 kcal/mol). qPCR validation supported the bioinformatics results. NFKB1, MYEF2, NSUN2, SQSTM1, and FKBP4 were identified as key genes linking SUMOylation and mitochondrial function in HF. These findings provide new insights into HF pathophysiology and may contribute to the development of novel diagnostic and therapeutic strategies.

Introduction

Loading...
$$\rightleftharpoonup{xx}$$ $$\longleftharp{xx}$$, $$\longrightharp{xx}$$,

Heart failure (HF), the end-stage of various cardiovascular diseases, is characterized by impaired cardiac function that fails to meet the body's metabolic demands1. This debilitating condition poses significant threats to patient health, leading to decreased quality of life and elevated mortality rates2. Current diagnostic modalities for HF primarily include biochemical marker detection3,4, echocardiography, and radiological imaging5. Although available treatments encompass pharmacological agents, device-based interventions, and surgical procedures, clinical outcomes remain unsatisfactory6. Limitations such as adverse drug reactions, limited applicability of devices, immune rejection, and other complications frequently hinder therapeutic efficacy7,8,9,10,11. Therefore, there is an urgent need to elucidate the underlying mechanisms of HF, identify early and precise diagnostic biomarkers, and develop more effective and safer therapeutic strategies.

Small ubiquitin-like modifier (SUMO) proteins are covalently conjugated to the lysine residues of substrate proteins through a dynamic and reversible process, regulating the structure and function of substrate proteins12. SUMOylation, a critical post-translational modification, serves as a key regulator of various cellular processes13,14. Mitochondria, as the energy-metabolism center of cells, are critically involved in the pathogenesis of HF. In the pathological process of HF, mitochondrial dysfunction, such as insufficient ATP production, reactive oxygen species (ROS) burst, and Ca2+ homeostasis imbalance, contributes significantly to progression15,16,17,18. Notably, emerging evidence suggests a potential interplay between SUMOylation and mitochondrial function. Mitochondrial stress can trigger SUMOylation-related pathways, while SUMO proteins and their specific proteases are essential for maintaining mitochondrial homeostasis19,20,21. Recent studies have further highlighted the importance of mitochondrial quality control and mitochondrial dynamics in cardiovascular diseases and HF progression22,23. However, the synergistic effect of SUMOylation and mitochondrial regulation on HF development remains unclear, particularly at the gene level.

In this study, we systematically investigated genes at the intersection of SUMOylation and mitochondrial dysfunction in HF, two key biological processes that have been individually implicated in HF but not yet comprehensively integrated. Overlapping genes were identified by intersecting HF-related differentially expressed genes, SUMOylation-related genes, and mitochondria-related genes. Key genes were then screened using machine learning algorithms and used to construct a diagnostic model. Functional enrichment and immune infiltration analyses were further performed to explore their potential biological roles in HF. This integrated approach may provide a systematic framework for exploring the crosstalk between SUMOylation and mitochondrial dysfunction in HF.

Access restricted. Please log in or start a trial to view this content.

Protocol

Loading...
$$\rightleftharpoonup{xx}$$ $$\longleftharp{xx}$$, $$\longrightharp{xx}$$,

The study was conducted in accordance with the Declaration of Helsinki, and the protocol was approved by the Ethics Committee of the Third Hospital of Hebei Medical University (W2025-065-1) in November 2024. Informed consent was obtained from all subjects involved in the study.

Data source and preprocessing

RNA-seq data associated with HF were obtained, including two microarray datasets from the Gene Expression Omnibus (https://www.ncbi.nlm.nih.gov/geo/). Two peripheral blood microarray datasets were selected: GSE59867 (34 HF samples and 30 controls) was used as the training dataset; GSE57338 (177 HF samples and 136 controls) was used as the validation dataset. Clinical information available for GSE57338, including age, gender, and disease status, was retrieved from GEO and is summarized in Supplementary Table 1. In addition, a total of 3,893 SUMOylation-related genes (SRGs) were obtained from the dbPTM database (https://awi.cuhk.edu.cn/dbPTM/index.php) (Supplementary Table 2), while 2,030 mitochondria-related genes (MRGs) were collected based on a previous study24 (Supplementary Table 3). Next, the R package GEOquery (v 2.72.0)25 was used to download datasets from the GEO database, extract the expression matrix, and obtain the sample phenotype information. Annotation was performed by mapping the annotation file and matching the gene IDs. Invalid gene IDs were removed, and the most highly expressed probes were retained.

Key gene selection via machine learning

A multi-step approach was used to select the genes related to the HF, SUMOylation, and mitochondria. First, the common genes between the training dataset, the SRGs, and the MRGs were identified using intersection analysis. The potential function of common genes was identified by Gene Ontology (GO) and Kyoto Encyclopedia of Genes and Genomes (KEGG) enrichment analysis using the R package ClusterProfiler (v 4.12.6)26. Then, three machine learning approaches, namely LASSO regression, XGBoost, and random forest (RF), were employed to further filter the genes. In LASSO regression, the optimal regularization parameter λ was selected via cross-validation to identify the genetic features with the greatest predictive value. The non-zero coefficient genes were selected for subsequent analysis. Then, XGBoost and RF algorithms were used to calculate the feature importance scores and screen the top 20 genes.

Construction and evaluation of diagnostic models

A diagnostic model was constructed using Logistic regression based on the GSE59867 dataset. The model was then applied to predict disease status and calculate probability scores. To validate the model, the same key genes were extracted from the GSE57338 dataset, normalized to match the training dataset, and used for external prediction. Model performance was assessed using Receiver Operating Characteristic (ROC) curves, Confusion Matrix, Calibration Curve, and Decision Curve Analysis (DCA).

Gene set enrichment analysis (GSEA) and subcellular localization

Spearman correlation analysis was used to identify correlated genes for each key gene. GSEA analysis was performed using the R package ClusterProfiler (v 4.12.6) on the key genes' related genes. Meanwhile, to determine the precise subcellular localization of the key genes within the cell, their subcellular localization was determined using the GeneCards database (https://www.genecards.org/).

Gene-disease association and drug prediction

To evaluate the clinical relevance of the identified key genes, systematic disease-association and drug-interaction analyses were performed. Disease-gene associations were interrogated using the Comparative Toxicogenomics Database (CTD; https://ctdbase.org/), with results ranked by both inference scores and reference counts (top 10 associations reported). Gene-drug interaction data for key genes were obtained from the Drug-Gene Interaction database (DGIdb), and drugs were excluded based on an interaction score < 0.5. Subsequently, we downloaded the 3D structures of proteins corresponding to key genes from the PDB database (https://www.rcsb.org/) and the molecular structures of potential drugs from PubChem (https://pubchem.ncbi.nlm.nih.gov/). Next, molecular docking analysis was performed using CB-Dock227 (https://cadd.labshare.cn/cb-dock2/php/index.php) to calculate the binding scores between the potential drugs and proteins. A lower binding free energy indicates a more stable interaction, suggesting the compound may have greater targeting potential.

Immune infiltration analysis

Immune cell infiltration was assessed using three complementary methods: Microenvironment Cell Populations-counter (MCP-counter)28, cell-type identification by estimating relative subsets of RNA transcripts (CIBERSORT)29 and single-sample enrichment analysis (ssGSEA)30. MCP-counter and CIBERSORT analysis was performed using the R package IOBR (v 0.99.0)31. MCP-counter was used to estimate immune and stromal cell abundance, while CIBERSORT was used to quantify the relative proportions of 22 immune cell types. ssGSEA was performed using the GSVA package (v1.52.3)32 to evaluate sample-level enrichment of immune cell subtypes.

Construction of the competing endogenous RNA (ceRNA) regulatory network

To investigate the potential miRNA–lncRNA regulatory roles associated with previously identified key genes, a ceRNA regulatory network was constructed. The R package multiMiR (v 1.26.0)33 was used to predict potential microRNA (miRNA)–mRNA interactions for key genes, integrating data from PITA (https://omictools.com/pita-tool/) and the miRDB database (https://mirdb.org/). miRNA–mRNA pairs with high confidence and consistency were selected. Subsequently, lncRNA–miRNA interactions were retrieved from the StarBase database (https://rnasysu.com/encori/) and filtered for interactions supported by ≥ 10 CLIP-seq experiments and categorized as lincRNAs. A ceRNA network was constructed by integrating lncRNA-miRNA-mRNA interactions.

qPCR validation

To validate the expression of key genes, Blood samples from patients with HF and healthy controls were collected from the clinical cohort (n = 6 per group) at the Third Hospital of Hebei Medical University (W2025-065-1) under approved protocols and informed consent. Total RNA was isolated using the TRIzol reagent in conjunction with chloroform and isopropanol. Following extraction, RNA was dissolved in DEPC-treated water, and its concentration and purity were assessed using a NanoDrop spectrophotometer. For transcriptional analysis, RNA was reverse transcribed into cDNA using the Fast First-Strand cDNA Synthesis Mix for RT (with dsDNase). Quantitative PCR was subsequently performed using the Fast Taq qPCR SYBR Green Mix. The specific primer sequences are detailed in the Table of Materials. Relative gene expression levels were calculated using the 2-ΔΔCT method, with appropriate normalization.

Statistical analysis

All statistical analyses were performed using R software and GraphPad Prism. Statistical comparisons between two independent groups were performed using either Student's t-test or the Mann-Whitney U test, depending on the data distribution. A p-value of less than 0.05 was considered to indicate statistical significance.

Access restricted. Please log in or start a trial to view this content.

Results

Loading...
$$\rightleftharpoonup{xx}$$ $$\longleftharp{xx}$$, $$\longrightharp{xx}$$,

Identification and functional enrichment of intersecting genes

To identify genes involved in SUMOylation and mitochondrial function in HF, quality control was first performed on the training set GSE59867 (Supplementary Figure 1A). A three-way intersection analysis was conducted among all genes in the training set, SRGs, and MRGs, identifying 113 overlapping genes (Figure 1A). To explore the potential biological functions of these ...

Access restricted. Please log in or start a trial to view this content.

Discussion

Loading...
$$\rightleftharpoonup{xx}$$ $$\longleftharp{xx}$$, $$\longrightharp{xx}$$,

HF, a progressive and terminal stage of various cardiovascular diseases, is characterized by highly complex and multifactorial pathophysiological mechanisms17,34. Although both SUMOylation and mitochondrial dysfunction have been individually implicated in HF, their potential synergistic roles remain insufficiently explored, particularly at the gene level. In the present study, we identified five key genes—NFKB1, MYEF2, NSUN2, SQSTM1, and FKBP4—and est...

Access restricted. Please log in or start a trial to view this content.

Disclosures

Loading...
$$\rightleftharpoonup{xx}$$ $$\longleftharp{xx}$$, $$\longrightharp{xx}$$,

This work was supported by the Medical Science Research Project of Hebei (Grant number: 20250084).

Acknowledgements

Loading...
$$\rightleftharpoonup{xx}$$ $$\longleftharp{xx}$$, $$\longrightharp{xx}$$,

The Authors have no conflicts of interest to declare.

Access restricted. Please log in or start a trial to view this content.

Materials

List of materials used in this article
NameCompanyCatalog NumberComments
Chloroform substituteServicebioG3014-02qPCR reagent
DEPC-treated water BiosharpBL510AqPCR reagent
Fast First-Strand cDNA Synthesis Mix for RT (with dsDNase)Albatross Biology500-101qPCR reagent
Fast Taq qPCR SYBR Green MixAlbatross Biology500-102qPCR reagent
FKBP4 primersTsingkeN/AForward: 5’-GAAGGCGTGCTGAAGGTCAT-3’
Reverse: 5’-TGCCATCTAATAGCCAGCCAG-3’
IsopropanolHushi80109218qPCR reagent
MYEF2 primersTsingkeN/AForward: 5’-CAGCTCCAATGGCGTTAAAATG-3’
Reverse: 5’-TGGCCTTCTTACTTCCTGTAGAT-3’
NanoDrop spectrophotometerThermo Fisher ScientificNanoDrop 2000CqPCR reagent
NFKB1 primersTsingkeN/AForward: 5’-AACAGAGAGGATTTCGTTTCCG-3’
Reverse: 5’-TTTGACCTGAGGGTAAGACTTCT-3’
NSUN2 primersTsingkeN/AForward: 5’-GAACTTGCCTGGCACACAAAT-3’
Reverse: 5’-TGCTAACAGCTTCTTGACGACTA-3’
SQSTM1 primersTsingkeN/AForward: 5’-GCACCCCAATGTGATCTGC-3’
Reverse: 5’-CGCTACACAAGTCGTAGTCTGG-3’
TRIzol reagentVazymeR401-01qPCR reagent
β-actin primersTsingkeN/AForward: 5’-CATGTACGTTGCTATCCAGGC-3’
Reverse: 5’-CTCCTTAATGTCACGCACGAT-3’

Reprints and Permissions

Request permission to reuse the text or figures of this JoVE article

Request Permission

Tags

Medicinemitochondriamachine learningImmune infiltration

Related Articles