This protocol integrates multidimensional transcriptomic data with machine learning to identify aging- and mitochondria-related genes in dilated cardiomyopathy for biomarker discovery and molecular subtyping.
Research Article
This protocol integrates multidimensional transcriptomic data with machine learning to identify aging- and mitochondria-related genes in dilated cardiomyopathy for biomarker discovery and molecular subtyping.
Dilated cardiomyopathy (DCM) is characterized by left ventricular dilation and systolic dysfunction and is associated with mitochondrial dysfunction and immune-inflammatory activation. However, aging-related molecular signatures and mitochondrial regulatory pathways in DCM remain incompletely understood. This study analyzed six bulk transcriptomic datasets and one single-cell RNA sequencing dataset from the Gene Expression Omnibus database. After data normalization, batch correction, and cell-type annotation, aging- and mitochondria-related candidate genes were identified using differential expression analysis, weighted gene co-expression network analysis, and protein-protein interaction network construction. Core genes were further screened using least absolute shrinkage and selection operator regression, random forest, and support vector machine-recursive feature elimination. Immune-cell infiltration, cell-cell communication, and molecular subtyping analyses were performed to characterize the cardiac immune microenvironment in DCM. A total of 66 aging-related genes and 16 mitochondria-related genes were associated with DCM and were mainly enriched in hypoxia-inducible factor-1 signaling, oxidative phosphorylation, and nitric oxide synthase-related pathways. Machine learning and single-cell RNA sequencing analyses identified SERPINE1, TGFB2, CYBB, and TLR2 as core genes. CYBB and TLR2 were highly expressed in monocytes and macrophages, whereas SERPINE1 and TGFB2 were predominantly expressed in stromal cells. Immune landscape analysis showed increased pro-inflammatory macrophage activation and altered cell-cell communication in DCM samples. Based on core gene expression, DCM samples were divided into two molecular subtypes associated with vascular endothelial growth factor signaling and primary bile acid biosynthesis, respectively. This protocol provides an integrated framework for identifying candidate biomarkers and molecular subtypes in DCM.
Dilated cardiomyopathy (DCM) is a myocardial disorder characterized by left ventricular dilation and impaired systolic function. It is the third most common cause of heart failure and the leading indication for cardiac transplantation worldwide1. Population-based studies estimate a prevalence of approximately 1 in 250 adults, with a higher prevalence in males and a substantial proportion of cases attributable to monogenic variants2. These findings indicate that both genetic susceptibility and environmental factors contribute to the onset and progression of DCM.
The pathogenesis of DCM involves interconnected processes, including inflammatory activation, oxidative stress, cardiomyocyte apoptosis, and dysregulated profibrotic signaling. Inflammatory genetic polymorphisms, including tumor necrosis factor-α promoter variants, have been associated with susceptibility to viral DCM3. Increased oxidative stress has also been associated with cardiomyocyte death and left ventricular dysfunction in human DCM subtypes4. In addition, aberrant activation of Wnt/β-catenin and calcineurin/nuclear factor of activated T cells signaling promotes myocardial hypertrophy and interstitial fibrosis, thereby contributing to disease progression5,6. Mitochondrial dysfunction is another important component of DCM because cardiomyocytes have high energy demands. Disruption of mitochondrial biogenesis, calcium homeostasis, mitophagy, and mitochondrial DNA integrity can impair oxidative phosphorylation and contribute to progressive cardiac dysfunction7,8,9,10.
Despite these mechanistic findings, important knowledge gaps remain. In particular, the temporal and causal relationships between mitochondrial structural remodeling and bioenergetic dysfunction during DCM initiation and progression have not been fully defined11. Several therapeutic strategies have been investigated. Stem cell therapy has shown regenerative potential through paracrine, cytoprotective, and immunomodulatory effects, but optimization of cell sources, delivery routes, and post-transplantation survival remains necessary12. Gene therapy approaches, including adeno-associated virus-based delivery and clustered regularly interspaced short palindromic repeats-based genome editing, also offer potential precision treatment strategies. However, limitations related to cardiac tropism, vector immunogenicity, and long-term safety remain unresolved13.
Public transcriptomic datasets from repositories such as the Gene Expression Omnibus (GEO) are widely used for biomarker discovery in DCM. These resources provide access to multicenter clinical cohorts, support cost-effective and reproducible investigations, and can improve statistical power through cross-dataset integration14. Transcriptomic profiling also enables genome-wide candidate gene screening, molecular subtyping, and pathway-level analysis15. However, public datasets have inherent limitations, including technical batch effects, clinical and etiological heterogeneity, limited capacity for causal inference, and incomplete longitudinal or prognostic information16. Therefore, findings derived from public transcriptomic datasets are most suitable for hypothesis generation and candidate biomarker prioritization and require validation in independent cohorts and experimental models.
Many bioinformatic studies of DCM rely primarily on differential expression analysis, which may generate false-positive findings and does not fully characterize gene co-expression networks or cellular heterogeneity within cardiac tissue. To address these limitations, the present study employed an integrated analytical strategy combining complementary methods. Bulk transcriptomic analysis provides tissue-level expression profiles suitable for case-control comparisons. Weighted gene co-expression network analysis (WGCNA) identifies gene modules associated with phenotypic traits and enables prioritization of functionally related gene sets rather than individual differentially expressed genes. Protein-protein interaction (PPI) network analysis identifies highly connected genes on the basis of network topology. Three machine learning algorithms—least absolute shrinkage and selection operator regression, random forest, and support vector machine-recursive feature elimination—were used to identify candidate biomarkers across the integrated datasets¹⁷. Single-cell RNA sequencing (scRNA-seq) was then used to examine cell-type-specific expression patterns and intercellular communication networks18.
Although mitochondrial dysfunction and aging-related molecular changes have each been investigated in DCM, their combined associations with disease-related transcriptional changes remain underexplored. The present study integrated multiple bulk transcriptomic and scRNA-seq datasets to identify aging- and mitochondria-related hub genes in DCM, characterize the cardiac immune microenvironment, and examine molecular subtypes based on the identified genes. This integrated approach was used to prioritize candidate biomarkers and provide a basis for subsequent mechanistic and validation studies.
All animal procedures were reviewed and approved by the Laboratory Animal Ethics Committee of the Second Affiliated Hospital of Henan University of Chinese Medicine (Approval No. HNSZYYYJS2023011150). All procedures were conducted in accordance with the Guidelines for Ethical Review of Laboratory Animal Welfare (GB/T 35892-2018) and the 3R principles of Replacement, Reduction, and Refinement. The reagents, databases, software, and equipment used in this study are listed in the Table of Materials.
1. Data resources and experimental materials
Male SPF-grade CTNTR141W transgenic mice with a spontaneous dilated cardiomyopathy (DCM) phenotype and a body weight of 25 ± 2 g were used as the model group. Age-matched male SPF-grade C57BL/6J mice with a body weight of 25 ± 2 g were used as the control group. Each group included 12 mice. All animals were obtained from institutions holding valid laboratory animal production licenses and were housed in an SPF-grade barrier environment at 22 ± 2 °C and 40%–60% relative humidity under a 12 h light/dark cycle, with free access to sterilized food and water. After 1 week of acclimatization, all mice were maintained under the same conditions for an additional 4 weeks before cardiac function assessment and sample collection. All mice were 6–8 weeks old at the start of the experiment. Mice were deeply anesthetized and euthanized by cervical dislocation.
Seven public transcriptomic datasets of left ventricular myocardial tissue from patients with DCM were retrieved from the Gene Expression Omnibus (GEO)19 database. These datasets included six bulk transcriptomic datasets and one single-cell RNA sequencing (scRNA-seq) dataset, GSE145154. Both CD45-positive and CD45-negative fractions were included in the analysis. Both CD45-positive and CD45-negative cell fractions were combined before clustering. Sample identity was used as the main batch variable for Harmony integration. Normal left ventricular and DCM left ventricular samples from GSE145154 were included, specifically GSM4307515, GSM4307516, GSM4307520, and GSM4307521. The datasets used in this study were GSE145154, GSE5406, GSE42955, GSE57338, GSE79962, GSE116250, and GSE141910. All samples other than DCM were excluded, and only control samples (Control group) and DCM samples (DCM group) were retained. No samples were removed after quality control. Sample information of included GEO datasets is summarized as follows: GSE5406 contained 102 samples (16 control and 86 DCM samples); GSE42955 contained 17 samples (5 control and 12 DCM samples); GSE57338 contained 231 samples (136 control and 95 DCM samples); GSE79962 contained 20 samples (11 control and 9 DCM samples); GSE116250 contained 51 samples (14 control and 37 DCM samples); and GSE141910 contained 322 samples (161 control and 161 DCM samples).
2. Bulk transcriptome data preprocessing
Raw expression matrices and clinical annotation files for the six bulk datasets were downloaded using the GEOquery package20. Raw CEL files were retrieved for the Affymetrix microarray datasets, and raw count matrices were retrieved for the RNA-seq datasets. Background correction, quantile normalization, and expression calculation for the microarray data were performed using the robust multi-array average algorithm implemented in the affy package21.
RNA-seq count data were normalized using the trimmed mean of M-values method in the edgeR package22 and were converted to log₂-transformed counts per million values. Probe identifiers were converted to official gene symbols using platform-specific annotation files. When multiple probes mapped to the same gene, the mean expression value was calculated.
Technical batch effects across datasets were removed using the ComBat algorithm in the sva package23. Dataset source and detection platform were specified as batch factors. Principal component analysis was performed before and after batch correction to evaluate the effectiveness of batch-effect removal.
3. Single-cell transcriptome data preprocessing and cell annotation
The gene expression matrix from GSE145154 was imported into Seurat to construct a Seurat object using Seurat version 524. Low-quality cells were excluded using the following thresholds: 200–6,000 detected genes per cell, total unique molecular identifier count greater than 500, and mitochondrial gene percentage below 25%. Cells outside these quality-control thresholds were excluded as low-quality or ruptured cells. We excluded low-quality cells solely using the quality control thresholds described above.
Log normalization was performed using the NormalizeData function with a scale factor of 10,000. The top 3,000 highly variable genes were selected using the FindVariableFeatures function with the vst method. The data were scaled using ScaleData, followed by principal component analysis for linear dimensionality reduction.
Batch effects were corrected using the Harmony algorithm25 through the RunHarmony function, with sample identity specified as the grouping variable. The first 15 principal components were used to cluster cells using the FindNeighbors and FindClusters functions. Clustering was performed using the Leiden algorithm at a resolution of 0.15. Nonlinear dimensionality reduction and visualization were performed using uniform manifold approximation and projection.
Cell types were annotated using canonical marker genes together with automated annotation using the SingleR package26. The marker genes were as follows: B cells, IGKC, MS4A1, and CD79A; cardiomyocytes, TNNI3, MYL2, and ACTC1; endothelial cells, VWF, PECAM1, and EGFL7; macrophages, C1QC, C1QB, and C1QA; monocytes, S100A8, S100A9, and G0S2; natural killer cells, NKG7, GNLY, and CCL5; smooth muscle cells, MYL9, TAGLN, and ACTA2; stromal cells, FBLN1, LUM, and DCN; and T cells, CD3E, CD3G, and CD3D.
4. Differential expression analysis and gene set enrichment scoring
A linear model was constructed using the limma package27 to compare gene expression between the DCM and healthy control groups. Genes with P value < 0.05 and an absolute fold change greater than 1.5, corresponding to an absolute log₂ fold change greater than 0.58, were defined as significantly differentially expressed.
Single-sample gene set enrichment analysis was performed to calculate enrichment scores for the aging-related and mitochondria-related gene sets in each sample28. Differences in enrichment scores between the DCM and healthy control groups were evaluated using the Wilcoxon rank-sum test, with P value < 0.05 considered statistically significant.
At the single-cell level, aging-related and mitochondrial module scores were calculated using the AddModuleScore function in Seurat. Differences in module scores between groups were assessed using the Wilcoxon rank-sum test.
Aging-related gene signatures were retrieved from the CellAge database (https://genomics.senescence.info/cells/), and mitochondria-related gene sets were obtained from GeneCards (https://www.genecards.org/). The complete gene lists used for scoring are provided in Supplementary File 1.
5. Weighted gene co-expression network construction
The top 5000 protein-coding genes with the highest expression variance in bulk transcriptomic data were retained for network construction. The pickSoftThreshold function was applied to calculate the scale-free topology fit index under multiple soft-thresholding powers. The optimal threshold was determined as the minimum power yielding a scale-free network with an R2 value above 0.9. Accordingly, a soft-thresholding power of β = 5 was adopted for subsequent network analysis.
A signed weighted co-expression network was constructed using the blockwiseModules function with a minimum module size of 30. Pearson correlation coefficients were calculated between each module eigengene and the aging-related or mitochondrial enrichment score. Modules with an absolute correlation coefficient greater than 0.4 and P < 0.001 were considered significantly associated modules.
Genes within significantly associated modules were intersected with differentially expressed genes to identify DCM-associated aging candidate genes and DCM-associated mitochondria candidate genes.
6. Functional enrichment analysis
Functional enrichment analyses, including Gene Ontology (GO) and Kyoto Encyclopedia of Genes and Genomes (KEGG) pathway analyses, were conducted on candidate genes using the clusterProfiler package29. GO enrichment covered three standard categories: biological process, cellular component, and molecular function.
All analyses were performed with human species annotation, false discovery rate (FDR) for P‑value correction, and a q‑value threshold of 0.05. Gene sets were restricted to a size range of 10–500 genes, and terms with an FDR < 0.05 were defined as statistically significant. Finally, GO enrichment results were visualized via grouped bar charts, while KEGG enrichment outcomes were displayed using bubble plots.
7. PPI network construction and hub gene screening
Candidate genes were submitted to the STRING version 11.5 database30, with the organism set to Homo sapiens and the interaction confidence threshold set to a combined score greater than 0.7. Disconnected nodes were hidden, and the interaction data were exported in tab-separated values format.
The interaction data were imported into Cytoscape version 3.9.1 for visualization31. Node topological scores were calculated using the CytoHubba plugin32 with three algorithms: Degree, maximum neighborhood component, and maximal clique centrality.
Core functional modules within the network were identified using the MCODE plugin33 with the following default parameters: degree cutoff, 2; k-core, 2; node score cutoff, 0.2; and maximum depth, 100. Genes ranked among the top 10 by all three topological algorithms were intersected with genes in the MCODE core subnetwork to identify the final protein-protein interaction hub genes.
8. Machine learning-based core gene selection and diagnostic model building
To ensure reproducibility and balanced representation, the integrated bulk transcriptomic dataset was randomly split into training and validation sets at a 7:3 ratio using a fixed random seed (seed = 123456). This split was stratified by disease group (DCM vs. control) to maintain consistent class proportions across both sets. Prior to splitting, batch effects from different dataset sources were corrected using the sva package, and the integrated samples were treated as a unified cohort during the random allocation.
Three machine learning algorithms were applied to screen candidate genes. First, LASSO logistic regression was performed via the cv.glmnet function in the glmnet package34. A 5-fold cross-validation binary classification model was constructed, with the AUC adopted as the evaluation metric. Genes with non-zero coefficients at lambda.min were reserved as candidate genes.
Second, a random forest classification model with 500 decision trees was built using the randomForest package35. The number of variables sampled for each split was set to the square root of the total number of features. Gene importance was quantified based on the Gini coefficient, and the top 10 genes with the highest importance scores were retained.
Third, SVM-RFE analysis was implemented using the rfe function in the caret package36. Feature numbers were set ranging from 1–10, and 5-fold cross-validation was adopted for model training. The gene subset with the optimal cross-validation accuracy was ultimately selected.
Genes identified by all three algorithms were defined as the final core aging- and mitochondria-related genes in DCM. Diagnostic models were then constructed using 10 classification algorithms: decision tree, gradient boosting machine, boosted generalized linear model, k-nearest neighbors, logistic regression, neural network, partial least squares, random forest, support vector machine, and extreme gradient boosting.
Receiver operating characteristic curves were generated using the pROC package37. Area under the curve, accuracy, sensitivity, and specificity were calculated to evaluate diagnostic performance in the training and validation sets.
SHapley Additive exPlanations analysis was performed to calculate the contribution of each core gene to model predictions38. Summary plots and per-sample waterfall plots were generated. A final diagnostic model with an area under the curve greater than 0.8 in the validation set was considered to have good diagnostic performance.
9. Cell-cell communication inference
Cell-cell communication networks in the cardiac microenvironment were inferred using the CellChat package39. A CellChat object was constructed using the CellChatDB.human database. Differentially expressed ligands and receptors were identified using identifyOverExpressedGenes, and significant interaction pairs were filtered using identifyOverExpressedInteractions.
Communication probabilities between cell types were calculated using computeCommunProb. The global cell-type-level communication network was aggregated using aggregateNet. The number of interactions and the communication strength between each pair of cell types were quantified and visualized using heatmaps and bar charts.
10. Immune cell infiltration quantification
Enrichment scores for 28 immune cell types were calculated for each bulk sample using single-sample gene set enrichment analysis28 and an immune cell signature gene set40. The Wilcoxon rank-sum test was used to compare immune cell enrichment scores between the DCM and healthy control groups. P < 0.05 was considered statistically significant.
Pearson correlation analysis was conducted to evaluate the association between core gene expression levels and immune cell enrichment scores. All correlations with P < 0.05 were regarded as statistically significant.
11. Consensus clustering for molecular subtyping
Unsupervised consensus clustering of DCM samples was performed using core gene expression profiles via the ConsensusClusterPlus package41. The clustering parameters were set as a maximum cluster number of 6, 1000 resampling iterations, and a resampling proportion of 0.8. Partitioning around medoids with Euclidean distance was adopted for clustering, and a fixed random seed was used to ensure reproducibility.
The optimal subtype number was determined according to delta area plot and consensus cluster stability scores, with K = 2 finally identified. Principal component analysis was further performed to verify the distinct separation of the two molecular subtypes.
Gene set variation analysis42 was applied to calculate sample-specific KEGG pathway enrichment scores. The limma package27 was utilized to detect differential pathway activation between subtypes, and a P value of less than 0.05 was defined as statistically significant.
12. Echocardiographic assessment of cardiac function
Mice were anesthetized via intraperitoneal injection of 1% sodium pentobarbital (30 mg/kg) and fixed in a supine position on a thermostatic operating table. After chest hair removal, ultrasound coupling gel was evenly applied to the precordial area.
Two-dimensional-guided M-mode echocardiography was performed at the level of the left ventricular papillary muscles using a small-animal ultrasound system. Three consecutive stable cardiac cycles were captured to measure left ventricular end-diastolic diameter, end-systolic diameter, ejection fraction, and fractional shortening. All echocardiographic assessments were conducted blindly by a professional ultrasonographer.
Three mice were randomly selected from each group for echocardiographic examination, and these 6 animals in total were subsequently sacrificed for myocardial tissue collection and ELISA measurement. The remaining experimental animals underwent additional parallel laboratory assays, and their data were not included in the present study.
13. Myocardial tissue collection, protein extraction, and enzyme-linked immunosorbent assay
After echocardiographic assessment, mice were euthanized under deep anesthesia. Heart tissues were rapidly harvested via median thoracotomy, and left ventricular myocardium was dissected on ice. Isolated tissues were thoroughly rinsed with ice‑cold phosphate‑buffered saline to eliminate residual intracardiac blood. After blotting excess liquid with sterile filter paper, samples were immediately snap‑frozen in liquid nitrogen and stored at −80 °C for subsequent protein extraction, with repeated freeze–thaw cycles strictly avoided.
Frozen myocardial tissues were weighed and cut into approximately 1 mm3 fragments on ice. Tissues were lysed in ice‑cold RIPA lysis buffer containing protease and phosphatase inhibitors at a standardized ratio of 100 µL buffer per 10 mg tissue. Samples were fully homogenized mechanically on ice and incubated for 30 min to achieve complete cell lysis.
Lysates were centrifuged at 12,000 × g for 15 min at 4 °C. The resulting supernatants were collected in enzyme‑free tubes, and total protein concentration was quantified using a bicinchoninic acid protein assay kit following the manufacturer’s protocols. All samples were normalized to an identical protein concentration with lysis buffer.
Protein expression levels of the four hub genes in myocardial lysates were measured using the corresponding enzyme‑linked immunosorbent assay (ELISA) kits. Serially diluted standards and normalized tissue lysates were added in duplicate (100 µL per well) to pre‑coated microplates. Plates were incubated for 2 h at room temperature and thoroughly washed with the kit‑supplied wash buffer.
Each well was supplemented with enzyme‑conjugated antibody and incubated for 1 h at room temperature, followed by a thorough wash. Substrate chromogen solution was then added, and the plates were incubated for 20 min at room temperature in the dark. The color reaction was terminated with the stop solution, and absorbance values were measured at 450 nm (reference wavelength: 570 nm) using a full‑wavelength microplate reader.
14. Statistical analysis
All statistical analyses and data visualizations were performed using R version 4.2.3. For ELISA concentration measurements of each target gene (TGFB2, SERPINE1, CYBB, TLR2), the Shapiro-Wilk test was first applied to assess the normality of the data in the Control and DCM groups separately. Subsequently, an F-test was used to evaluate the homogeneity of variances between the two groups. The method for intergroup comparison was determined according to the results of the variance homogeneity test: if variances were homogeneous (P ≥ 0.05), an unpaired Student’s t-test was adopted to compare mean values between groups; if variances were heterogeneous (P < 0.05), the corrected Welch’s t-test was utilized for analysis. All tests were two-tailed, and the threshold for statistical significance was set at P < 0.05. Data were visualized as boxplots overlaid with jittered individual data points. The P values of all tests and the type of t-test implemented were annotated in detail on each plot.
Data preprocessing and differential expression analysis
All six bulk transcriptomic datasets underwent standardized preprocessing and batch-effect correction before downstream analysis. Microarray data were normalized using the robust multi-array average algorithm, whereas RNA-seq count data were normalized using the trimmed mean of M-values method. The ComBat algorithm was applied to remove technical batch effects associated with the dataset source and detection platform. Principal component analysis showed that samples clustered by dataset source before correction but were more uniformly distributed after correction, with no visible separation by batch.
Differential expression analysis between the dilated cardiomyopathy (DCM) and healthy control (HC) groups was performed using the limma package. The heatmap of the 20 most significantly differentially expressed genes showed separation of the expression profiles between the two groups (Figure 1A). A total of 1,473 differentially expressed genes were identified using thresholds of P value < 0.05 and |log₂ fold change| > 0.58. Of these, 819 genes were upregulated, and 654 were downregulated in DCM myocardial samples (Figure 1B).
Single-sample gene set enrichment analysis was then used to calculate enrichment scores for the aging-related and mitochondria-related gene sets in each sample. Both scores differed significantly between the DCM and HC groups (Figure 1C).
Weighted gene co-expression network analysis
Weighted gene co-expression network analysis was performed to identify gene modules associated with the aging-related and mitochondrial enrichment scores. The 5,000 protein-coding genes with the highest expression variance in the bulk dataset were used to construct the network. At a soft-thresholding power of β = 5, the scale-free topology fit index exceeded R2 = 0.9, meeting the scale-free network criterion (Figure 1D).
Hierarchical clustering and module merging identified three gene modules. All three modules were significantly correlated with the aging-related score. The turquoise module showed the strongest correlation with the aging-related score (r = 0.69, P < 0.001). For the mitochondrial score, the blue and grey modules were significantly correlated, with the blue module showing the strongest association (r = 0.56, P < 0.001; Figure 1E). The turquoise module was therefore selected for aging-related gene screening, and the blue module was selected for mitochondria-related gene screening.

Figure 1: Differential expression analysis and weighted gene co-expression network construction. (A) Heatmap of the 20 most significantly differentially expressed genes between the dilated cardiomyopathy (DCM) and healthy control (HC) groups. (B) Volcano plot of all differentially expressed genes. Red indicates upregulated genes, green indicates downregulated genes, and gray indicates nonsignificant genes. The thresholds were P value < 0.05 and |log₂ fold change| > 0.58. (C) Box plots of single-sample gene set enrichment analysis scores for the aging-related and mitochondria-related gene sets. (D) Soft-threshold selection for weighted gene co-expression network analysis, showing the scale-free topology fit index and mean connectivity across different soft-thresholding powers. (E) Heatmap of correlations between module eigengenes and the aging-related and mitochondrial scores. Please click here to view a larger version of this figure.
Identification of aging- and mitochondria-related candidate genes
Candidate genes were identified by intersecting the differentially expressed genes, genes in the selected weighted gene co-expression network analysis modules, and the corresponding reference gene sets. This analysis identified 66 DCM-associated aging candidate genes (Figure 2A) and 16 DCM-associated mitochondria candidate genes (Figure 2B).
Gene Ontology enrichment analysis showed that the aging-related candidate genes were enriched in biological processes, including nitric oxide synthase biosynthesis and collagen-containing extracellular matrix organization (Figure 2C). The mitochondria-related candidate genes were enriched in terms associated with mitochondrial energy metabolism, including the mitochondrial inner membrane and respiratory chain complex (Figure 2D).
Kyoto Encyclopedia of Genes and Genomes analysis showed that the aging-related candidate genes were enriched in the hypoxia-inducible factor-1, phosphoinositide 3-kinase-protein kinase B, and advanced glycation end product-receptor for advanced glycation end product signaling pathways (Figure 2E). The mitochondria-related candidate genes were enriched in pathways including oxidative phosphorylation (Figure 2F).
The differential expression patterns of the 66 aging-related candidate genes between the DCM and HC groups were visualized using an expression heatmap (Figure 2G). The expression patterns of the 16 mitochondria-related candidate genes were visualized using box plots (Figure 2H).

Figure 2: Screening and functional enrichment of candidate genes. (A) Venn diagram showing the intersection of differentially expressed genes, weighted gene co-expression network analysis module genes, and the aging-related reference gene set. (B) Venn diagram showing the intersection of differentially expressed genes, weighted gene co-expression network analysis module genes, and the mitochondria-related reference gene set. (C) Gene Ontology enrichment analysis of aging-related candidate genes. (D) Gene Ontology enrichment analysis of mitochondria-related candidate genes. (E) Kyoto Encyclopedia of Genes and Genomes pathway enrichment analysis of aging-related candidate genes. (F) Kyoto Encyclopedia of Genes and Genomes pathway enrichment analysis of mitochondria-related candidate genes. (G) Expression heatmap of the 66 aging-related candidate genes in the DCM and HC groups. (H) Expression box plots of the 16 mitochondria-related candidate genes in the DCM and HC groups. Please click here to view a larger version of this figure.
Cell-type annotation of the single-cell RNA sequencing dataset
The GSE145154 single-cell RNA sequencing dataset was used for single-cell resolution validation. Following quality-control filtering, log normalization, and Harmony batch correction, cells from different samples were distributed across the uniform manifold approximation and projection space without visible sample-specific separation. Using the first 15 principal components and a clustering resolution of 0.15, the cells were divided into 9 clusters (Figure 3A).
Canonical marker genes and automated annotation using SingleR identified 9 major cell types: macrophages, natural killer cells, T cells, B cells, endothelial cells, smooth muscle cells, monocytes, stromal cells, and cardiomyocytes (Figure 3B). The expression patterns of cell-type-specific marker genes supported these annotations (Figure 3C).
Mitochondrial module scores were calculated for each cell using the AddModuleScore function and differed significantly between the dilated cardiomyopathy and healthy control groups (P < 2.22 × 10⁻16; Figure 3D). Aging-related module scores also differed significantly between the two groups (P < 2.22 × 10⁻16; Figure 3E). Projection of the mitochondrial scores onto the uniform manifold approximation and projection space showed that high scores were mainly observed in cardiomyocytes (Figure 3F). In contrast, high aging-related scores were predominantly observed in macrophages (Figure 3G).

Figure 3: Single-cell transcriptome annotation and module-score analysis. (A) Uniform manifold approximation and projection plot of cell clusters generated using the first 15 principal components and a clustering resolution of 0.15. (B) Uniform manifold approximation and projection plot of the annotated cell types. (C) Bubble plot showing the expression of canonical marker genes across cell types. (D) Violin plot of mitochondrial module scores in the DCM and HC groups. (E) Violin plot of aging-related module scores in the DCM and HC groups. (F) Uniform manifold approximation and projection plot showing the distribution of mitochondrial module scores across cells. (G) Uniform manifold approximation and projection plot showing the distribution of aging-related module scores across cells. Please click here to view a larger version of this figure.
Protein-protein interaction network construction
The combined set of 66 aging-related and 16 mitochondria-related candidate genes was submitted to the STRING version 11.5 database to construct a protein-protein interaction network using a high-confidence threshold of combined score > 0.7. The network was imported into Cytoscape for visualization and topological analysis (Figure 4A).
Degree, maximal clique centrality, maximum neighborhood component, and MCODE analyses were used to identify highly connected nodes and core subnetworks. The subnetworks identified by these methods are shown in Figure 4B–E.
The top 10 genes ranked by Degree, maximal clique centrality, and maximum neighborhood component were intersected with genes in the MCODE core subnetwork. This analysis identified 10 candidate genes: TGFB2, TLR2, SERPINE1, CYBB, KDR, TLR4, HIF1A, CCL2, MMP9, and CXCR2.

Figure 4: Protein-protein interaction network construction and hub gene screening.
(A) Overall protein-protein interaction network of the candidate genes. (B) Core subnetwork identified using MCODE. (C) Core subnetwork identified using maximal clique centrality. (D) Core subnetwork identified using maximum neighborhood component. (E) Core subnetwork identified using Degree. Please click here to view a larger version of this figure.
Machine learning-based screening of hub genes
Three machine learning algorithms—least absolute shrinkage and selection operator logistic regression, random forest, and support vector machine-recursive feature elimination—were applied to screen hub genes from the 10 protein-protein interaction candidates. All analyses were performed using a fixed random seed (set.seed(12345)) and 5-fold cross-validation. In the LASSO model, genes with nonzero coefficients at the optimal lambda (lambda.min) were retained as candidates (Figure 5A).
In the support vector machine-recursive feature elimination model, the highest cross-validation accuracy of 0.859 was achieved when 10 features were included (Figure 5B), with a corresponding minimum error rate of 0.141 (Figure 5C). The random forest model with 500 decision trees showed stable convergence of the out-of-bag error rate (Figure 5D). Gene importance ranking based on the Gini coefficient placed TGFB2, TLR2, SERPINE1, and CYBB among the highest-ranked genes (Figure 5E). The intersection of genes selected by all three algorithms yielded four final hub genes: CYBB, SERPINE1, TGFB2, and TLR2 (Figure 5F).
SHapley Additive exPlanations analysis was performed to evaluate the contribution of each hub gene to the model's predictions. TGFB2 had the highest mean absolute SHapley Additive exPlanations value at 0.249, followed by SERPINE1 at 0.103, CYBB at 0.083, and TLR2 at 0.078 (Figure 6A). The summary plot showed the distribution and direction of gene contributions across samples (Figure 6B). Dependence plots illustrated the relationship between individual gene values and model contributions (Figure 6C), whereas per-sample waterfall plots showed the contribution of each gene to individual predictions (Figure 6D).
Diagnostic classification models based on the four hub genes were subsequently constructed using 10 classification algorithms. In the training set, most algorithms achieved area under the curve values above 0.85 (Figure 6E). In the internal validation set, most algorithms achieved area under the curve values above 0.78 (Figure 6F).

Figure 5: Machine learning-based screening of hub genes. (A) Least absolute shrinkage and selection operator regression coefficient trajectory and optimal lambda selection. (B) Cross-validation accuracy curve for the support vector machine-recursive feature elimination model. (C) Cross-validation error curve for the support vector machine-recursive feature elimination model. (D) Out-of-bag error rate curve for the random forest model. (E) Gene importance ranking based on the Gini coefficient in the random forest model. (F) Venn diagram showing the hub genes identified by the three machine learning algorithms. Please click here to view a larger version of this figure.

Figure 6: Diagnostic model evaluation and SHapley Additive exPlanations analysis. (A) Mean absolute SHapley Additive exPlanations values for the four hub genes. (B) SHapley Additive exPlanations summary plot showing the distribution and direction of gene contributions. (C) SHapley Additive exPlanations dependence plots for each hub gene. (D) SHapley Additive exPlanations waterfall plot for a representative sample. (E) Diagnostic performance heatmap of 10 classification algorithms in the training set. (F) Diagnostic performance heatmap of 10 classification algorithms in the validation set. Please click here to view a larger version of this figure.
Single-cell validation and cell-cell communication analysis
The expression patterns of the four hub genes were evaluated at the single-cell level. Cell-type distribution analysis showed that CYBB and TLR2 were highly expressed in monocytes and macrophages, whereas SERPINE1 and TGFB2 were predominantly expressed in stromal cells (Figure 7A).
Violin plots showed that CYBB expression differed significantly between the dilated cardiomyopathy and healthy control groups (Figure 7B). SERPINE1 (Figure 7C), TGFB2 (Figure 7D), and TLR2 (Figure 7E) also differed significantly between groups. All four genes were significantly upregulated in the dilated cardiomyopathy group relative to healthy controls, with P < 0.0001 for each comparison.
Cell-cell communication networks in the cardiac microenvironment were inferred using CellChat and a ligand-receptor database. The number and overall strength of cell-cell interactions differed between the dilated cardiomyopathy and control groups (Figure 7F). Differential communication strength between cell types was also observed (Figure 7G). Monocytes, macrophages, cardiomyocytes, and stromal cells were major participants in the communication network.

Figure 7: Single-cell validation of hub genes and cell-cell communication analysis. (A) Bubble plot showing the expression of the four hub genes across cell types. (B) Violin plot of CYBB expression in the DCM and HC groups. (C) Violin plot of SERPINE1 expression in the DCM and HC groups. (D) Violin plot of TGFB2 expression in the DCM and HC groups. (E) Violin plot of TLR2 expression in the DCM and HC groups. (F) Bar plot showing the number and overall strength of cell-cell interactions. (G) Heatmap showing the differential strength of cell-cell communication between groups. Please click here to view a larger version of this figure.
Immune cell infiltration analysis
Enrichment scores for 28 immune cell subsets were calculated for each bulk sample using single-sample gene set enrichment analysis. The abundance of most immune cell types differed significantly between the dilated cardiomyopathy and healthy control groups (Figure 8A).
Pearson correlation analysis was then performed to assess the relationship between hub gene expression and immune cell enrichment scores. CYBB expression was significantly correlated with the abundance of multiple immune cell types (Figure 8B). Similar correlations were observed for SERPINE1 (Figure 8C), TGFB2 (Figure 8D), and TLR2 (Figure 8E). CYBB, SERPINE1, and TLR2 showed positive correlations with several innate immune cell populations, including monocytes and macrophages.

Figure 8: Immune cell infiltration and correlation analysis. (A) Box plots of enrichment scores for 28 immune cell types in the DCM and HC groups. (B) Lollipop plot showing correlations between CYBB expression and immune cell abundance. (C) Lollipop plot showing correlations between SERPINE1 expression and immune cell abundance. (D) Lollipop plot showing correlations between TGFB2 expression and immune cell abundance. (E) Lollipop plot showing correlations between TLR2 expression and immune cell abundance. Please click here to view a larger version of this figure.
Molecular subtyping of dilated cardiomyopathy
Unsupervised consensus clustering was performed on dilated cardiomyopathy samples based on the expression profiles of the four hub genes. The consensus clustering matrix supported separation at K = 2 (Figure 9A). The delta area plot further supported K = 2 as the optimal cluster number, dividing the samples into two molecular subtypes, C1 and C2 (Figure 9B).
The expression levels of CYBB, SERPINE1, and TLR2 differed significantly between the two subtypes (Figure 9C). The abundance of multiple immune cell subsets also differed between subtypes (Figure 9D). Gene set variation analysis showed relative activation of the vascular endothelial growth factor signaling pathway in the C1 subtype, whereas primary bile acid biosynthesis and glycosphingolipid biosynthesis were enriched in the C2 subtype (Figure 9E). Principal component analysis showed separation between samples assigned to the two subtypes (Figure 9F).

Figure 9: Consensus clustering for molecular subtyping of dilated cardiomyopathy. (A) Consensus clustering matrix at K = 2. (B) Delta area plot used to determine the optimal number of clusters. (C) Box plots of hub gene expression in the two molecular subtypes. (D) Box plots of immune cell abundance in the two molecular subtypes. (E) Heatmap of differentially enriched Kyoto Encyclopedia of Genes and Genomes pathways between the two molecular subtypes. (F) Principal component analysis plot showing separation between the two molecular subtypes. Please click here to view a larger version of this figure.
In vivo validation in the dilated cardiomyopathy mouse model
CTNTR141W transgenic mice with a spontaneous dilated cardiomyopathy phenotype were used for in vivo validation. Compared with age-matched wild-type C57BL/6J control mice, the transgenic mice showed significantly increased left ventricular end-diastolic diameter and decreased left ventricular ejection fraction, consistent with ventricular dilation and systolic dysfunction (Figure 10A).
Total protein was extracted from left ventricular myocardial tissue, and the concentrations of the four hub gene-encoded proteins were measured by enzyme-linked immunosorbent assay (ELISA) after bicinchoninic acid-based total protein normalization. All assays were performed in duplicate. Standard curves had correlation coefficients (R2) ≥ 0.99, and the coefficients of variation between duplicate wells were below 10%. Statistical significance between the Control and DCM groups was evaluated using either Student's t-test or Welch's t-test, depending on the equality of variances assessed by the F-test (normality was confirmed by the Shapiro-Wilk test). Myocardial protein levels corresponding to all four hub genes were significantly increased in the dilated cardiomyopathy (DCM) mice compared with controls (Figure 10B). For ELISA validation, 3 biological replicates (individual mice) were included in each group. For ELISA validation, three independent biological replicates were included per group. These findings should be considered preliminary and require confirmation in a larger cohort.

Figure 10: In vivo validation in the CTNTR141W transgenic dilated cardiomyopathy mouse model. (A) Representative M-mode echocardiographic images of CTNTR141W transgenic DCM mice and wild-type control mice. (B) ELISA quantification of four hub gene-derived proteins in mouse left ventricular myocardial tissues. Box-and-whisker plots display protein concentrations for Control and DCM groups (n = 3 biological replicates per group). For each box plot: the solid horizontal line inside the box denotes the median value; the upper and lower boundaries of the box represent the 75th and 25th percentiles (interquartile range, IQR); the upper and lower whiskers extend to the maximum and minimum non-outlier data points within 1.5 × IQR; individual black solid dots correspond to independent biological replicates from single animals. The y-axis indicates absolute protein concentration: pg/mL for TGFB2 and CYBB, ng/mL for TLR2 and SERPINE1. Statistical comparisons between two groups were performed using Student’s t-test (equal variance) or Welch’s t-test (unequal variance), with normality verified by the Shapiro-Wilk test and homogeneity of variance assessed by the F-test. Limitation statement: The ELISA results obtained from n = 3 replicates are preliminary exploratory findings, and future validation with enlarged sample size is warranted. Please click here to view a larger version of this figure.
Data Availability:
The six bulk transcriptomic datasets and one single-cell RNA sequencing dataset analyzed in this study are publicly available in the Gene Expression Omnibus database under accession numbers GSE5406, GSE42955, GSE57338, GSE79962, GSE116250, GSE141910, and GSE145154. The single-cell analysis included samples GSM4307515, GSM4307516, GSM4307520, and GSM4307521 from GSE145154. All other data generated or analyzed during this study, along with the computational code, are included in this published article and its supplementary information files. Specifically, Supplementary File 1 contains the complete aging and mitochondria-related gene lists, the 28-cell immune signature, custom analytical scripts, normalized transcriptomic data matrices, source data for ELISA assays, and the raw source data underlying all manuscript figures.
Supplementary File 1: Aging- and Mitochondria-Related Gene Sets, Immune Signatures, Analysis Scripts, Normalized Transcriptomic Data, and Figure Source Data. Please click here to download this file.
The integrated multilayered workflow combined bulk transcriptomic meta-analysis, weighted gene co-expression network construction, ensemble machine learning, single-cell transcriptomic validation, and in vivo animal-model verification. Four aging- and mitochondria-related hub genes CYBB, SERPINE1, TGFB2, and TLR2 were identified as candidate diagnostic biomarkers for dilated cardiomyopathy (DCM). Integration of six independent left ventricular transcriptomic datasets from the Gene Expression Omnibus repository, including microarray and RNA-sequencing platforms, reduced single-dataset bias and increased the statistical basis of the analysis43,44,45. Weighted gene co-expression network analysis combined with predefined aging-related and mitochondrial gene sets enabled identification of trait-associated functional modules rather than reliance on differential expression analysis alone46. Ensemble machine learning reduced the algorithm-specific bias associated with individual feature-selection methods47,48, while SHapley Additive exPlanations analysis quantified the contribution of each hub gene to model predictions49. Validation across bulk myocardial transcriptomes, single-cell transcriptomes, and a transgenic mouse model further characterized the cellular distribution and myocardial protein levels of the selected genes50.
Batch correction was a critical step in the integrated analysis because residual dataset-specific variation could affect differential expression analysis and module-trait associations. Dataset source and detection platform were therefore included as batch factors in the ComBat model. Residual dataset-dependent clustering in principal component analysis plots would indicate incomplete correction and potential systematic bias44. The soft-thresholding power was also important for weighted gene co-expression network construction. The minimum value that produced a scale-free topology fit index of R2 > 0.9 was selected, yielding β = 5. A lower value may produce fragmented or functionally uninformative modules, whereas a higher value may weaken gene connectivity and reduce the statistical power of module-trait correlation analysis46. Single-cell quality-control thresholds were adapted to cardiac tissue because cardiomyocytes have high metabolic activity. A stringent filtering strategy with a mitochondrial gene percentage cutoff below 25% and a detected-gene range of 200–6,000 was implemented to remove ruptured and low-quality cells while retaining cardiomyocytes50. A fixed random seed, set. seed(12345), was used for dataset splitting, model training, and cross-validation to reduce variation across repeated machine learning analyses47.
Genotype confirmation, standardized housing, and consistent echocardiographic measurement were important for maintaining phenotypic stability in the animal experiments. CTNTR141W transgenic mice develop left ventricular dilation and systolic dysfunction after the specified acclimatization and feeding period51. Genotype verification before grouping is necessary to exclude nontransgenic animals and prevent phenotype misclassification. Echocardiographic measurements should be acquired consistently at the level of the left ventricular papillary muscles, with measurements averaged across three consecutive stable cardiac cycles. Variation in imaging position or anesthesia depth may increase variability in left ventricular ejection fraction measurements51. Enzyme-linked immunosorbent assay quality was evaluated using standard-curve correlation coefficients, R2 ≥ 0.99, and coefficients of variation < 10% between duplicate wells. Poor standard-curve linearity or inconsistent duplicate measurements may introduce systematic error into protein concentration estimates.
Several analytical problems may arise during the implementation of the workflow. Persistent batch separation after ComBat correction may reflect collinearity between batch variables and clinical factors, insufficient filtering of low-expression genes, or unmodeled technical variation. Clinical covariates such as age and sex may be included as protected variables when available, and genes with zero expression in more than 70% of samples may be removed to reduce noise44. Supplementary correction with removeBatchEffect may be considered when residual separation remains. An unexpectedly high or low number of differentially expressed genes may require evaluation of sample heterogeneity, normalization, outliers, and threshold selection27. Low module-trait correlations may be addressed by reassessing the variance threshold, soft-thresholding power, and extreme trait values. Expansion from the 5,000 to the 7,500 most variable genes or replacement of single-sample gene set enrichment analysis with gene set variation analysis may improve module detection46. Excessive isolated nodes in the protein-protein interaction network may require adjustment of the STRING confidence threshold or expansion of the candidate gene set30. Poor machine learning performance may reflect distributional differences between the training and validation sets, feature redundancy, or group imbalance. Stratified sampling, reduction of redundant features, or oversampling of the minority class may reduce these effects47. Ambiguous single-cell clustering may require reassessment of Harmony correction, principal component selection, and marker gene annotation50.
Several limitations should be considered. The transcriptomic datasets were obtained retrospectively from public repositories, and the original study designs and clinical confounders could not be controlled. Clinical annotations were incomplete across datasets, and most datasets lacked detailed information on etiology, medication history, patient age, and long-term outcomes. These limitations prevented evaluation of associations between the selected genes and prognosis, treatment response, or chronological aging52. Residual technical variation may also remain despite batch correction. The analysis was based primarily on messenger RNA expression and did not include integrated epigenomic, proteomic, or metabolomic data. Consequently, protein activity, post-translational regulation, and upstream mechanisms could not be determined. Only one single-cell dataset was included, limiting evaluation of cellular heterogeneity across DCM etiologies50. The CTNTR141W transgenic model primarily represents hereditary DCM associated with a cardiac troponin T mutation and may not reproduce idiopathic, viral, or ischemic forms of the disease51. Species differences between mice and humans also limit direct clinical translation. Protein-level validation was restricted to myocardial tissue from mice, and large clinical cohorts and comparisons with established biomarkers were not performed. The four hub genes are not specific to DCM and may also be altered in other cardiovascular or inflammatory conditions. In addition, candidate selection was based on predefined aging-related and mitochondrial gene sets. This hypothesis-driven strategy may exclude genes outside the selected reference sets, while intersection across three machine learning algorithms may omit genes identified by only one method48.
The analytical framework may support future molecular subtyping, biomarker validation, and multi-omics studies in DCM. The four-gene panel may be evaluated in independent peripheral blood or myocardial cohorts before assessment as a diagnostic or subtyping tool. The C1 and C2 subtypes showed different immune and metabolic pathway profiles, providing a basis for subsequent validation of subtype-specific biological features15. The selected genes may also be examined in molecular docking, cellular, and functional studies. TLR2 and CYBB are associated with inflammatory signaling and reactive oxygen species production, whereas TGFB2 and SERPINE1 are associated with fibrosis and cardiac remodeling53. Integration with proteomic, metabolomic, epigenomic, genome-wide association, and Mendelian randomization data may help evaluate regulatory relationships and potential causal associations54. The workflow may also be adapted to transcriptomic studies of hypertrophic cardiomyopathy, ischemic cardiomyopathy, and heart failure by replacing the disease-specific datasets and reference gene sets45. Future incorporation of single-cell assays for transposase-accessible chromatin sequencing and spatial transcriptomics may provide additional information on cellular regulation and spatial expression. The observed enrichment of aging-related signatures in macrophages and mitochondrial signatures in cardiomyocytes was consistent with previous reports on inflammatory and mitochondrial processes in cardiac disease55,56,57.
This study has several limitations that should be acknowledged. Notably, the commercial ELISA kits utilized for protein quantification were officially validated for the detection of target proteins in serum samples. In the present study, myocardial tissue lysates were adopted as the detection matrix instead of serum. Although consistent sample pretreatment and experimental operation procedures were strictly implemented throughout the assay to ensure the reliability and comparability of experimental data, the lack of official manufacturer validation for these ELISA kits in myocardial tissue lysate samples may lead to potential subtle deviations in protein quantitative results. Therefore, the application of serum-specific ELISA kits to myocardial tissue lysates constitutes a methodological limitation of this study.
The author declares no competing interests.
The publicly available data provided through the Gene Expression Omnibus database are gratefully acknowledged. The reviewers and editors are also acknowledged for their constructive comments on the manuscript. This work was supported by the Provincial Department-Level Scientific Research Project (Grant No. 2021JDZX2026), “Mechanism of Yiqi Huoxue Formula in Attenuating Atherosclerotic Vascular Remodeling via KLF2-Nrf2-Mediated Inflammatory Regulation.”
| Name | Company | Catalog Number | Comments |
|---|---|---|---|
| Aquasonic Clear Ultrasound Gel | Parker Laboratories, Inc. | Mar-34 | Used for small-animal echocardiographic imaging. |
| BCA Protein Assay Kit | Thermo Fisher Scientific | 23227 | Detection at 562 nm; range, 20–2,000 µg/mL; used for total protein quantification of mouse heart lysates. |
| CellAge database | Human Ageing Genomic Resources | https://genomics.senescence.info/cells/ | Source of aging-related gene signatures. |
| CytoHubba, Cytoscape plugin | Cytoscape App Store | Version 0.1 | Used for node-topology scoring in protein-protein interaction networks. |
| Cytoscape | Cytoscape Consortium | Version 3.9.1 | Used for visualization of protein-protein interaction networks. |
| Full-wavelength microplate reader | Thermo Fisher Scientific | Multiskan FC | Used to measure absorbance in ELISA assays. |
| Gene Expression Omnibus | National Center for Biotechnology Information | https://www.ncbi.nlm.nih.gov/geo/ | Public repository used to obtain transcriptomic datasets. |
| GeneCards | Weizmann Institute of Science | https://www.genecards.org/ | Source of mitochondria-related gene sets. |
| Halt Protease and Phosphatase Inhibitor Cocktail, 100×, EDTA-free | Thermo Fisher Scientific | 78441 | Stored at 4 °C; added to RIPA buffer at 10 µL/mL immediately before use. |
| Liquid nitrogen | Local laboratory gas supplier | Not applicable | Used for snap-freezing myocardial tissue. |
| Male SPF-grade C57BL/6J mice, 6–8 weeks old, 25 ± 2 g | Beijing Vital River Laboratory Animal Technology Co., Ltd. | Not applicable | Animal Production License No. SCXK (Jing) 2021-0006; used as normal controls. |
| Male SPF-grade CTNTR141W transgenic DCM mice, 6–8 weeks old, 25 ± 2 g | Institute of Laboratory Animal Science, Chinese Academy of Medical Sciences | Not applicable | Animal Production License No. SCXK (Jing) 2021-0065; used as the spontaneous DCM model. |
| MCODE, Cytoscape plugin | Cytoscape App Store | Version 2.0.2 | Used to identify core functional subnetworks in protein-protein interaction networks. |
| Mouse CYBB ELISA Kit | Biogradetech | A-QEK09250-96wells | Used in this study to measure CYBB in mouse myocardial tissue lysates. |
| Mouse PAI-1 ELISA Kit | EK-BIO | ML30970 | Used in this study to measure PAI-1, the SERPINE1-encoded protein, in mouse myocardial tissue lysates. |
| Mouse TGF-β2 ELISA Kit | ElaBoX | SEKM-0036 | Used in this study to measure TGF-β2 in mouse myocardial tissue lysates. |
| Mouse TLR-2 ELISA Kit | Solarbio | SEKM-0163 | Used in this study to measure TLR-2 in mouse myocardial tissue lysates. |
| Phosphate-buffered saline, pH 7.4, calcium- and magnesium-free | Biological Industries | 02-024-1ACS | Sterile 1× solution; stored at 4 °C; used for tissue washing and dilution. |
| R package: caret | CRAN | Version 6.0-94 | Used for support vector machine-recursive feature elimination. |
| R package: CellChat | CellChat developers | Version 1.6.1 | Used for cell-cell communication inference from single-cell RNA sequencing data. |
| R package: clusterProfiler | Bioconductor | Version 4.8.3 | Used for functional enrichment analysis. |
| R package: ConsensusClusterPlus | Bioconductor | Version 1.64.0 | Used for unsupervised consensus clustering. |
| R package: edgeR | Bioconductor | Version 3.42.4 | Used for RNA sequencing data normalization with the trimmed mean of M-values method. |
| R package: GEOquery | Bioconductor | Version 2.68.0 | Used to download data from the Gene Expression Omnibus. |
| R package: glmnet | CRAN | Version 4.1-8 | Used for least absolute shrinkage and selection operator logistic regression. |
| R package: limma | Bioconductor | Version 3.56.2 | Used for differential expression analysis and statistical modeling. |
| R package: pROC | CRAN | Version 1.18.5 | Used for receiver operating characteristic curve analysis. |
| R package: randomForest | CRAN | Version 4.7-1.2 | Used for random forest machine learning. |
| R package: Seurat | CRAN | Version 5.0.1 | Used for single-cell RNA sequencing data analysis. |
| R package: SingleR | Bioconductor | Version 2.2.0 | Used for automated cell-type annotation. |
| R package: sva | Bioconductor | Version 3.48.0 | Used for ComBat batch-effect correction. |
| Refrigerated centrifuge | Sigma-Aldrich | SIGMA 3-K | Used for centrifugation of myocardial tissue lysates. |
| RIPA Lysis and Extraction Buffer | Thermo Fisher Scientific | 89900 | Ready-to-use 1× solution; stored at 4 °C; supplemented with protease and phosphatase inhibitors before use. |
| Small-animal ultrasound imaging system | VINNO Technology Co., Ltd. | VINN06LAB | Used for echocardiographic assessment of cardiac function. |
| Sodium pentobarbital | Sinopharm Chemical Reagent Co. | 20040428 | Prepared as a 1% solution, 10 mg/mL, in sterile saline; used for intraperitoneal anesthesia at 30 mg/kg. |
| STRING database | STRING Consortium | Version 11.5 | Used for protein-protein interaction network construction. |
| TGrinder H24 Tissue Homogenizer | TIANGEN | OSE-TH-01 | Used to homogenize mouse myocardial tissue in RIPA buffer at 6.0 m/s for 30–60 s over 2–3 cycles. |
| Thermostatic animal platform/heated operating table for small animals | Shanghai Yuyan Scientific Instrument Co., Ltd. | T-30350 | Used to maintain mice at 37 °C during echocardiography; operating range, room temperature to 50 °C. |