Method Article

Bibliometrics and Bioinformatics Exploration of Research Hotspots and Key Targets in Comorbidity: Interstitial Lung Disease and Pulmonary Hypertension

DOI:

10.3791/71631

August 7th, 2026

In This Article

Summary

Loading...
$$\rightleftharpoonup{xx}$$ $$\longleftharp{xx}$$, $$\longrightharp{xx}$$,

This protocol uses interstitial lung disease and pulmonary hypertension as examples and integrates bibliometrics with bioinformatics to systematically and preliminarily explore their research trends, hotspots, key targets, and related pathways. This protocol could be replicated to analyze any disease pairs by modifying the prompts.

Abstract

Loading...
$$\rightleftharpoonup{xx}$$ $$\longleftharp{xx}$$, $$\longrightharp{xx}$$,

Interstitial lung disease (ILD) and pulmonary hypertension (PH) frequently coexist as comorbidities, leading to poor clinical outcomes and limited therapeutic options. Identifying important research hotspots and uncovering key comorbidity targets is essential for improving disease management and developing novel therapeutic strategies. This protocol retrieves relevant publications from the Web of Science Core Collection and Scopus databases, and uses CiteSpace and VOSviewer to map research hotspots and evolving trends. Subsequently, overlapping genetic targets associated with both diseases are identified from the GeneCards database. Protein-protein interaction networks are constructed using STRING to identify hub genes, and KEGG pathway enrichment analysis is performed using the R language to elucidate key signaling pathways. The results show that bibliometrics identifies the developmental trajectory of this interdisciplinary field from a macro perspective; bioinformatics analysis reveals FN1, IL6, and TNF as potential core comorbidity targets; and KEGG enrichment analysis indicates that immune-related pathways centered on PI3K-Akt and MAPK are deeply involved in the pathogenesis of the comorbidity. By integrating the two approaches, this protocol systematically delineates the development trajectory and research hotspots in ILD-PH comorbidity research, while preliminarily screening potential comorbidity targets and pathways. It provides direction and a computational basis for future experimental research and clinical validation. This integrated framework offers a reproducible methodology applicable to the study of other complex disease comorbidities.

Introduction

Loading...
$$\rightleftharpoonup{xx}$$ $$\longleftharp{xx}$$, $$\longrightharp{xx}$$,

Interstitial lung disease (ILD) refers to a group of diffuse lung diseases characterized pathologically by inflammation and fibrosis of the pulmonary interstitium. During its disease progression, it is often complicated by multiple comorbidities, among which pulmonary hypertension (PH) is one of the common comorbidities that significantly affects the prognosis1,2. Epidemiological data indicate that the incidence of PH in patients with ILD varies depending on the subtype and disease stage, ranging from 13% to 86%3,4,5. This comorbid condition not only significantly reduces patients' quality of life and increases their risk of mortality but also exacerbates the socioeconomic and healthcare burden6,7. In contrast to pulmonary arterial hypertension, the pathogenesis of interstitial lung disease-associated pulmonary hypertension is more complex, involving multiple pathological processes such as hypoxic pulmonary vasoconstriction, pro-fibrotic and inflammatory factor-mediated vascular remodeling, and extracellular matrix deposition8,9,10. At present, clinical diagnosis and treatment strategies for this comorbid condition remain relatively limited, and specific targeted drugs are lacking. Its underlying molecular mechanisms, key signaling pathways, and early biomarkers still need to be systematically elucidated11,12.

As an objective tool for detecting research focal points and evolutionary trajectories, bibliometrics has gained extensive application in recent years13,14. Bioinformatics analysis identifies key targets and core pathways by integrating previously discovered disease-associated targets15. For comorbidity research, integrating macroscopic research hotspots with molecular-level pathway and target information deepens understanding of comorbidity mechanisms and provides computational predictions and directional guidance for subsequent experimental studies. However, research on the comorbidity of interstitial lung disease and pulmonary hypertension still lacks a systematic integrated analysis that combines macroscopic knowledge mapping with microscopic molecular networks. Accordingly, this study used the Web of Science Core Collection (WoSCC) and Scopus databases to comprehensively outline the development trajectory and research hotspots of ILD-PH comorbidity research through bibliometric methods. At the same time, bioinformatics tools were employed to identify core genes and potential molecular pathways associated with the comorbidity, providing a candidate molecular basis and experimental design direction for subsequent experimental research and targeted exploration.

Protocol

Loading...
$$\rightleftharpoonup{xx}$$ $$\longleftharp{xx}$$, $$\longrightharp{xx}$$,

The study used publicly available secondary databases, including GeneCards, STRING, and KEGG. According to Article 32 of the Regulations on Human Genetic Resources of China, these databases do not involve direct collection of human genetic resources within China; therefore, no ethical approval or informed consent was required for this study. Detailed version information of the software and platforms used in this section is provided in the Table of Materials.

1. Literature collection and organization

  1. Access the MeSH database (Table of Materials).
  2. Search for the disease terms “Interstitial Lung Disease” and “Pulmonary Hypertension” separately to retrieve their associated subheadings.
  3. Develop a comprehensive search strategy based on the identified MeSH terms and subheadings.
  4. On March 9, 2026, open the WoSCC database (Table of Materials).
  5. Select Advanced Search. Enter the search strategy: TS=("Diffuse Parenchymal Lung Disease" OR "Diffuse Parenchymal Lung Diseases" OR "Interstitial Lung Disease" OR "Interstitial Lung Diseases") AND TS=("Pulmonary Hypertension"). Set the publication date range from January 1, 2016, to December 31, 2025, and click Search.
  6. Select Articles and Review Articles as the document types, and choose English as the language. Select Full Record and Cited References as the record content. Export the records in txt format and save as wos.txt, then convert to csv format and save as wos.csv.
  7. On March 9, 2026, open the Scopus database (Table of Materials).
  8. Set the search field to” Title”/ ”Abstract” / “Keywords”. Enter the search strategy: TITLE-ABS-KEY(("Diffuse Parenchymal Lung Disease" OR "Diffuse Parenchymal Lung Diseases" OR "Interstitial Lung Disease" OR "Interstitial Lung Diseases")) AND TITLE-ABS-KEY("Pulmonary Hypertension"). Select the year range from January 1, 2016, to December 31, 2025, choose Medicine as the subject area, and click Search (Supplementary Table 1).
  9. Click Export, select CSV format, export all retrieved records, and save as scopus.csv.
  10. Open wos.csv, copy the DOI column, and paste it into scopus.csv. Perform data comparison, use the Remove Duplicates function in Excel to mark duplicate records, delete the duplicates, and obtain scopus1.csv. For the very few records lacking a DOI, we manually cross-checked using a combination of "title + author + year" to ensure complete inclusion of the literature required for this study.
  11. Create new folders named input directory and output directory. Place scopus1.csv obtained after deduplication into the input directory. Open CiteSpace, click Data, select Scopus, specify the corresponding folders, and click the Scopus (csv) > WoS button to convert the Scopus data format, unifying the format of the two databases into txt format. Name the resulting file scopus1.txt.
  12. Merge the data from the two databases. Place scopus1.txt and wos.txt into CiteSpace for deduplication and merging to obtain uniformly formatted, deduplicated literature data: dataset.txt, for subsequent analysis (Supplementary Table 2).
  13. Check the obtained wos.txt and scopus.csv files to ensure that the deduplication process uses DOI as the matching criterion and that the dataset.txt file is successfully generated.

2. Annual publication output trend analysis

  1. In Microsoft Excel, count the number of publications by year to generate two columns of data (Year, Publications).
  2. Select the dataset, access Insert > Chart > Column Chart to generate a column chart of annual publication output.
  3. Designate the horizontal axis as Year and the vertical axis as Publications.
  4. Enable data labels to display the specific number of publications for each year.
  5. Export all figures in Tagged Image File Format (TIFF), ensuring a resolution of at least 300 dpi.

3. National/regional publication output and international collaboration analysis

  1. Launch VOSviewer, select Create > Create a map based on bibliographic data > Read data from bibliographic database files16.
  2. Import the plain text dataset generated in step 1.12.
  3. Select Co-authorship as the analysis type and Countries as the unit of analysis.
  4. Set the minimum number of documents for a country to 40, retaining the top 20 countries by publication output. To ensure that the collaboration network focuses on high-output countries while maintaining the clarity and readability of the map, making it easier for readers to intuitively grasp the core international research forces and collaboration patterns.
  5. In the Verify selected items interface, select all countries, copy the list to Excel, and export the number of publications and citation counts for each country.
  6. In the resulting Excel table, select the top 10 countries by publication output and generate a chart displaying their publication and citation counts.
  7. Select the "Finish" button to build the national collaboration network and save the network in GML format.
  8. Launch Scimago Graphica and import the GML file.
  9. Map the label field to Country and set the cluster field type to String.
  10. In the visualization configuration, place the publication output field (typically weight<Documents> or a similar name) into the Size box, set the cluster field as the Color box, and map the label field to the Label, Tooltip, and Unit boxes.
  11. Go to Marks > Map. Use Edges for line curvature adjustment, and Edge color for node color configuration.
  12. Export the final map in PNG format.
  13. Import the plain text dataset generated in step 1.12 into the Bibliometric Online Analysis Platform (https://bibliometric.com), click Upload citation data, and select Country relations to generate a national collaboration chord diagram for visualization.
  14. A key checkpoint is to adjust the minimum number of documents appropriately so that the resulting network is well-balanced and visually clear, while still reflecting the research objectives.

4. Journal publication output and citation analysis

  1. Import the same dataset into VOSviewer, select Citation as the analysis type, and select Sources as the unit of analysis.
  2. Set the minimum number of documents for a source to 5 and the minimum number of citations to 20.
  3. In the Verify selected items interface, copy the journal list to Excel to obtain the number of publications and total citation counts for each journal.
  4. Choose Finish, and the journal co-citation network map will be generated. Adjust the Attraction and Repulsion parameters to enhance the layout, refine the appearance using the Visualization panel, and save the map as a TIFF with a resolution of at least 300 dpi.
  5. In Excel, sort journals by the number of publications, select the top 10 journals by publication output. On March 9, 2026, access the official JCR platform (Table of Materials), search for each journal respectively, and record their latest Impact Factor and JCR quartile. Generate a chart displaying the top 10 journals by publication output.
  6. In Excel, sort journals by total citation count, select the top 10 journals by citation frequency. On March 9, 2026, retrieve and record their latest Impact Factor and JCR quartile from the official JCR platform. Generate a chart displaying the top 10 journals by citation frequency.

5. High-productivity author collaboration network analysis

  1. Import the dataset into VOSviewer, select Co-authorship as the analysis type, and select Authors as the unit of analysis.
  2. Set the minimum number of documents for an author to 3 to filter authors with a certain level of activity.
  3. In the Verify selected items interface, copy the author list and save the publication counts and citation counts of authors in an Excel file.
  4. Choose Finish to produce the author collaboration network, adjust the network layout and appearance, and save as TIFF format.
  5. In Excel, sort authors according to their number of publications, choose the 10 most productive researchers, and generate a donut chart to display their publication counts.
  6. From the original dataset, retrieve the top 10 authors by citation count, identify their respective countries, assign distinct colors based on country, and generate a chart displaying the top 10 authors by citation frequency.

6. Keyword co-occurrence and thematic evolution analysis

  1. Launch CiteSpace and navigate to Data > Import/Export.
  2. Place the txt file exported in step 1.12 into the input folder. In the Import/Export interface, set the input and output paths, select Article and Review as the document types, and click Start to perform data format conversion.
  3. After conversion, copy the files from the output folder to the data folder. In the CiteSpace main interface, click New, set the project directory and data directory, select WoS as the data source, and choose English as the language.
  4. Go to the time slicing panel on the right, set the span to 2016–2025, adopt a 1-year window, and choose Keyword as the node type.
  5. Apply the g-index (k=10) for network scaling, and select the three pruning strategies: Pathfinder, Pruning sliced networks, and Pruning the merged network.
  6. Click Start to run the analysis. After completion, use the One-Click Clustering Label Optimization function to create the keyword co-occurrence map.
  7. Set the display of node and cluster labels from the control panel, and optimize visual elements such as node size.
  8. Select Layout > Timeline to generate the keyword timeline view.
  9. In the control panel, select Burstness, click Refresh and View sequentially, set the number of top 25 keywords to display, and generate the chart of the top 25 keywords by burst strength. Export the generated images in TIFF format for storage.

7. Reference co-citation analysis

  1. Load the same dataset into CiteSpace. Set the time range and pruning strategies as described in steps 6.4–6.5, and select Cited Reference as the node type.
  2. After running the analysis, use the One-Click Clustering Label Optimization function to generate the reference co-occurrence map.
  3. Adjust node and cluster labels, and select Clusters to generate the reference clustering map.
  4. In the Burstness panel, set the display to show the top 20 cited references by citation frequency, and generate the chart of the top 20 co-cited references by burst strength.
  5. Export all images in TIFF format for storage.

8. PPI network analysis

  1. Access the GeneCards database (https://www.genecards.org)17on March 12, 2026, and search using “Interstitial Lung Disease” and “Pulmonary Hypertension” as keywords, respectively. The disease terms used for the search are consistent with those in the literature retrieval, and no additional term variants were expanded.
  2. Export the list of targets with a relevance score ≥ 1 for each search. The extraction method involved sorting by Relevance Score and exporting all targets with a score ≥ 1. The threshold of ≥ 1 was chosen because it captures the vast majority of literature-supported associated genes while minimizing false positives.
  3. Use R to obtain the intersection of the two target sets, resulting in the ILD-PH comorbidity-related targets.
  4. Access the Multiple Proteins module of the STRING database (https://cn.string-db.org/)18, enter the list of comorbidity targets into the search box, and select Homo sapiens as the species.
  5. Click Continue, then in the Settings panel, set the minimum required interaction score to High confidence (0.700), and enable hide disconnected nodes in the network19.
  6. After updating the network, download the interaction data as a TSV file from the Export menu.
  7. Start Cytoscape, and load the TSV file through File > Import > Network from File System.
  8. In Tools > NetworkAnalyzer > Network Analysis > Network Interpretation, select Treat the network as undirected, and calculate topological parameters, including node degree.
  9. Configure node shape, color, size, and other appearance settings in the Style panel, and manually adjust node positions to achieve a clear layout.
  10. Save the PPI network image via File > Export > Network Image to File, and export the network analysis results via File > Export > Table to File.
  11. Open the exported table in Excel, order by the Degree column descending, choose the top 20 key targets along with their degree values, create a bar chart with data labels, and export in TIFF format.

9. KEGG Pathway enrichment analysis

  1. Launch R, and enter the commands to install and load the required packages (Supplementary Table 3):
    install.packages("BiocManager")
    BiocManager::install(c("clusterProfiler", "org.Hs.eg.db"))
    library(clusterProfiler)
    library(org.Hs.eg.db)
  2. Convert the gene symbols of ILD-PH comorbidity targets to Entrez IDs using the bitr() function. In R, enter:
    coma_entrez <- bitr(coma_genes, fromType = "SYMBOL", toType = "ENTREZID", OrgDb = org.Hs.eg.db)
  3. Perform enrichment analysis using the enrichKEGG() function with the following parameters: organism = "hsa", pvalueCutoff = 0.05, qvalueCutoff = 0.05. The Benjamini-Hochberg (BH) method was used for multiple testing correction. Enter the following code:
    kegg_result <- enrichKEGG(gene = coma_entrez$ENTREZID,
    organism = "hsa",
    pvalueCutoff = 0.05,
    qvalueCutoff = 0.05,
    use_internal_data = FALSE)

    NOTE: p-value: the original p-value, representing the probability of observing the given level of enrichment; q-value: the FDR (False Discovery Rate) after Benjamini-Hochberg correction, controlling for the false positive rate. In this study, q < 0.05 was used as the criterion for significant enrichment.
  4. Filter the results using the keywords “Human Diseases” and “Metabolism” to remove pathways associated with human diseases and metabolism, retaining pathways associated with biological processes. Human Diseases and Metabolism pathways mainly describe the overall mechanisms of diseases and metabolic networks, which are weakly associated with the fibrosis, inflammation, and vascular remodeling mechanisms of ILD-PH comorbidity. Therefore, they were excluded to focus the results on key signaling pathways. Enter:
    kegg_result_df <- as.data.frame(kegg_result)
    kegg_filtered <- kegg_result_df[!grepl("Human Diseases|Metabolism", kegg_result_df$Description), ]
  5. Sort the pathways in descending order by gene count, select the top 10 pathways, generate a dot plot using the dotplot() function, and export the plot in PNG format. Enter:
    top10_pathways <- head(kegg_filtered[order(kegg_filtered$Count, decreasing = TRUE), ], 10)
    dotplot_result <- dotplot(enrichResult(select(top10_pathways)), showCategory = 10)
    ggsave("KEGG_dotplot.png", dotplot_result, width = 8, height = 6, dpi = 300)
  6. Extract the top 10 pathways and their associated target genes, construct a two‑column table (pathway name, target gene), and save the table as a tab‑delimited TXT file.
  7. Import the file into Cytoscape, configure the shape and color differences between pathway nodes and target nodes via the Style panel, manually refine the layout, and export the pathway‑target network image20.

Results

Loading...
$$\rightleftharpoonup{xx}$$ $$\longleftharp{xx}$$, $$\longrightharp{xx}$$,

Data retrieval and screening

In total, 2,400 records were collected from the Scopus database, while 696 were identified in the WoSCC database. After merging the two datasets and removing 533 duplicate records, a total of 2,563 unique publications were finally included for subsequent bibliometric analysis (Figure 1).

Annual publication output trends

The annual publication output in the field of interstitial lung disease and pulmonary hypertension from 2016 to 2025 is shown in Figure 2. Over the ten-year period, the number of publications exhibited a fluctuating upward trend. It remained relatively stable during the first four years, followed by rapid increases during 2020–2022 and 2023–2025, reaching two peaks in 2022 and 2025, respectively. Overall, the number of publications increased from 142 in 2016 to 402 in 2025, representing a nearly threefold increase (Figure 2).

National publication output and international collaboration

A total of 20 countries met the threshold of at least 40 publications and were included in the national collaboration network analysis. The United States led in both publication output and citation frequency, with 823 publications and 21,943 citations. Italy ranked second with 271 publications (7,037 citations), followed by France (237 publications, 8,172 citations), the United Kingdom (231 publications, 9,764 citations), and Japan (205 publications, 3,221 citations). China ranked sixth with 185 publications and 2,978 citations, while Germany ranked 7th with 184 publications and 7,932 citations (Figure 3A, Table 1).

In terms of international collaboration, total link strength indicators revealed that the United States (total link strength = 523), Italy (448), the United Kingdom (447), Germany (443), and France (405) exhibited the closest collaborative relationships (Figure 3B). The collaboration network analysis identified three main clusters. The first cluster consisted of 10 European countries, including Belgium, France, Germany, Greece, Italy, the Netherlands, Poland, Spain, Switzerland, and Turkey, forming a highly interconnected European research network. The second cluster comprised nine countries, including Australia, Brazil, Canada, China, India, Japan, South Korea, the United Kingdom, and the United States. This cluster was centered around the United States while integrating major research forces from Asia, the Americas, and Oceania, presenting a cross-regional collaboration pattern. The third cluster included only Austria, showing a relatively independent collaboration pattern (Figure 3C).

Journal publication output and citation analysis

Among the journals publishing in this field, Journal of Clinical Medicine (54 documents), Pulmonary Circulation (53 documents), and Clinical Rheumatology (51 documents) ranked highest in publication volume, with Chest (IF = 8.6, Q1) and Frontiers in Immunology (IF = 5.9, Q1) demonstrating the highest impact factors among the leading 10 journals in terms of productivity (Figure 4A). In terms of citation impact, European Respiratory Journal (1,885 citations), Journal of Heart and Lung Transplantation (1,817 citations), and European Respiratory Review (1,600 citations) emerged as the most influential journals (Figure 4B). The journal co-occurrence network revealed a high degree of clustering among respiratory medicine, rheumatology, and transplantation journals, and multidisciplinary collaboration is required for managing the comorbidity of ILD and PH (Figure 4C).

Author publication output and collaboration network

Among the most productive authors, Nathan, Steven D from the United States led with 45 publications, followed by a cluster of French authors, including Launay, David (27), Cottin, Vincent (27), and Humbert, Marc (26), highlighting France as a key contributor in this field (Figure 5A, Table 2). In terms of citation impact, Nathan, Steven D also led with 2,446 citations, while Humbert, Marc (1,675), Cottin, Vincent (1,347), and Shlobin, Oksana A (1,153) demonstrated substantial influence, with the United States and France dominating the list of highly cited authors (Figure 5B). The author collaboration network identified a total of six clusters. Overall, the collaboration pattern was predominantly domestic, with relatively limited international collaboration (Figure 5C).

Keyword co-occurrence and thematic evolution

Among the high-frequency keywords, apart from common database index terms such as "human" and "article", disease-specific keywords like "pulmonary hypertension", "interstitial lung disease", and "systemic sclerosis" occupied central positions, reflecting the main research themes in this field (Table 3). Centrality analysis revealed that terms related to therapeutic drugs (azathioprine, corticosteroid therapy), study design terms (randomized controlled trial, clinical article, case report), and symptom descriptors (dyspnea) exhibited high betweenness centrality, serving as bridging links connecting different themes such as etiology, symptoms, diagnosis, treatment, and prognosis. Overall, research in this field heavily relies on clinical evidence, particularly clinical studies related to immunosuppressive therapy (Figure 6A, Table 4).

According to the analysis of the top 25 keywords with the strongest citation bursts illustrated in Figure 6B, a temporal evolution in research focus can be observed. Early bursts included terms such as “pathophysiology” and “lung diffusion capacity” (2016–2020), followed by a mid-period emphasis on “differential diagnosis” and “lung lavage” (2018–2022), with recent bursts highlighting “COVID-19,” “fatigue,” “NT-proBNP,” and “intensive care unit” (2021–2025), reflecting a shift from fundamental mechanisms toward diagnostic procedures, biomarker identification, and critical care management.

The timeline view illustrates the temporal evolution of key research themes (Figure 6C). Early bursts (2016–2020) included terms such as "priority journal," "pathophysiology," "lung diffusion capacity," "randomized controlled trial (topic)," and "tomography," reflecting an emphasis on fundamental pathophysiological mechanisms and diagnostic imaging. Mid-period bursts (2018–2022) featured "differential diagnosis," "lung lavage," "steroid," and "practice guideline," indicating a transition toward diagnostic procedures and clinical management. Recent bursts (2021–2025) encompassed "coronavirus disease 2019," "laboratory test," "fatigue," "amino terminal pro brain natriuretic peptide," "intensive care unit," and "receiver operating characteristic," highlighting emerging interests in COVID-19-related impacts, biomarker identification, and critical care, demonstrating a progression from basic pathophysiology and pulmonary function testing toward clinical assessment tools and targeted therapies.

Reference co-citation analysis

The reference co-citation network identified 16 major clusters, with the largest clusters centered on idiopathic pulmonary fibrosis, systemic sclerosis, connective tissue disease, pulmonary hypertension, and progressive pulmonary fibrosis, reflecting the core knowledge domains in this field (Figure 7A). The most frequently cited reference was a study on the clinical management of systemic sclerosis, followed by a study on the efficacy of inhaled treprostinil for pulmonary hypertension associated with interstitial lung disease, and a study updating the hemodynamic definition and clinical classification of pulmonary hypertension4,21,22(Table 5). In terms of betweenness centrality, a randomized controlled trial comparing mycophenolate mofetil with oral cyclophosphamide for systemic sclerosis-related interstitial lung disease ranked highest, closely followed by a study on initial combination therapy for connective tissue disease-associated pulmonary arterial hypertension, indicating the pivotal bridging roles of these two studies in connecting different research domains23,24 (Figure 7B, Table 6). Citation burst analysis revealed a temporal evolution of influential works in this field: early bursts (2016-2019) included the ATS/ERS classification of idiopathic interstitial pneumonias and the DETECT study for screening systemic sclerosis-associated pulmonary arterial hypertension; mid-period bursts (2017-2021) encompassed the 2015 ESC/ERS guidelines for the diagnosis and treatment of pulmonary hypertension and the SLS II trial for scleroderma-related interstitial lung disease; recent bursts (2020-2025) included the 2022 ATS/ERS/JRS/ALAT clinical practice guideline for idiopathic pulmonary fibrosis and the study of nintedanib in progressive fibrosing interstitial lung disease. This evolution reflects a shift from diagnostic classification toward targeted therapeutic strategies25,26,27,28 (Figure 7C).

PPI analysis of overlapping genetic targets

Through Venn diagram analysis, 271 overlapping genes associated with ILD and PH were identified (Figure 8A). Based on it, a PPI network was established to explore the functional relationships among these shared targets (Supplementary Table 4). The PPI network comprised 271 nodes and 3,414 edges, with an average node degree of 25.2, indicating a densely interconnected network. The PPI network visualization reveals a highly interconnected architecture with multiple functional clusters (Figure 8B). The top 20 genes ranked by degree centrality include FN1, IL6, TNF, AKT1, EGFR, TGFB1, CTNNB1, IL1B, TP53, MMP9, ALB, STAT3, INS, IL10, CXCL8, CCL2, ICAM1, IFNG, SMAD4, and CRP, highlighting these molecules as central hubs within the interaction network (Figure 8C). Inflammatory and fibrotic processes may serve as candidate drivers of the comorbidity, which require experimental validation. The enriched pathways suggest potential core candidate pathways involved in the comorbidity mechanism.

KEGG pathway enrichment analysis

To further elucidate the functional mechanisms underlying the comorbidity of ILD and PH, KEGG pathway enrichment analysis was performed on the overlapping genetic targets (Supplementary Table 5). The most significantly enriched pathways included the PI3K-Akt signaling pathway, Relaxin signaling pathway, Integrin signaling pathway, FoxO signaling pathway, Focal adhesion, HIF-1 signaling pathway, Cellular senescence, TGF-beta signaling pathway, MAPK signaling pathway, and Phospholipase D signaling pathway (Figure 8D). The pathway-target network visualizes the interconnections between these enriched pathways and the hub genes identified in the PPI network (Figure 8E). Notably, core genes such as AKT1, EGFR, TGFB1, IL6, TNF, and FN1 were found to be involved in multiple pathways, particularly PI3K-Akt, MAPK, and TGF-beta signaling, suggesting that these pathways represent candidate mechanisms that need further experimental or clinical validation.

Protocol validation and reproducibility checkpoints

Protocol validation requires confirming that each module achieves the expected outputs and meets the key reproducibility checkpoints. For the bibliometrics module, the expected outputs include a column chart of annual publication output, a national collaboration network map, and ranking tables of the top 10 journals and authors. The reproducibility checkpoints are: the number of deduplicated records should account for more than 80% of the original retrieved records, CiteSpace data conversion should complete without errors, and major country nodes should be clearly distinguishable in the collaboration network. For the PPI network analysis module, the expected outputs include a PPI network graph with ≥250 nodes and ≥3,000 edges, as well as a list of the top 20 hub genes ranked by degree. The checkpoints are: the confidence threshold in STRING should be set to 0.700, the average node degree calculated by Cytoscape should be ≥20, and the network should contain no disconnected isolated nodes. For the KEGG enrichment analysis module, the expected outputs are a dot plot of the top 10 significant pathways and a pathway–target network graph. The checkpoints require that the enrichment analysis yields p < 0.05 and q < 0.05, that after excluding the “Human Diseases” and “Metabolism” categories, at least 5 pathways remain, and that the dot plot shows a clear decreasing trend in gene counts with ranking.

Venn diagram: WOS vs. Scopus overlap, data comparison, research database intersections.
Figure 1: Venn diagram of database retrieval results. This diagram illustrates the number of records retrieved from the Scopus and Web of Science Core Collection (WoSCC) databases, as well as the number of duplicate records removed. The overlapping area represents duplicate records identified by DOI, and the final unique records used for bibliometric analysis are indicated. Please click here to view a larger version of this figure.

Annual publications growth bar chart 2016-2025.
Figure 2: Interstitial lung disease-pulmonary hypertension research output and citations (2016–2025). Annual publication output and citation trends in ILD and PH research (2016–2025). The bar chart depicts the yearly number of articles, while the line graph represents the cumulative citation frequency. Please click here to view a larger version of this figure.

Global scientific collaboration visualization; bar chart, chord diagram, network map; citation data.
Figure 3: National/regional publication output and international collaboration. (A) National publication and citation ranking; (B) Inter-country collaboration chord diagram; (C) International collaboration network clusters. National publication and citation rankings, inter-country collaboration linkages via a chord diagram, and clustering patterns of international collaboration networks are illustrated in this figure. Please click here to view a larger version of this figure.

Journal impact and citation analysis with bar graphs and network diagram of respiratory research.
Figure 4: Journal publication output and citation. (A) Top 10 journals by publication volume; (B) Top journals by citation impact; (C) Journal co-occurrence network. Ranked by publication volume and citation counts, the top 10 journals are presented with their impact factors and quartiles. The network intuitively displays journal publication volume and collaborative relationships. Please click here to view a larger version of this figure.

Research collaboration analysis with citation donut chart, bar graph by country, and network diagram.
Figure 5: High-productivity author collaboration network. (A) Top 10 productive authors; (B) Top 10 authors by citation impact; (C) Author collaboration network. Authors were sorted by their publication count and citation influence. The country affiliations of the top 10 authors were retrieved from the downloaded Web of Science files and presented in a figure. The clustering network clearly and intuitively illustrates author collaboration patterns. Please click here to view a larger version of this figure.

Keyword cluster analysis, citation burst, and connectivity diagram in pulmonary research analysis.
Figure 6: Keyword co-occurrence and thematic evolution. (A) Keyword co-occurrence clustering; (B) Top 25 burst keywords; (C) Keyword evolution timeline. Keyword analysis was performed employing CiteSpace with the g-index configured to 10 and the random forest algorithm. The co-occurrence map was generated by adjusting node size and color, and clustering analysis was performed on the nodes to obtain the cluster map. The timeline view mainly illustrates the research trends and changes over the past decade. Please click here to view a larger version of this figure.

Pulmonary disease relationship and citation burst network diagrams with reference analysis table.
Figure 7: Reference co-citation analysis. (A) Cluster analysis of co-cited references. (B) Co-citation network of highly cited references. (C) Top 20 references with the strongest citation bursts. Co-cited reference analysis was performed using CiteSpace with the g-index set to 10 and the random forest algorithm. The co-citation network was generated by adjusting node size and color, and clustering analysis was performed on the nodes to obtain the cluster map, revealing different thematic research areas. Citation burst detection identified the top 20 references with the strongest citation bursts and their active periods, with an emphasis on reflecting the shifts in research hotspots over the past decade. Please click here to view a larger version of this figure.

Venn diagram analysis, gene network mapping, bar chart, pathway enrichment dot plot, signaling pathways.
Figure 8: Identification of ILD-PH comorbidity-related hub targets and enriched pathways. (A) Venn diagram showing overlapping genes associated with ILD and PH. (B) Protein-protein interaction (PPI) network. (C) Top 20 hub targets identified from the PPI network. (D) KEGG pathway enrichment analysis of overlapping genetic targets. (E) Interaction network illustrating the links between signaling pathways and their corresponding target genes. Comorbidity-related genes were obtained as the intersection of ILD and PH targets from GeneCards. A PPI network (confidence ≥ 0.700) was constructed using STRING, visualized in Cytoscape, and the top 20 hubs were selected by node degree. KEGG enrichment (q < 0.05, BH correction) was performed, excluding human disease and metabolism pathways. A pathway-target network was built to show associations between signaling pathways and their corresponding target genes. Please click here to view a larger version of this figure.

RankCountryDocumentsCitationsAverage citation per paperTotal link strength
1USA8232194326.7523
2Italy271703726.0448
3France237817234.5407
4UK231976442.3447
5Japan205322115.7147
6China185297816.168
7Germany184793243.1443
8Canada153598239.1282
9Spain133512238.5294
10Australia123433435.2217

Table 1: Top 10 countries in publication output and citation frequency. Data were retrieved from the Web of Science Core Collection and Scopus databases on March 9, 2026, and merged after deduplication. Countries were ranked by the total number of publications (Documents). “Total link strength” indicates the sum of collaboration link strengths with other countries in the VOSviewer co-authorship network. Average citation per paper is calculated as Citations / Documents.

RankAuthorDocumentsCitationsCountry
1nathan, steven d452446USA
2launay, david27784France
3cottin, vincent271347France
4humbert, marc261675France
5nikpour, mandana22250Australia
6allanore, yannick21482France
7stevens, wendy20236Australia
8hachulla, eric20706France
9khanna, dinesh20755USA
10proudman, susanna17163Australia

Table 2: Top 10 productive authors in the field of ILD and PH. Authors were ranked by the total number of publications. Only authors with at least 3 publications were included. Citation counts represent the total number of times the author’s publications in the dataset have been cited. Country affiliation is based on the author’s primary institution as recorded in the retrieved articles.

RankKeywordsCount
1human2311
2pulmonary hypertension2289
3interstitial lung disease2201
4article1719
5humans1544
6female1444
7male1410
8adult1276
9middle aged903
10major clinical study897

Table 3: The top 10 keywords by frequency. Keywords were retrieved from the deduplicated dataset and analyzed using CiteSpace. Frequency indicates the number of occurrences of each keyword in the title, abstract, or keyword fields of the retrieved literature. Generic indexing terms are automatically assigned by the databases and reflect indexing practices rather than specific research focus.

RankKeywordsCentrality
1human0.95
2article0.71
3azathioprine0.44
4clinical feature0.43
5randomized controlled trial (topic)0.42
6corticosteroid therapy0.41
7humans0.39
8clinical article0.35
9case report0.35
10dyspnea0.28

Table 4: The top 10 keywords by centrality. Centrality was calculated using CiteSpace to identify keywords that serve as bridges between different research clusters. Higher centrality values indicate stronger connectivity and a greater mediating role in the co-occurrence network.

RankCountCited References
168Khanna D, 2017, SYSTEMIC SCLEROSIS @ LANCET, V390, P1685-1699
241Thenappan T, 2021, INHALED TREPROSTINIL IN PULMONARY HYPERTENSION DUE TO INTERSTITIAL LUNG DISEASE @ N ENGL J MED, V384, P325-334
336Celermajer DS, 2019, HAEMODYNAMIC DEFINITIONS AND UPDATED CLINICAL CLASSIFICATION OF PULMONARY HYPERTENSION @ EUR RESPIR J, V0, P53
432Cottin V, 2020, SPECTRUM OF FIBROTIC LUNG DISEASES @ N ENGL J MED, V383, P958-968
529Shlobin OA, 2020, THE TROUBLE WITH GROUP 3 PULMONARY HYPERTENSION IN INTERSTITIAL LUNG DISEASE: DILEMMAS IN DIAGNOSIS AND THE CONUNDRUM OF TREATMENT @ CHEST, V158, P1651-1664
627Gahlemann M, 2019, NINTEDANIB FOR SYSTEMIC SCLEROSIS-ASSOCIATED INTERSTITIAL LUNG DISEASE @ N ENGL J MED, V380, P2518-2528
726Cottin V, 2019, NINTEDANIB IN PROGRESSIVE FIBROSING INTERSTITIAL LUNG DISEASES @ N ENGL J MED, V381, P1718-1727
826Richeldi L, 2022, IDIOPATHIC PULMONARY FIBROSIS (AN UPDATE) AND PROGRESSIVE PULMONARY FIBROSIS IN ADULTS: AN OFFICIAL ATS/ERS/JRS/ALAT CLINICAL PRACTICE GUIDELINE @ AM J RESPIR CRIT CARE MED, V205, P0
924Clements PJ, 2016, MYCOPHENOLATE MOFETIL VERSUS ORAL CYCLOPHOSPHAMIDE IN SCLERODERMA-RELATED INTERSTITIAL LUNG DISEASE (SLS II): A RANDOMISED CONTROLLED DOUBLE-BLIND PARALLEL GROUP TRIAL @ LANCET RESPIR MED, V4, P708-719
1023Hoeper MM, 2022, 2022 ESC/ERS GUIDELINES FOR THE DIAGNOSIS AND TREATMENT OF PULMONARY HYPERTENSION @ EUR HEART J, V43, P3618-3731

Table 5: The top 10 most frequently cited references. Cited references were retrieved from the deduplicated dataset and analyzed using CiteSpace. “Count” refers to the frequency with which a reference was co-cited within the literature. Only references with the highest co-citation counts are shown.

RankCentralityCited References
10.42Clements PJ, 2016, MYCOPHENOLATE MOFETIL VERSUS ORAL CYCLOPHOSPHAMIDE IN SCLERODERMA-RELATED INTERSTITIAL LUNG DISEASE (SLS II): A RANDOMISED CONTROLLED DOUBLE-BLIND PARALLEL GROUP TRIAL @ LANCET RESPIR MED, V4, P708-719
20.41Barbera JA, 2017, INITIAL COMBINATION THERAPY WITH AMBRISENTAN AND TADALAFIL IN CONNECTIVE TISSUE DISEASE-ASSOCIATED PULMONARY ARTERIAL HYPERTENSION (CTD-PAH): SUBGROUP ANALYSIS FROM THE AMBITION TRIAL @ ANN RHEUM DIS, V76, P1219-1227
30.32Hachulla E, 2013, SURVIVAL IN SYSTEMIC SCLEROSIS-ASSOCIATED PULMONARY ARTERIAL HYPERTENSION IN THE MODERN MANAGEMENT ERA @ ANN RHEUM DIS, V72, P1940-1946
40.26Visovatti S, 2019, PREVALENCE TREATMENT AND OUTCOMES OF COEXISTENT PULMONARY HYPERTENSION AND INTERSTITIAL LUNG DISEASE IN SYSTEMIC SCLEROSIS @ ARTHRITIS RHEUMATOL, V71, P1339-1349
50.25Dimopoulos K, 2014, BOSENTAN IN PULMONARY HYPERTENSION ASSOCIATED WITH FIBROTIC IDIOPATHIC INTERSTITIAL PNEUMONIA @ AM J RESPIR CRIT CARE MED, V190, P208-217
60.24Smith P, 2021, INHALED TREPROSTINIL IN PULMONARY HYPERTENSION DUE TO INTERSTITIAL LUNG DISEASE @ N. ENGL. J. MED, V384, P325-334
70.24Pausch C, 2022, PHENOTYPING OF IDIOPATHIC PULMONARY ARTERIAL HYPERTENSION: A REGISTRY ANALYSIS @ LANCET RESPIR MED, V10, P937-948
80.21Cottin V, 2019, NINTEDANIB IN PROGRESSIVE FIBROSING INTERSTITIAL LUNG DISEASES @ N ENGL J MED, V381, P1718-1727
90.21Beghetti M, 2016, 2015 ESC/ERS GUIDELINES FOR THE DIAGNOSIS AND TREATMENT OF PULMONARY HYPERTENSION: THE JOINT TASK FORCE FOR THE DIAGNOSIS AND TREATMENT OF PULMONARY HYPERTENSION OF THE EUROPEAN SOCIETY OF CARDIOLOGY (ESC) AND THE EUROPEAN RESPIRATORY SOCIETY (ERS): ENDORSED BY: ASSOCIATION FOR EUROPEAN PAEDIATRIC AND CONGENITAL CARDIOLOGY (AEPC) @ INTERNATIONAL SOCIETY FOR HEART AND LUNG TRANSPLANTATION (ISHLT), VEur. Heart J, P67-119
100.21Pope JE, 2018, TREATMENT ALGORITHMS FOR SYSTEMIC SCLEROSIS ACCORDING TO EXPERTS @ ARTHRITIS RHEUMATOL, V70, P1820-1828

Table 6: The top 10 references in terms of centrality. Centrality was calculated using CiteSpace. It measures the importance of a reference as a bridge connecting different co-citation clusters. Higher values indicate that the reference plays a key role in integrating diverse research topics.

Supplementary Table 1: Database search parameters for Web of Science Core Collection and Scopus. This table details the database versions, access dates, search fields, date ranges, document types, language restrictions, and subject area filters applied during the literature retrieval process.Please click here to download this file.

Supplementary Table 2: Full literature records retrieved from Web of Science and Scopus before and after duplicate removal. This table presents the complete exported records of all documents identified from the Web of Science Core Collection and Scopus. The table contains three sections: records from WoSCC alone, records from Scopus alone, and the deduplicated merged list. Duplicate records were identified and removed based on DOI.Please click here to download this file.

Supplementary Table 3: R scripts used for data processing, analysis, and visualization in this study. This table lists all R scripts used in this study, including those for bibliometric analysis, data merging and deduplication, Venn diagram generation, and statistical and plotting tasks for intersecting genes.Please click here to download this file.

Supplementary Table 4: Protein-protein interaction (PPI) raw data. This table presents the specific data downloaded from the STRING database and visualized during the PPI analysis step.Please click here to download this file.

Supplementary Table 5: Detailed raw data of KEGG pathways retrieved for this study. This table presents the complete raw data of KEGG (Kyoto Encyclopedia of Genes and Genomes) pathway enrichment analysis, including pathway ID, pathway name, the number of input genes mapped to the pathway, the list of involved genes, and the specific enrichment parameters.Please click here to download this file.

Discussion

Loading...
$$\rightleftharpoonup{xx}$$ $$\longleftharp{xx}$$, $$\longrightharp{xx}$$,

This study represents the first integrated application of bibliometrics and bioinformatics to explore the comorbidity of ILD and PH. It systematically reviews the developmental trajectory, knowledge structure, and cutting-edge trends in this field over the past decade, and identifies the hub genes and key signaling pathways underlying this comorbid state at the molecular level. From a macro-trend perspective, the annual number of publications in this field increased from 142 in 2016 to 402 in 2025, representing a nearly threefold rise, reflecting a sustained growth in academic attention to ILD-PH comorbidity. Significant surges were observed in two periods: 2020–2022 and 2023–2025. The growth inflection point around 2020 coincided with the COVID-19 pandemic; it is possible that the pandemic may have indirectly stimulated a surge in related research by inducing or exacerbating ILD and PH. The peak from 2022 to 2023 may be attributed to the increased exploration, validation, and preliminary clinical application of PH-ILD treatment methods29,30,31. The latter surge might be partly explained by the development of novel targeted therapies and the application of artificial intelligence in diagnosis32,33,34.

Regarding national contributions and international collaboration, the United States led with 823 publications and 21,943 total citations. Its central position is reflected not only in output volume but also in its significantly higher total link strength in the international collaboration network compared with other countries, forming an extensive collaborative network. European countries, including Italy, France, and the United Kingdom, followed closely, constituting a highly interconnected research cluster with close cooperative ties. Among Asian countries, Japan and China ranked fifth and sixth in publication volume, respectively, but their citations per paper remain markedly lower than those of European countries such as Germany and France, indicating that their academic influence in this field requires further improvement. In terms of journal distribution, high-impact journals, including the European Respiratory Journal and the Journal of Heart and Lung Transplantation, publish the most influential findings in this field, while specialized journals such as the Journal of Clinical Medicine and Pulmonary Circulation contribute the majority of publication output. This distribution pattern demonstrates that research in this field is primarily published in respiratory specialist journals while also involving multidisciplinary collaboration. At the author level, the American scholar Nathan, Steven D ranked first in both publication and citation counts, making him the most influential core researcher in this field. French scholars Launay, David, Cottin, Vincent, and Humbert, Marc formed the second-largest research force centered in France. This distribution is highly consistent with the dominance of the United States and France observed in country and journal analyses, further confirming the leading roles of these two nations in ILD-PH comorbidity research.

The integration of bibliometrics and bioinformatics in this protocol offers a dual-layer analytical framework that goes beyond either method alone. Bibliometrics captures macro-level research trends, collaboration networks, and evolving hotspots from the literature, providing a bird's-eye view of the field. Bioinformatics, on the other hand, mines molecular data from public databases to identify candidate genes, protein interactions, and enriched pathways. By linking the two layers, this approach allows researchers to assess whether a molecular target or pathway aligns with emerging clinical interests, thereby generating hypotheses that are both literature-supported and mechanistically plausible. This synergy is particularly valuable for complex comorbidities like ILD-PH, where cross-disciplinary integration is often lacking.

Regarding the evolutionary trajectory of research themes, co-occurrence and burst detection analyses of keywords and references collectively reveal a developmental path in this field: from understanding disease mechanisms, to exploring diagnosis and treatment norms, and then to precision management and prognosis evaluation. During 2016–2020, the emergence of burst keywords such as "pathophysiology", "lung diffusion capacity" and "tomography" reflected fundamental research on the pathological mechanisms of ILD-PH and pulmonary function/imaging evaluation methods. From 2018 to 2022, burst keywords including "differential diagnosis", "practice guideline" and "steroid" indicated that standardized clinical diagnosis and treatment became the focus of attention. Between 2021 and 2025, burst keywords such as "COVID-19", "fatigue", "NT-proBNP", "intensive care unit" and "receiver operating characteristic" demonstrated a shift in research hotspots toward precision and critical care management. This evolutionary trend is further supported by reference analysis: early landmark studies established the diagnostic foundation, including the classification of idiopathic interstitial pneumonias and the DETECT study for screening. Mid‑term guideline documents, such as the 2015 ESC/ERS pulmonary hypertension guidelines, promoted standardization of diagnosis and treatment. More recently, burst publications, including the updated IPF clinical practice guideline and the trial of nintedanib in progressive fibrosing ILD, signaled a shift from empirical treatment to targeted precision interventions.

At the molecular mechanism level, bioinformatics analysis identified 271 ILD-PH comorbidity-related genes, and subsequent network analysis revealed the potential molecular basis of this comorbidity. A PPI network showed high connectivity among 20 core genes, including FN1, IL6, TNF, AKT1, EGFR, and STAT1. Among these, FN1 is a key component of the extracellular matrix (ECM) and participates in matrix assembly and remodeling35,36,37. The fibrotic process in ILD may lead to excessive ECM deposition, which is speculated to trigger pulmonary vascular remodeling through mechanotransduction and growth factor sequestration. Inflammatory factors such as IL6, TNF, and IL1B not only drive inflammatory injury in ILD but may also participate in the initiation and progression of PH by activating endothelial cells and promoting smooth muscle proliferation38,39,40,41. KEGG pathway enrichment analysis further mapped these core genes to critical signaling networks, including the PI3K-Akt, TGF-β, MAPK, HIF-1, and integrin signaling pathways. These pathways have been previously verified to participate in the pathogenesis of pulmonary fibrosis and pulmonary hypertension. Specifically, the TGF-β signaling pathway is involved in both pulmonary fibrosis and vascular remodeling, serving as a key molecular bridge linking ILD and PH42,43,44; the HIF-1 signaling pathway regulates vasoconstriction and vascular remodeling45,46; and the PI3K-Akt pathway is widely implicated in cell survival, proliferation, and inflammation47,48,49. Based on the above bioinformatics analysis, this study preliminarily suggests that ILD and PH may share activation of multiple signaling pathways. Future research should focus on the following directions: first, exploring subtype-stratified mechanisms of ILD-PH comorbidity, including molecular characteristics of distinct subgroups such as IPF-PH and CTD-ILD-PH; second, developing targeted drugs for comorbidity and promoting clinical translation; third, strengthening the construction of global multicenter collaborative networks.

To enhance the practical reproducibility of this protocol, this study discusses the following steps that are prone to issues. First, database export incompatibility may occur when different databases use inconsistent field tags or file formats. We recommend exporting records in plain text format with a standardized tag set such as “Full Record and Cited References” for Web of Science and “CSV export with all available fields” for Scopus. Second, duplicate matching conflicts often arise when the DOI is missing or contains minor variations like extra spaces or lowercase letters. Our protocol uses DOI as the primary deduplication key. If a record lacks a DOI, use a composite key combining title, author names, publication year, and journal. For mismatches due to inconsistent formatting, apply a string normalization step that removes punctuation, converts to lowercase, and trims whitespace. Third, threshold sensitivity in bioinformatics analyses, such as the combined score cutoff in STRING or the p‑value and logFC cutoffs in differential expression, can strongly influence the resulting PPI network and hub gene lists. If instability is observed, select a more conservative threshold or integrate evidence from multiple databases. Fourth, KEGG filtering and output errors may appear as missing pathway annotations or failure to retrieve updated pathway maps. This typically results from an outdated KEGG API version or incorrect organism code. Ensure that the latest version of the enrichment tool, such as clusterProfiler, is installed and that the organism parameter is set correctly, for example, “hsa” for human. If the output contains pathways without any mapped genes, verify that the input gene identifiers match the annotation database version. By attempting these adjustments, researchers can preliminarily identify and resolve the corresponding issues.

Overall, this protocol combines macro-level hotspot identification through bibliometrics with comorbidity target mining through bioinformatics as a preliminary exploratory approach, which may provide useful insights for research in the field of comorbidities. Nevertheless, several limitations of this protocol should be acknowledged. Regarding reproducibility, the transferability of this framework to other disease comorbidities may depend on factors such as database coverage, search strategy, disease terminology, and the availability of relevant molecular datasets. The bibliometric component excludes non-English articles, which may limit the comprehensiveness of the knowledge map. The bioinformatics component relies on public databases with heterogeneous data sources and sample compositions, which may affect the reproducibility and cross-dataset consistency of the results. In addition, the date of database updates may also influence the results. Finally, all bioinformatics findings are associative and hypothesis-generating rather than causal, and further validation through animal models or clinical cohort studies is still required before they can be applied to new research questions.

Disclosures

Loading...
$$\rightleftharpoonup{xx}$$ $$\longleftharp{xx}$$, $$\longrightharp{xx}$$,

The authors report no conflicts of interest in this work.

Acknowledgements

Loading...
$$\rightleftharpoonup{xx}$$ $$\longleftharp{xx}$$, $$\longrightharp{xx}$$,

The authors gratefully acknowledge the financial support from the Noncommunicable Chronic Diseases-National Science and Technology Major Project (2024ZD0522300, 2024ZD0522301) and the Non‑profit Central Research Institute Fund of Chinese Academy of Medical Sciences (2022‑ZHCH330‑01).

Materials

List of materials used in this article
NameCompanyCatalog NumberComments
Bibliometric Online Analysis PlatformK-Synth Srl (Bibliometric)https://bibliometric.comOnline analysis platform; module: Country relations; generates country collaboration chord diagram; accessed March 10, 2026
CiteSpaceDrexel University6.4.R1Bibliometric analysis; time slicing: 2016–2025; node types: keywords, references; pruning: Pathfinder + Pruning sliced networks
clusterProfiler (R package)Guangchuang Yu et al.4.13.2KEGG enrichment analysis; thresholds: p < 0.05, q < 0.2
CytoscapeCytoscape Consortium3.10.3Network visualization and analysis; plugin: CytoHubba; algorithm: degree centrality; used for hub gene identification
GeneCardsWeizmann Institute of Sciencehttps://www.genecards.orgGene information database; database version 5.22; accessed March 12, 2026; filtering criterion: Relevance score ≥ 1
Journal Citation Reports (JCR)Clarivate Analyticshttps://jcr.clarivate.comJournal impact factor and quartile retrieval; accessed March 9, 2026
KEGGKyoto Encyclopedia of Genes and Genomeshttps://www.kegg.jpPathway database; accessed via clusterProfiler API on March 13, 2026; species: hsa
MeSH DatabaseNational Library of Medicine (NLM)https://meshb.nlm.nih.govMeSH term query database; accessed March 9, 2026; used to construct search strategy
Microsoft ExcelMicrosoft Corporation16.88 (Microsoft 365)Data statistics and chart generation; used for publication count, deduplication, chart plotting
org.Hs.eg.db (R package)Bioconductor Core Team3.20.0Human genome annotation database; used for gene ID conversion
RR Core Team4.4.2Statistical computing and plotting environment; run date: March 13, 2026
Scimago GraphicaScimago Lab2Visualization tool; layout: map projection; edge curvature: 0.5; color mapping: by cluster
ScopusElsevierhttps://www.scopus.comLiterature retrieval database; Elsevier Scopus 2026 complete edition; retrieval date: March 9, 2026
STRINGEMBL12.0 / https://cn.string-db.orgProtein-protein interaction database; accessed March 12, 2026; interaction score threshold: 0.7 (high confidence); network type: physical + functional
VOSviewerLeiden University1.6.20Bibliometric analysis; co-authorship as analysis type; normalization method: association strength
Web of Science Core CollectionClarivate Analyticshttps://www.webofscience.comLiterature retrieval database; covers SCIE and SSCI sub-databases; retrieval date: March 9, 2026

Reprints and Permissions

Request permission to reuse the text or figures of this JoVE article

Request Permission

Tags

Biology

Related Articles