$$\rightleftharpoonup{xx}$$
$$\longleftharp{xx}$$,
$$\longrightharp{xx}$$,
The overall workflow of this study is illustrated in Figure 1, outlining the sequential steps of literature search, bibliometric analysis, identification of shared genes, construction of the protein-protein interaction network, and pathway enrichment analysis.
Literature search strategy
The literature search was conducted on June 22, 2025, on the Web of Science (WOS) Core Collection database19. The search was limited to the Science Citation Index Expanded (SCI-Expanded) and Social Sciences Citation Index (SSCI), covering publications from January 1, 2015, to December 31, 2024. The specific search terms and the corresponding number of results are provided in Table 1. After the initial search, document types such as meeting abstracts, conference papers, editorial materials, book chapters, and retracted publications were excluded. Subsequently, only records in the English language and of the "Article" or "Review Article" types were retained for further analysis.
Bibliometric analysis
Bibliometric analysis and visualization require the support of specific software tools, details of which are provided in Table of Materials. This study aims to construct a knowledge map of the interdisciplinary research field bridging myositis and rheumatoid arthritis. The analytical scope encompasses publication trend analysis over the past decade, patterns of international collaboration and collaborative networks, academic influence among institutions/journals/authors, core keywords reflecting research hotspots, and citation network analysis.
Annual publication count data were first exported from the Web of Science database and used to generate statistical charts illustrating publication trends. The retrieved literature data was then imported into VOSviewer software for analysis: within the "Co-authorship" module, options for "Authors," "Organizations," and "Countries" were sequentially selected to generate author collaboration networks, institutional collaboration networks, and country collaboration networks respectively; within the "Citation" module, the "Sources" option was selected to create journal collaboration networks; within the "Co-citation" module, "Cited Sources" and "Cited Authors" were chosen to obtain cited journal collaboration networks and cited author collaboration networks. To more intuitively visualize international collaborative relationships, the country collaboration network was further aesthetically optimized using Scimago Graphica software to enhance visual clarity and interpretability20.
Analysis of keyword hotspots and cross-disciplinary thematic terms was conducted using CiteSpace software21. The Pathfinder algorithm was selected for network pruning, with the time span set from January 2015 to December 2024, divided into monthly time slices. The node type was configured as "Keyword," yielding a keyword-based cluster map of research hotspots, representative thematic terms for each cluster, and the top 20 most frequently occurring hotspot keywords. Citation network analysis followed a similar procedural workflow, requiring only the modification of the node type to "Cited Reference" to accomplish the corresponding analysis.
Identification of shared genes
Based on bibliometric analysis, the keyword "myositis" showed significant prominence in the interdisciplinary research domain of muscle pathology and rheumatoid arthritis. Therefore, this study selected myositis and rheumatoid arthritis as research subjects to conduct overlapping gene analysis. The specific procedure was as follows: First, target genes associated with myositis and rheumatoid arthritis were retrieved from the GeneCards database22. To ensure the comprehensiveness of the analysis, this study did not set a relevance score threshold and included all reported genes associated with the diseases, thereby constructing two corresponding gene sets. By comparing these two sets, overlapping genes were identified. These overlapping genes were ultimately defined as "RA-Myositis comorbidity-related genes" and served as the core dataset for subsequent analyses.
Protein-protein interaction (PPI) network construction
Using the "Multiple proteins" query function in the STRING database, we submitted the "RA-myositis comorbidity-related genes" list for analysis23. The organism was set to "Homo sapiens" with a minimum required interaction score threshold of ≥0.700 (medium confidence). The resulting protein-protein interaction network data were then imported into Cytoscape for visualization. By customizing node shapes and colors through the STYLE panel, a clearly visualized protein-protein interaction network diagram was generated.
Pathway enrichment analysis
The specific workflow for pathway enrichment analysis proceeded as follows. First, Kyoto Encyclopedia of Genes and Genomes (KEGG) pathway enrichment analysis was performed on the shared gene set using the clusterProfiler package in R software24. Statistical significance was assessed by the hypergeometric test, and p-values were adjusted for multiple testing using the Benjamini-Hochberg method to control the false discovery rate. Pathways with an adjusted p-value < 0.05 were considered significantly enriched and were retained for further analysis. Pathways related to specific human diseases or general metabolic processes were then manually excluded to focus the analysis on core signaling mechanisms. The filtered pathways and their associated genes were subsequently imported into Cytoscape software to construct an interaction network25. The "yFiles Organic Layout" algorithm was applied to generate a clear hierarchical structure. Visual attributes were further optimized in the STYLE panel: node color intensity was mapped to enrichment significance, node size was weighted by the number of core genes in each pathway, and edge thickness was weighted by the Jaccard similarity coefficient between pathways to reflect the degree of gene overlap. Ultimately, a KEGG pathway interaction network was obtained for comprehensive analysis.