$$\rightleftharpoonup{xx}$$
$$\longleftharp{xx}$$,
$$\longrightharp{xx}$$,
Study design:
This study was a systematic review and meta-analysis conducted according to the Preferred Reporting Items for Systematic Reviews and Meta-Analyses (PRISMA) guidelines. The completed PRISMA 2020 checklist is provided as Supplementary File 1. The protocol aimed to quantitatively synthesize evidence on nurses' perceptions, attitudes, and intentions regarding the use of Artificial Intelligence (AI) in patient care, comparing those with prior AI knowledge to those without18. Figure 1 illustrates the study selection process.
Eligibility criteria
Studies were selected based on the following PICOS framework19:
Population (P): Registered nurses or nurse managers of any clinical specialty or setting.
Intervention (I)/Exposure: Having knowledge, training, or awareness of AI use in nursing practice.
Comparison (C): Nurses or nursing students who report not knowing how AI is used in nursing practice.
Outcomes (O): Quantitative measures of perception, attitude, or intention toward AI in patient care, reported as mean scores with standard deviations (SD) or data convertible to such.
Study design (S): Observational studies (cross-sectional or cohort) and survey-based studies. Interventional studies (e.g., randomized controlled trials) were excluded because the exposure of interest—knowledge of how AI is used in nursing practice—is an existing characteristic rather than a randomized intervention.
Exclusion criteria:
Qualitative studies, review articles, editorials, letters, conference abstracts without full data, and non-peer-reviewed literature were excluded from this study. Studies not providing comparative quantitative data (e.g., mean, SD, sample size) for the two groups (with vs. without AI knowledge) were excluded. Studies where the population was not exclusively or primarily nurses (e.g., mixed healthcare professional groups without separable nurse data) were also excluded. No studies were excluded based on language.
Information sources and search strategy
A systematic search was performed across five electronic databases from their inception until July 31, 202520: PubMed, Cochrane Library, Embase, OVID, and Google Scholar.
The search strategy combined controlled vocabulary (e.g., MeSH terms) and free-text keywords related to three core concepts: (1) Nurses, (2) Artificial Intelligence, and (3) Perception/Attitude. The Boolean operators "AND" and "OR" were used to link terms within and between concepts. A sample search strategy for PubMed is provided in Table 1. The reference lists of all included studies and relevant reviews were manually screened to identify additional eligible publications21. The complete reproducible search strings for each database are provided below. Syntax was adapted to each database's required format, including appropriate field tags, controlled vocabulary (MeSH, EMTREE), and Boolean operators. No language or date restrictions were applied. The search was performed on July 31, 2025.
Study selection process
The study selection process followed the PRISMA flow diagram (see Figure 1). All retrieved records were imported into EndNote X9 (Clarivate Analytics) for deduplication. Two independent reviewers screened titles and abstracts against the eligibility criteria. Studies deemed potentially relevant by either reviewer proceeded to full-text review. The same two reviewers independently assessed the full texts of the shortlisted studies. Disagreements at any stage were resolved through discussion and consensus. No third reviewer was required. The final list of studies meeting all criteria was agreed upon by consensus.
Data extraction and management22
A standardized, piloted data extraction form was developed in a spreadsheet. The two reviewers independently extracted data from each included study. Extracted data included study characteristics, population details, exposure definition, and outcome data.
Study characteristics: First author, publication year, country, study design, and sample size.
Population details: Nurse type (e.g., clinical, student, manager), clinical setting, mean age, gender distribution.
Exposure definition: How "knowledge of AI use in nursing practice" was defined and measured (e.g., specific training course, self-reported familiarity on a Likert scale).
Outcome data: For each relevant outcome (perception, attitude, intention), the mean score, standard deviation (SD), and sample size (n) for both the "AI-knowledgeable" and "non-AI-knowledgeable" groups were extracted. If means and standard deviations (SDs) were not directly reported, they were calculated from available statistics using methods described in the Cochrane Handbook for Systematic Reviews of Interventions (Version 6.4)19. Specifically:
From medians and interquartile ranges (IQR): The method of Wan et al. (2014)23 was used to estimate means and SDs.
From 95% confidence intervals (CIs) and sample sizes: SDs were back-calculated using the formula: SD = √n × (upper CI limit – lower CI limit) / (2 × 1.96).
From p-values and sample sizes: SDs were estimated using the method described in the Cochrane Handbook (Section 6.5.2.3) when t-statistics or exact p-values were available.
Among the 9 included studies, three required data conversion because means and SDs were not reported in the format required for meta-analysis:
Study [Author, Year]: Reported medians and IQRs; converted using Wan et al.23 method.
Study [Author, Year]: Reported only 95% CIs; SDs back calculated.
Study [Author, Year]: Reported means without SDs but provided p-values for group comparisons; SDs estimated from p-values.
The remaining 6 studies reported means and SDs directly and required no conversion. All converted values were verified by two reviewers independently.
Risk of bias (quality) assessment24
The methodological quality of the included observational studies was assessed independently by two reviewers using the Joanna Briggs Institute (JBI) critical appraisal checklist for analytical cross-sectional studies25. This tool evaluates domains such as sample representativeness, exposure and outcome measurement, confounding, and statistical analysis. Each item was scored as "Yes," "No," "Unclear," or "Not Applicable." An overall study quality rating (High, Moderate, Low) was assigned based on consensus. Disagreements were resolved as described above.
Data synthesis and statistical analysis
Effect measure: The primary effect measure was the mean difference (MD) with a 95% confidence interval (CI). An MD > 0 indicated a higher score (e.g., more positive attitude) in the group with AI knowledge.
Meta-analysis model: A random-effects model was used for all primary analyses due to anticipated clinical and methodological heterogeneity across studies. A random-effect model was also applied as a sensitivity analysis26,27.
Heterogeneity assessment: Statistical heterogeneity was quantified using the I2 statistic. I2 values of 25%, 50%, and 75% were interpreted as low, moderate, and high heterogeneity, respectively. Cochran's Q test (p < 0.10, indicating significant heterogeneity) was also consulted.
Subgroup and sensitivity analysis: Planned subgroup analyses were not feasible due to the limited number of studies (<10) per outcome. Sensitivity analyses were conducted by switching to a fixed-effect model. Sensitivity analyses were planned to assess the robustness of the pooled estimates, including excluding studies rated as low quality and switching to a fixed-effect model. A sensitivity analysis excluding the two studies rated as low quality (JBI score ≤4)27,28 was performed. The pooled effect estimates for the remaining seven studies (n = 3,161) were similar to the main analysis for all outcomes, and all associations remained statistically significant (p < 0.001). Therefore, all 9 studies were retained in the primary analysis. Because excluding these studies did not materially alter the findings, they were retained in the primary analysis to maximize sample size and generalizability.
Publication bias assessment: Visual inspection of funnel plots for asymmetry was planned for outcomes including ≥10 studies. Because all outcomes included fewer than 10 studies, formal tests such as Egger's regression test are underpowered and typically not recommended. However, an exploratory Egger's regression test was performed for the outcome with the largest number of studies (attitude, n = 9) as a post-hoc analysis. Results should be interpreted with extreme caution due to low statistical power.
Software: All statistical analyses were performed using Review Manager (RevMan) software, version 5.4 (The Cochrane Collaboration, Copenhagen, Denmark).