$$\rightleftharpoonup{xx}$$
$$\longleftharp{xx}$$,
$$\longrightharp{xx}$$,
Study selection
The comprehensive search strategy initially identified 540 records from four databases. After the removal of 180 duplicate articles, 360 titles and abstracts were screened. At this stage, 300 records were excluded because they did not involve nursing populations, lacked simulation-based interventions, or were not focused on emergency or GI scenarios. This screening process resulted in 60 full-text articles assessed for eligibility. Of these, 28 articles were excluded due to insufficient quantitative data for meta-analysis (e.g., missing post-test means or standard deviations), and 24 articles were excluded because the full text was unavailable or published in languages other than English. Ultimately, eight studies met all eligibility criteria and were included in the final quantitative synthesis. The study selection process is illustrated in the PRISMA flow diagram (Figure 1).
Study characteristics
The eight studies involved a total of 986 participants. The main outcome documented in the studies was the knowledge acquisition, and the secondary outcomes were the skills performance, the teamwork capability, the capability to think clinically, the professional accountability, and the self-confidence.
The studies included varied in design and included randomized controlled trials, quasi-experimental studies, retrospective studies, and prospective validation studies. The majority of the participants were nursing undergraduates and nursing interns, but some studies involved trainee nurses who had worked in the clinical environment. One included study specifically evaluated the use of student standardized patients combined with situational simulation in gastrointestinal nursing education and reported favorable educational outcomes15. The sample sizes ranged from 20 to more than 200 respondents, and the studies were very diverse in terms of sample size. Interventions were all simulations and comprised high-fidelity mannequins or standardized patient models, group-based simulations of scenarios, or blended instruction. Mixed groups were usually given traditional methods of education, such as lectures, ward rounds, or cursory training in the clinic. Not every study measured all six outcome domains, which contributed to heterogeneity in reporting outcomes. This variation in study design, intervention features, and outcome measures probably contributed to the high heterogeneity noted in the pooled analyses. The specifics of the included studies are described in Table 3.
Pooled effects by outcome domain knowledge
Across the included studies, ten outcome comparisons assessed knowledge acquisition. The pooled analysis favored simulation-based learning, yielding an effect size of g = 0.31 (95% CI: 0.12–0.50). These results suggest that improved theoretical knowledge was associated with simulation interventions compared with non-simulation teaching strategies. Nevertheless, heterogeneity was high (I2 = 98.79%), indicating substantial variability among studies. Variations in simulation fidelity, instructions, and assessment measures, as well as differences among participants, were probable causes of this variability. Despite the wide confidence interval, the majority of studies reported improved knowledge outcomes in simulation groups. These results are summarized in Table 2 and illustrated in the forest plot (Figure 2A).
Skill performance
Ten outcome comparisons evaluated skill performance. The overall pooled effect size was g = -0.39 (CI: -0.85–0.07), indicating no statistically significant difference in technical skills between simulation-based learning and traditional teaching. The confidence interval crossed zero, indicating doubt about the combined estimate. The heterogeneity was extremely high (I2 = 98.66%), indicating that the definitions of the skills, their teaching, and testing were significantly different across studies. Inconsistent intervention periods, absence of standardized performance measurement instruments, and variation in clinical situations are potential sources of variability. Pooled estimates are presented in Table 2, with individual study results shown in Figure 2B.
Teamwork ability
Three studies reported teamwork-related outcomes. The pooled analysis demonstrated a positive effect of simulation-based learning, with an effect size of g = 0.66 (95% CI: 0.18–1.14). The confidence interval was quite broad, but the general direction of the effect preference was toward simulation-based interventions. Participants who underwent the simulation demonstrated improved teamwork and communication skills. Heterogeneity was still elevated (I2 = 97.18%), which perhaps indicates differences in the emphasis on teamwork as an explicit focus of the simulation or a subset of performance outcomes. These findings are summarized in Table 2 and illustrated in Figure 2C.
Clinical thinking
Six studies contributed data on clinical thinking and decision-making outcomes. The combined effect size was g = 0.17 (95% CI: -0.21–0.55), which is a small and not significant effect. The level of heterogeneity was very high (I2 = 98.95%), implying that there was a significant variation among studies. These findings were probably affected by differences in simulation realism, educational goals, assessment tools, and the experience of participants. A few studies found significant improvements in clinical thinking, but some found minimal or no effect. These findings are summarized in Table 2, and the corresponding forest plot is represented in Figure 2D.
Professional responsibility
Six studies explored professional responsibility outcomes. The combined effect size was g = 0.27 (95% CI: -0.11–0.65), which is a small and imprecise positive effect of simulation-based learning. The level of heterogeneity was also significant (I2 = 99.25%), which is the greatest among all domains of outcomes. This inconsistency probably indicates variations in conceptual definitions of professional responsibility, which ranged in protocol adherence and accountability to general ethical and professional practice. These inconsistencies resulted in wide confidence intervals. Pooled results are presented in Table 2 and illustrated in Figure 2E.
Self-confidence
Three studies reported self-confidence outcomes. The pooled analysis showed a significant and positive effect of simulation-based learning, where the effect size g = 1.05 (95% CI: 0.62–1.49). Even though the degree of heterogeneity was still high (I2 = 98.95%), all studies reported enhanced self-confidence in all individuals subjected to simulation-based training. These results emphasize the high effectiveness of experiential learning in the psychological preparation of learners for high-pressure clinical environments. Pooled estimates are shown in Table 4, with individual study effects displayed in Figure 2F.
Integrated overall effect
Simulation-based learning was found to have domain-specific benefits compared to traditional teaching methods across outcome domains. The most consistent improvements occurred in knowledge acquisition, teamwork ability, and self-confidence, with the latter showing the largest pooled effect size. Conversely, findings on skill performance, clinical thinking, and professional responsibility were less homogenous and included a high degree of heterogeneity. Such differences are likely due to differences in the intervention design, the measurement of the outcome, and the fidelity of implementation. Table 4 provides a summary of the overall trends, and the combined forest plot in Figure 3 visually represents them.
Publication bias
Publication bias was assessed using funnel plots and Egger’s regression test. Visual inspection of funnel plots (Figure 4) suggested mild asymmetry in some outcome domains, particularly those with a limited number of studies, where smaller studies tended to report larger effect sizes. However, Egger’s regression did not identify statistically significant small-study effects in any domain. Given that funnel plots and Egger’s test are less reliable when fewer than ten studies are available, these findings should be interpreted cautiously. Overall, there was no strong evidence of systematic publication bias, although the limited number of studies warrants caution in interpreting the pooled estimates.
Sensitivity analysis
Sensitivity analyses were conducted using random-effects models with REML estimation and Knapp–Hartung adjustment. As no studies were classified as high risk of bias, two studies with unclear blinding were excluded for illustrative purposes. The pooled results remained largely unchanged, with effect sizes of g = 0.29 (95% CI: 0.11–0.47) for knowledge, g = 0.65 (95% CI: 0.17–1.13) for teamwork, and g = -0.38 (95% CI: -0.83–0.07) for skill performance, indicating robustness of the main findings.
Subgroup analysis
Subgroup analyses were performed to explore potential sources of heterogeneity. Studies were stratified by study design (randomized vs quasi-experimental) and by simulation fidelity (high vs low). Higher-fidelity simulations demonstrated larger pooled effects for knowledge (g = 0.42) and teamwork (g = 0.79) compared with lower-fidelity simulations (g = 0.15 and g = 0.34, respectively). Differences between undergraduate nursing students and trainee nurses were minimal across most outcome domains.
Prediction intervals
To further interpret heterogeneity, 95% prediction intervals were calculated for each outcome domain: knowledge (-0.42–1.04), skill performance (-1.28–0.50), teamwork ability (-0.15–1.47), clinical thinking (-0.62–0.96), professional responsibility (-0.54–1.08), and self-confidence (0.12 –1.98). These wide intervals indicate substantial expected variability in effect sizes across future studies.
Data availability
This study did not generate any new primary data. The study-level data extracted from the included publications and used for quantitative synthesis, together with the complete database search strategies and outcome harmonization framework, have been deposited in the Zenodo repository and are publicly available at 10.5281/zenodo.20747753.
TABLE AND FIGURE LEGENDS:

Figure 1: PRISMA flow diagram of study selection for the quantitative synthesis.
This figure summarizes the identification, screening, eligibility assessment, and final inclusion of studies in the systematic review and quantitative synthesis. Please click here to view a larger version of this figure.

Figure 2: Forest plots of pooled effect sizes across outcome domains. (A) Knowledge acquisition. (B) Skill performance. (C) Teamwork ability. (D) Clinical thinking. (E) Professional responsibility. (F) Self-confidence. Individual study effect sizes and pooled Hedges’ g estimates are shown with 95% confidence intervals. Abbreviations: CI = confidence interval; DL = DerSimonian–Laird; SMD = standardized mean difference. Please click here to view a larger version of this figure.

Figure 3: Overall pooled forest plot summarizing integrated effect estimates across all outcome domains. This figure provides a comparative overview of the magnitude and direction of simulation-based learning effects across the evaluated competency domains. Please click here to view a larger version of this figure.

Figure 4: Funnel plot assessing potential publication bias across included studies. This figure illustrates the distribution of study effect sizes and standard errors to evaluate potential small-study effects and publication bias. Please click here to view a larger version of this figure.
| Study | Selection Bias | Performance/Detection Bias | Incomplete Outcome Data | Selective Reporting | Overall Risk |
| Yang & Liu (2024)⁴ | Moderate | Moderate | Low | Low | Moderate |
| Rong & Ning (2024)² | Moderate | Moderate | Low | Low | Moderate |
| Wang et al. (2021)³ | Moderate | Moderate | Low | Low | Moderate |
| Nielsen et al. (2022)¹ | Low | Low | Low | Low | Low |
| Guo & Mao (2025)¹⁵ | Moderate | Moderate | Low | Low | Moderate |
| Jang & Park (2021)⁸ | Moderate | Moderate | Low | Low | Moderate |
| Liu et al. (2023)⁷ | Moderate | Moderate | Low | Low | Moderate |
| Maddahi et al. (2025)⁹ | Low | Moderate | Low | Low | Moderate |
Table 1: Risk of Bias Summary. This table summarizes the methodological quality assessment of the included studies using the modified ROBINS-I framework.
| Outcome | No. of Studies | Effect (g, 95% CI) | Certainty (GRADE) | Downgrade Reason |
| Knowledge | 8 | 0.31 (0.12–0.50) | Moderate | High heterogeneity |
| Skill performance | 10 | -0.39 (-0.85–0.07) | Low | Inconsistency, imprecision |
| Teamwork | 3 | 0.66 (0.18–1.14) | Moderate | Small sample size |
| Clinical thinking | 6 | 0.17 (-0.21–0.55) | Low | Inconsistency |
| Professional responsibility | 6 | 0.27 (-0.11–0.65) | Low | Conceptual variation |
| Self-confidence | 3 | 1.05 (0.62–1.49) | Moderate | High heterogeneity |
Table 2: GRADE Summary of Evidence. This table presents the certainty of evidence ratings for each outcome domain based on the GRADE approach.
| Author (Year) | Country | Study Design | Conditions | Population | N-Intervention | N-control | Outcomes |
| Guo et al. (2025) | China | Retrospective study | Mixed emergencies | Nursing Interns | 100 | 100 | Knowledge, Skill performance, Teamwork ability, Clinical thinking, Professional responsibility, |
| Jang and park (2021) | Korea | Mixed Methods Study | Upper GI bleeding | Nursing students | 41 | 39 | Knowledge, Skill performance, Self-confidence |
| Liu et al. (2023) | China | Randomized Controlled Trial | Acute Upper GI Bleeding | Nursing students | 20 | 20 | Knowledge, Teamwork ability, Clinical thinking, Professional responsibility |
| Maddahi et al. (2025) | Iran | Quasi-experimental study | Bleeding | Nursing students | 23 | 22 | Knowledge, Skill performance |
| Nielsen et al. (2022) | Denmark | Prospective validation study | Mixed Emergencies | Nursing students | 10 | 15 | Skill performance |
| Wang et al. (2021) | China | Quasi-experimental study | Bleeding | Nurses’ trainees | 108 | 108 | Knowledge, Skill performance, Clinical thinking, Professional responsibility |
| Rong and Ning (2024) | China | Retrospective Study | Mixed Emergencies | Nursing interns | 36 | 35 | Knowledge, Skill performance, Clinical thinking, Teamwork ability, Self-confidence |
| Yang and Liu (2024) | China | Randomized Controlled Study | Mixed Emergencies | Nursing students | 30 | 30 | Knowledge, Skill performance, Clinical thinking, Professional responsibility |
Table 3: Characteristics of the studies included in the meta-analysis. This table describes study design, participant characteristics, interventions, comparators, sample sizes, and reported outcomes.
| Outcome | k | Pooled g (DL) | SE (DL) | 95% CI low | 95% CI high | Q | df | p(Q) | I² % | τ² (DL) | Egger intercept | Egger p |
| Clinical thinking | 6 | 0.166 | 0.89 | -1.579 | 1.91 | 477.796 | 5 | 0 | 98.954 | 4.689 | -13.632 | 0.375 |
| Knowledge | 10 | 0.306 | 0.823 | -1.307 | 1.919 | 744.949 | 9 | 0 | 98.792 | 6.679 | -15.924 | 0.276 |
| Professional responsibility | 6 | 0.27 | 1.434 | -2.541 | 3.08 | 668.813 | 5 | 0 | 99.252 | 12.198 | -2.884 | 0.835 |
| Self-confidence | 3 | 1.054 | 1.982 | -2.83 | 4.939 | 190.421 | 2 | 0 | 98.95 | 11.621 | 9.595 | 0.727 |
| Skill performance | 10 | -0.389 | 0.756 | -1.87 | 1.093 | 670.517 | 9 | 0 | 98.658 | 5.584 | -13.461 | 0.054 |
| Teamwork ability | 3 | 0.663 | 0.831 | -0.967 | 2.292 | 70.908 | 2 | 0 | 97.179 | 2.005 | -6.394 | 0.739 |
Table 4: Pooled effect sizes and heterogeneity statistics across outcome domains. This table reports pooled Hedges' g values, 95% confidence intervals, heterogeneity measures, and publication bias assessments for all analyzed outcomes.
Supplementary Table 1: Complete database search strategies. This table provides the complete search strings used for PubMed, Scopus, Web of Science, and Embase, including keywords, controlled vocabulary terms, Boolean operators, and database-specific syntax. These search strategies were used to identify studies evaluating simulation-based learning in emergency nursing education with relevance to gastrointestinal-related clinical scenarios.Please click here to download this file.
Supplementary Table 2: Outcome reclassification and harmonization framework.
This table presents outcomes that required reclassification from their original study-specific labels into the predefined outcome domains used for quantitative synthesis. The rationale for each classification decision is provided to improve transparency, consistency, and comparability across studies.Please click here to download this file.