A subscription to JoVE is required to view this content. Sign in or start your free trial.

Review Article

The Landscape of Virtual Reality and Generative AI in Oncology Education: A Scoping Review of Clinical Reasoning Assessment Tools

159 views

DOI:

10.3791/70722

May 15th, 2026

In This Article

Summary

This scoping review finds VR and GenAI offer new ways to assess clinical reasoning in oncology training. However, both require more robust validation and ethical frameworks before widespread educational use, as current evidence is preliminary.

Abstract

Oncology is a setting in which clinical reasoning is a fundamental competency necessitating the synthesis of complex clinical information in a high-stakes setting. Conventional means of assessment, such as written tests, Objective Structured Clinical Examinations (OSCEs) and workplace-based assessments, offer valuable knowledge, but are still not effective in reflecting the dynamic and iterative quality of clinical reasoning. New opportunities to overcome these limitations are presented by emerging digital technologies, especially Virtual Reality (VR) and Generative Artificial Intelligence (GenAI). VR facilitates simulation-based environments, which can be immersive, and which allow the evaluation of procedural and situational factors of reasoning based on observable clinical action. GenAI, in turn, provides a more dialogue-based interaction, which enables assessing cognitive and analytical processes, including key cognitive processes. Literature shows that, these technologies have complementary advantages but are conceptually disjointed. VR mainly records behavioral performance but GenAI emphasizes cognitive processes, and the two are not as integrated. Moreover, the present evidence base is typified by methodological heterogeneity, dependence on preliminary research, and lack of validity evidence, particularly regarding reliability and real-world impact. Other factors to take into account are effect of cognitive load, user variability, and ethical issues, e.g., algorithmic bias and transparency. The feasibility of implementation is further influenced by organizational factors such as infrastructure and institutional readiness. In general, VR and GenAI are encouraging, yet developing methods of assessment of clinical reasoning in oncology. In the future, the evolution of the field will require coherent assessment frameworks, rigorous validation, and clear ethical structures.

Introduction

Clinical reasoning, the cognitive and metacognitive processes underlying diagnosis and management, constitutes the cornerstone of safe and effective oncology practice1. It is a distinctively complicated process in the context of oncology, as clinicians must navigate probabilistic decision-making in the face of uncertainty, synthesize quickly changing evidence of therapies, and participate in multidisciplinary discussions in tumor boards. Moreover, oncologists have to find a balance between biomedical and patient-centered thought, including telling patients of their poor prognoses, dealing with ethical issues, and adjusting treatment plans to the changing paths of disease2,3. These layered demands make the development and assessment of clinical reasoning particularly critical in oncology education. Traditionally, clinical reasoning has been assessed using written examinations, Objective Structured Clinical Examinations (OSCEs), and workplace-based assessments such as the Mini-Clinical Evaluation Exercise (Mini-CEX)4. Although such methods can offer background information about the learner's competence, they have significant drawbacks. Written tests are more inclined to the application of static knowledge instead of dynamic processes; OSCEs are resource-demanding and might be ecologically invalid; as well as workplace-based tests are usually episodic and subjective and prone to contextual biasing. All these constraints together make it difficult to study the iterative, context-dependent, and longitudinal character of clinical reasoning, especially in specialties that involve high stakes and cognitive load, like oncology5. Collectively, these limitations hinder the ability to capture the iterative, context-dependent, and longitudinal nature of clinical reasoning, particularly in high-stakes and cognitively demanding specialties such as oncology6.

New developments in digital technologies in recent times present good opportunities to overcome these challenges. Two of them include Virtual Reality (VR) and Generative Artificial Intelligence (GenAI), which have become powerful medical education tools. VR provides simulation-based environments that are immersive, highly spatial and contextual, and that recreate the complex clinical scenario. In oncology education, these simulations may simulate scenarios such as oncologic emergencies, complications during chemotherapy, and sensitive patient interactions, and enable learners to practice decision-making in an experiential manner in a safe and controlled setting7,8,9. Importantly, VR systems can capture detailed behavioral data—such as action sequences, timing, and decision pathways—thereby enabling the assessment of procedural and situational dimensions of clinical reasoning.

Simultaneously, GenAI, especially large language models (LLMs) have opened up novel opportunities to evaluate the cognitive architecture of reasoning. GenAI systems have the capability to interact with learners through natural language to engage them in adaptative, dialogue-based settings that query generation of hypotheses, formation of differential diagnosis and justification of management. These systems can examine responses of learners based on techniques of natural language processing that provide a potentially personalized and scalable assessment of reasoning processes that have been previously challenging to measure objectively10,11,12.

Moreover, the existing literature is disjointed and heterogeneous in terms of methodology. Much of the existing evidence is based on pilot studies, technical reports, or applications that are mainly aimed at training as opposed to an actual evaluation. Consequently, most arguments on the evaluation abilities of VR and GenAI are based on technological functionality and not on the basis of stringent empirical validation. Reported evidence of validity is varied, with content validity predominantly reported and minimal reference to internal structure, reliability, or correlations with previously established assessment measures.

These constraints are further exacerbated by the new issues concerning usability, cognitive load, and ethical issues. As an example, GenAI systems provide an interactive solution that is scalable; however, the problems of information redundancy, system response variability, and familiarity of the user can affect the learner interaction and performance. Likewise, the lack of transparency, fairness, and accountability in the assessment scenarios can be considered a problem of the black box nature of AI-driven scoring. Implementation-wise, VR systems need a massive infrastructure investment, whereas GenAI adoption is based on the institutional readiness, governance, and alignment with the educational goals. These factors underscore the importance of a more conceptual and organizational framework in assessing how such technologies can be incorporated into oncology education.

In line with this, this review will offer an in-depth narrative synthesis of current literature on VR and GenAI in oncology education, with a specific interest in their application in evaluating clinical reasoning. Particularly, it aims to (1) describe technological designs and contexts of implementation of these tools; (2) review the evidence of the reported validity and educational outcomes; and (3) outline the main challenges, such as cognitive, ethical, and organizational concerns, and the future research and practice directions. Major academic databases were searched systematically to find pertinent literature, but as it is customary to review articles, this paper focuses on interpretation, synthesis, and not on procedural reporting.

Figure 1 integrated model combining VR-based assessment of clinical actions with GenAI-based assessment of clinical cognition. This dual approach captures both behavioral and analytic dimensions of clinical reasoning for comprehensive evaluation and feedback.

figure-introduction-1
Figure 1: Integrated Model of Clinical Reasoning Assessment Using VR and GenAI. This Figure shows an integrated model combining Virtual Reality (VR) for assessing action-based clinical reasoning (e.g., workflow, procedures, teamwork) with Generative AI (GenAI) for assessing cognitive and dialogic reasoning (e.g., diagnosis, justification, reflection). Please click here to view a larger version of this figure.

Access restricted. Please log in or start a trial to view this content.

Review and Perspective

Technology typology and implementation
The existing situation in the field of digital technologies of clinical reasoning evaluation in oncology education is marked by two paradigms of digital technologies. Current approaches rely on VR- and GenAI-based systems. Although both methods strive to improve the evaluation of complex clinical competencies, they have significant differences in terms of technological underpinnings, interaction, and assessment logic13,14. These systems put learners into realistic oncology settings, whether in an outpatient consultation, inpatient management of oncologic emergencies, or in a multidisciplinary decision-making context using head-mounted displays and interactive interfaces6.

An advantage of VR systems is that they can record granular behavioral data using embedded analytics. These data comprise a series of actions, a decision schedule, navigation, and compliance with clinical guidelines. These measures enable educators to assess the procedural elements of clinical reasoning, such as data obtaining, prioritization of diagnostic actions, and management pathways implementation15. However, implementing VR-based assessment tools is accompanied by significant infrastructural and logistical challenges. High initial costs associated with hardware acquisition, software development, and maintenance can limit accessibility, particularly in resource-constrained settings16. Additionally, the need for dedicated physical space, technical support, and faculty training further complicates large-scale deployment. These constraints often confine VR applications to specialized simulation centers, restricting their scalability and widespread integration into oncology curricula.

Conversely, GenAI-based tools are run via natural language interaction, providing a new evaluation method. Developed based on large language model architectures, they involve learners in interactive, text- or voice-based conversations, which mimic clinical reasoning mechanisms. The learners are encouraged to describe the differential diagnoses, the rationale behind management decisions and answer the changing clinical situations, with the system offering formative feedback and assessment10,11.

The main benefit of GenAI systems is that they can evaluate the cognitive and analytic aspects of clinical reasoning. These tools can assess the structure, coherence and depth of reasoning by a learner by looking at the language output and going beyond correctness to get qualitative information about the process of thinking. Furthermore, GenAI solutions have significant benefits of scalability and accessibility. They can also be implemented in remote and asynchronous learning contexts, unlike VR systems, which can only be implemented on web-based interfaces5.

In spite of the advantages, there are significant drawbacks to GenAI-based assessment tools. The dependence on natural language processing brings about issues related to scoring reliability, interpretability, and transparency. Many systems are black boxes, i.e., the logic of the scoring decision is not always available to the educators or learners17. The main theme that is apparent in the literature is the intrinsic trade-off between ecological validity and scalability between the two technological paradigms. VR systems are good in their ability to recreate realistic clinical settings and capture embodied decision-making, but are resource-intensive and thus limited in terms of scalability. GenAI systems on the other hand can be used to assess the cognitive processes in a scaleable and dialogue-based fashion7.

Critically, there is a significant lack of the existing literature that entails integrated or hybrid platforms that align the advantages of both VR and GenAI. This type of integration may facilitate concomitant measurement of procedural performance and cognitive reasoning, as the nature of clinical decision-making is interconnected18. From an implementation perspective, technical capabilities are not the only factors that determine technology adoption; organizational and environmental factors also play a role. Models like the Technology-Organization-Environment (TOE) model are a good perspective in explaining these dynamics19.

In general, the typology of VR and GenAI tools illuminates a discontinuous yet rapidly evolving space. Although both technologies have their unique benefits in measuring a particular aspect of clinical reasoning, their existing separation does not provide the opportunity to fully evaluate them. Closing this gap by means of conceptual integration, technological innovation, and context-sensitive implementation strategies will be necessary to take the field towards more comprehensive and effective assessment models.

Clinical reasoning constructs and iterative assessment
Clinical reasoning is a multidimensional concept that involves cognitive, metacognitive, and behavioral processes that support diagnostic and therapeutic decision-making. In oncology, uncertainty, dynamic evidence, and the need for interdisciplinary coordination further complicate these processes. The literature is increasingly conceptualizing clinical reasoning as a continuum that entails the acquisition of information, the generation of hypothesis, the refinement of diagnosis, the decision-making process and the reflective evaluation20,21.

The VR-based systems are mostly consistent with the evaluation of the observable and action-oriented aspects of clinical reasoning. Placing learners in simulated clinical settings, these platforms allow testing how people acquire knowledge, focus on diagnostic processes, and perform a series of operations. It focuses on the translation of reasoning into action, in which the clinical decisions are based on patterns of interaction in the simulated environment22.

GenAI-based systems, in contrast, concentrate on the processes of eliciting and analyzing internal cognitive processes underlying clinical reasoning. Natural language interaction encourages learners to express the different diagnosis, to justify clinical management choices, and to ponder clinical uncertainties. Such systems determine reasoning by measuring semantic coherence, logical structure, and depth of responses by the learner, typically by applying natural language processing methods to provide an approximation of expert judgment13,10.

Although these are complementary strengths, a weakness that has been found to be most critical throughout the literature is the adversarial fragmentation of assessment between cognitive and behavioral domains. A real-life practice of clinical reasoning is inherently iterative, which means constant feedback between thinking and action. This is a key cyclic process in an effective clinical decision-making process, especially in oncology, whereby the conditions and reactions of the patients to the treatment change over time11.

This dynamic interaction is often not reflected in current VR and GenAI tools. VR systems generally derive reasoning from observable behavior without necessarily reaching the underlying mental processes, whereas GenAI systems analyze articulated reasoning outside the context of clinical behavior. Consequently, both strategies present biased images of clinical competence and cannot adequately evaluate the entire range of reasoning activities. The other significant weakness associated with the use of automated analytics as the main evaluation tool. In VR, the analysis of log-files is employed to understand the behavior of users, but these measurements might not be sufficient to understand the reasoning behind clinical choices. Equally, algorithmic interpretation of language is commonly used to score GenAI systems and could ignore language nuances like clinical intuition, uncertainty handling, or context-specific reasoning23.

Besides, the literature notes the absence of standardized methods of mapping technology outputs to known models of clinical reasoning. Although there are studies that implicitly connect VR to procedural reasoning and GenAI to cognitive reasoning, not many studies place these tools in the context of validated theoretical constructs. This disparity restricts the possibilities of making comparisons on the results of other studies and forming a consistent structure of technology-enhanced assessment24.

The focus of future research in this area should be on integrating cognitive and behavioral modalities of assessment. Combination Hybrid systems Hybrid systems that integrate immersive simulation with real-time, AI-mediated dialogue can better embrace the iterative character of clinical reasoning. With these, the behavior of learners in a virtual environment may prompt a questioning response on the part of an AI agent, and thus allow assessing both what learners do and why they do it at the same time25,26.

Altogether, although VR and GenAI technologies provide useful insights into specific aspects of clinical reasoning, their application remains disjointed at the moment. Integration of concepts, methodological rigor, and technological innovation will be needed to address this fragmentation and develop comprehensive, meaningful assessment strategies in oncology education. Table 1 highlights key differences between Virtual Reality (VR) and Generative AI (GenAI) for assessing clinical reasoning, comparing their respective foci, interactions, strengths, limitations, and data outputs. Moreover, Figure 2 illustrates the iterative clinical reasoning cycle, which shows how VR supports data collection, action, and reflection, while GenAI aids hypothesis generation, refinement, and diagnostic explanation. Together, they enable a continuous loop of informed decision-making and experiential learning.

DimensionVirtual Reality (VR)Generative AI (GenAI)
Assessment focusProcedural / behavioralCognitive / analytical
Interaction modeImmersive simulationDialogue-based
StrengthHigh ecological validityHigh scalability
LimitationResource-intensiveLimited transparency
Data typeBehavioral logsNatural language responses

Table 1: Comparison of VR vs GenAI in Clinical Reasoning Assessment. Comparison of Virtual Reality (VR) and Generative AI (GenAI) for clinical reasoning assessment across key dimensions, including assessment focus, interaction mode, strengths, limitations, and data type.

figure-protocol-1
Figure 2: Iterative Clinical Reasoning Cycle: Contributions of VR and GenAI. The iterative clinical reasoning cycle illustrates how VR captures actions and performance during simulation-based tasks (information gathering and decision-making), while GenAI supports hypothesis generation, differential diagnosis, explanation, and refinement, together enabling continuous reflection and learning from experience. Please click here to view a larger version of this figure.

Validity evidence and measurement limitations
Developing strong validity evidence is core to implementing any assessment tool, especially in high-stakes areas like oncology education. Modern conceptualization of validity theory views validity as a single construct that is supported by several sources of evidence such as content, response process, internal structure, relations to other variables, and consequences of assessment27,28.

The majority of studies focus on content validity, which is typically demonstrated through expert review of clinical situations, task design, or scoring criteria. Although these methods ensure that the assessment material is pertinent and consistent with oncology practice, they do not provide much information about the tools' ability to measure the underlying constructs of clinical reasoning. In most situations, the validation of the content is regarded as adequate evidence of effectiveness, even though more rigorous psychometric practices are lacking29.

There is less consistent reporting of evidence on response process validity- how learners perceive and interact with assessment activities. Other studies use usability testing, learner feedback, or think-aloud protocols to imply that participants are engaging with VR environments or GenAI systems in a manner desired. In addition, it is not yet clear to what degree these interactions can accurately capture true clinical reasoning, especially in simulated or AI-mediated scenarios30.

One of the most important constraints on both VR and GenAI modalities is the lack of internal-structure evidence. Measures of reliability, including internal consistency, inter-rater agreement, or test-retest stability, are seldom reported. This is particularly troublesome for tools that are based on automated scoring. GenAI systems are further variable in terms of scoring algorithms that are driven by natural language processing, where results can vary depending on prompt design, model parameters, or contextual interpretation31.

Likewise, there is limited evidence of relations to other variables. Not many studies have sought to relate performance on VR or GenAI platforms to existing assessment techniques, including OSCEs, written tests, or workplace assessments. Without such comparisons, it becomes hard to know whether these technologies measure constructs consistent with current frameworks of clinical competence. Such a deficiency in convergent and discriminant validity limits the interpretation and believability of assessment results32.

The greatest gap in the literature, however, is its lack of transfer validity. Little to no data support the concept that practice in VR simulations or GenAI-based tests is correlated with better clinical reasoning in the actual oncology practice. Since the ultimate aim of medical training is to improve patient care, this restriction is one of the biggest obstacles to implementing such tools in formal assessment programs. In the absence of longitudinal or outcome studies, the educational effect of these technologies is limited to the simulated or experimental situation in which they are implemented33.

Along with these psychometric drawbacks, there are other challenges posed by the growing use of automated analytics. Both VR and GenAI rely heavily on algorithmic interpretation of user behavior or language, raising questions about transparency and interpretability. The black-box nature of large language models in the context of GenAI-based assessments complicates understanding how scores are produced and how to guarantee unbiased scores. Such inexplicability may destroy the trust between teachers and students, especially in high-stakes environments34.

Another essential issue is bias in AI-based evaluation. The training data used to create language models might capture differences in clinical knowledge, communication patterns, or cultural practices, thereby leading to systematic differences in scoring across learner groups. Unaddressed, such biases may have serious consequences to the aspects of fairness and equity in the medical education35.

In addition, the implications of assessment, another important aspect of validity, are rarely discussed in the existing literature. Very little research examines the effects of VR or GenAI tool use on learners' behavior, educational paths, or decision-making in clinical environments. Lack of studies on the downstream effects restricts knowledge on the positive and the possible unintended effects of these technologies36.

To conclude, VR and GenAI technologies are promising for providing new methods to evaluate clinical reasoning, but the existing evidence base is insufficient to implement them in high-stakes situations. In all areas, enhancing validity will ensure the next step: transforming these tools from experimental use to effective elements of oncology education evaluation. Table 2 summarizes validity evidence for VR and GenAI across four domains—content, internal structure, relations to other variables, and consequences—while highlighting key limitations, including expert-driven designs, limited reliability testing, sparse comparative studies, and the absence of long-term outcome data.

Validity DomainVR EvidenceGenAI EvidenceKey Limitation
ContentStrongModerateMostly expert-driven
Internal structureLimitedLimitedLack of reliability testing
Relations to other variablesSparseSparseFew comparative studies
ConsequencesMinimalMinimalNo long-term outcomes

Table 2: Validity Evidence Across Technologies. This table shows validity evidence for VR and GenAI in clinical reasoning assessment across content, internal structure, relations to other variables, and consequences domains

Educational outcomes and interpretive caution
The increased adoption of Virtual Reality (VR) and Generative Artificial Intelligence (GenAI) in oncology education has been met with the literature that documents overall positive education results. These are improved learner engagement, motivation, perceived realism of learning environments, and the ability to receive immediate, personalized feedback37,38,39. Such findings suggest that these technologies hold promise for enriching educational experiences and supporting the development of clinical reasoning skills.

A significant percentage of the reported results are self-reported like learner satisfaction, perceived confidence, or subjective learning evaluation. Although these indicators serve as insightful information about user experience, they are not always associated with any tangible clinical reasoning or clinical performance improvements40. In the case of objective reporting, they tend to concentrate on the results of performance on the technological platform itself. Though these results can suggest that the tool was successfully engaged with, they do not determine whether such interaction can be further translated into competence in general in clinical reasoning. The gains realized in controlled or simulated settings can be associated with familiarity with the system as opposed to real cognitive or behavioral growth. The other limitation is associated with the methodological features of the existing studies. Much of the literature comprises small-scale, single-institution studies that have limited sample sizes and short intervention periods. These limitations decrease statistical power and restrict the generalizability of the results to a wide range of educational settings33.

The transferability problem is of special concern in cancer education. Oncological clinical reasoning is a complex and longitudinal decision-making process that occurs within a time frame and within clinical contexts. Nonetheless, there is a lack of literature on whether skills trained in VR or GenAI interventions remain after the immediate learning context or affect the real clinical practice. Moreover, the literature tends to mix training and assessment results. Most VR and GenAI applications have the primary aim of learning facilitation as an educational tool, but their success is occasionally viewed as a sign of assessment validity. To interpret accurately, it is important to differentiate the benefits of formative learning and utility of summative assessment42.

There are also cognitive reasons that make the interpretation of educational outcomes more complicated. Advanced technologies present a range of cognitive loads, and it may affect the performance of learners. As an example, immersive VR systems can have extraneous cognitive load in terms of navigation, interacting with the interface, or sensory processing. With the right design and familiarity with the user, these factors may either be facilitating or hinder learning36. The user experience and previous experience with technology are also factors. Students who are more accustomed to computer interfaces or AI-driven systems can perform better, not necessarily because they have more clinical reasoning, but because of a more convenient interaction with the system43.

The possibility of unintended consequences is another factor that should be taken into account. To illustrate, excessive dependence on AI-developed feedback can decrease chances of logical thinking or critical analysis. Likewise, well-articulated VR situations can unintentionally restrict decision-making by reducing the number of potential actions. The current literature does not extensively investigate these effects yet they have significant implications on educational design44.

Although these restrictions exist, the positive tendencies that were recorded in terms of learner engagement and experiential learning cannot be neglected. Safe practice environments, real-time feedback, and the ability to be exposed to intricate situations again and again are affordances of VR and GenAI technologies. These characteristics can be associated with the current educational theories, such as experiential learning and deliberate practice, which prioritize active learning and continuous enhancement45.

Overall, although VR and GenAI technologies show a significant promise to transform oncology education, the evidence base is currently limited in methodology and interpretation. A critical and cautious attitude is thus necessary so as to prevent overgeneralization of the findings and to be sure that claims about the effectiveness of education have strong empirical evidence.

Cognitive, ethical, and organizational considerations
Cognitively, the development of these technologies and their application may have a significant impact on how learners perceive information and practice clinical reasoning. Cognitive load theory offers a practical framework of understanding these effects with three types of cognitive load namely, intrinsic, extraneous and germane cognitive load43. Although VR environments are immersive and engaging, they can place a significant amount of extraneous cognitive load on users, requiring navigation, sensory stimulation, and the complexity of the interface. Unless well planned, such factors may shift the cognitive resources to other areas other than essential reasoning processes, thus negating the learning and assessment performance38.

The same way, GenAI systems present new cognitive challenges regarding information presentation and dynamics of interaction. An example is that excessive or redundant feedback produced by AI systems can lead to information overload, which can impede successful learning. On the other hand, properly designed, adaptive feedback can augment the germane cognitive load through schema formation, and reflective thought46. The balance between guidance and cognitive autonomy is therefore critical, as excessive reliance on AI-generated responses may reduce opportunities for independent problem-solving and critical thinking.
Cognitive engagement with such technologies is also mediated by user familiarity and digital literacy44.

Ethics is also a key aspect in the implementation of VR and GenAI in assessing education. Among the most salient issues is the concept of algorithmic bias in AI-based systems. Machine learning models, such as large language models, are trained on large amounts of data that may be biased towards existing social, cultural, or professional norms. Consequently, assessment results can lead to unwanted discrimination against specific groups of learners, and fairness and equity are called into question35. Other ethical issues are transparency and explainability. Most GenAI systems are black boxes, with no explicit descriptions of the decision-making process. This interpretability in assessments may cause a lack of trust between learners and educators, especially when the end product has serious academic implications34.

Security and privacy of the data is also a significant factor to consider as educational and, possibly, clinical data is sensitive. VR can be used to record detailed behavioral data, whereas GenAI platforms can handle large quantities of textual data. To uphold ethical standards and trust in the institution, it is important to ensure that data protection rules are followed and that the information of the users remains protected47.

In addition to individual-level issues, organizational issues are decisive in the successful adoption of these technologies. Implementing VR and GenAI in oncology education needs to be consistent with the aims of the institutions, access to resources, and the supportive governance frameworks. The Technology-Organization-Environment (TOE) model is a valuable tool to examine these factors and focus on the interaction of technological preparedness, organizational capabilities, and the impact of the external environment48.

The organizational readiness includes faculty competence, training, and the acceptance of new technologies. The possible resistance to adoption can be related to the fears of reliability, work overload, or simply the unfamiliarity with AI-driven systems. These constraints can be overcome through focused faculty development programs and by having explicit rules on how technology is used in assessment49.

A new approach to the field focuses on the idea of enlightened autonomy, in which AI systems are created to aid and not substitute human judgment. This method proposes the harmonious combination of automated analytics and human control where technology would improve decision-making without interfering with accountability or professionalism50. In the context of clinical reasoning assessment, this may involve combining AI-generated insights with expert evaluation to achieve a more comprehensive and trustworthy assessment process.

Overall, the introduction of VR and GenAI into oncology education involves more than mere technical innovation; it also explores complex cognitive, ethical, and organizational aspects. These considerations are crucial in ensuring that these technologies are maximized in terms of their educational value and reduced risks. An interdisciplinary strategy that considers the user-centered design, ethical concerns, and institutional preparedness will be paramount to the development of the responsible use of technology in clinical reasoning assessment. Figure 3 Technology-Organization-Environment (TOE) framework identifying key implementation factors for integrating VR and GenAI into oncology education. Successful adoption requires alignment across system capabilities, organizational support, and external environmental conditions.

figure-protocol-2
Figure 3: Technology-Organization-Environment (TOE) Framework for Implementing VR and GenAI in Oncology Education. The Technology-Organization-Environment (TOE) framework applied to integrating VR and GenAI in oncology education, highlighting key factors across system capability, leadership and curriculum, as well as regulatory, funding, and cultural considerations for successful implementation. Please click here to view a larger version of this figure.

Access restricted. Please log in or start a trial to view this content.

Conclusions

Clinical reasoning in oncology is an expanding and challenging issue that has long been a complex, dynamic, high-stakes, and context-specific aspect of oncological decision-making. This review has discussed the new functions of Virtual Reality (VR) and Generative Artificial Intelligence (GenAI) as innovative solutions to the limitations of conventional assessment techniques. Taken together, the evidence suggests that these technologies offer non-overlapping advantages but have yet to achieve conceptual fragmentation, ove...

Access restricted. Please log in or start a trial to view this content.

Acknowledgements

Funding: Fund: General Research Project (Natural Science Category) of Zhejiang Provincial Department of Education(Grant No.Y202558353).

Fund: Research Project on Higher Education Special Project: Artificial Intelligence Empowering Education and Teaching Applications(KT2025426)

Access restricted. Please log in or start a trial to view this content.

References

  1. Shimizu, H., Nakayama, K. I. Artificial intelligence in oncology. Cancer Sci. 111 (5), 1452-1460 (2020).
  2. Ginsburg, K. B., Curtis, G. L., Timar, R. E., George, A. K., Cher, M. L. Delayed radical prostatectomy is not associated with adverse oncologic outcomes: implications for men experiencing surgical delay due to the COVID-19 pandemic. J Urol. 204 (4), 720-725 (2020).
  3. Egbuna, I., Olatokun, T., Ozo-Ogueji, P., Ekechi, C., Akinbo, O., et al. Revolutionizing cancer surgery: Harnessing artificial intelligence and augmented reality for next-generation precision oncology. Int J Life Sci Res Arch. 9 (1), 147-176 (2025).
  4. Sminia, P., et al. Clinical radiobiology for radiation oncology. Radiobiology textbook. , (2023).
  5. Wood, E. A., Ange, B. L., Miller, D. D. Are we ready to integrate artificial intelligence literacy into medical school curriculum? Students and faculty survey. J Med Educ Curric Dev. 8, 23821205211024078 (2021).
  6. Kok, D. L., Dushyanthen, S., Peters, G., Sapkaroski, D., Barrett, M., et al. Virtual reality and augmented reality in radiation oncology education: A review and expert commentary. Tech Innov Patient Support Radiat Oncol. 24, 25-31 (2022).
  7. Rahimi, F., Sadeghi-Niaraki, A., Choi, S. M. Generative AI meets virtual reality: A comprehensive survey on applications, challenges, and future directions. IEEE Access. 13, 1-20 (2025).
  8. von Ende, E., Ryan, S., Crain, M. A., Makary, M. S. Artificial intelligence, augmented reality, and virtual reality advances and applications in interventional radiology. Diagnostics (Basel). 13 (5), 892 (2023).
  9. Esmail, S., Concannon, B. Immersive virtual reality and AI (generative pretrained transformer) to enhance student preparedness for objective structured clinical examinations: Mixed methods study. JMIR Serious Games. 13, e69428 (2025).
  10. Gilson, A., Safranek, C. W., Huang, T., Socrates, V., Chi, L., et al. How does ChatGPT perform on the United States Medical Licensing Examination (USMLE)? The implications of large language models for medical education and knowledge assessment. JMIR Med Educ. 9 (1), e45312 (2023).
  11. Kung, T. H., Cheatham, M., Medenilla, A., Sillos, C., De Leon, L., et al. Performance of ChatGPT on USMLE: Potential for AI-assisted medical education using large language models. PLoS Digit Health. 2 (2), e0000198 (2023).
  12. Xu, L., Sanders, L., Li, K., Chow, J. C. Chatbot for health care and oncology applications using artificial intelligence and machine learning: Systematic review. JMIR Cancer. 7 (4), e27850 (2021).
  13. Pottle, J. Virtual reality and the transformation of medical education. Future Healthc J. 6 (3), 181-185 (2019).
  14. Farina, E., Nabhen, J. J., Dacoregio, M. I., Batalini, F., Moraes, F. Y. An overview of artificial intelligence in oncology. Future Sci OA. 8 (4), FSO787 (2022).
  15. Hamilton, A. Artificial intelligence and healthcare simulation: The shifting landscape of medical education. Cureus. 16 (5), e00000 (2024).
  16. Siddiqui, F. M., Jabeen, S., Alwazzan, A., Vacca, S., Dalal, L., et al. Integration of augmented reality, virtual reality, and extended reality in healthcare and medical education: A systematic review. J Med Educ Curric Dev. 12, 23821205251342315 (2025).
  17. Wang, S., Zhang, H. Industrial information integration through autonomous AI agents: Paradoxical effects on transparency, dehumanization, and responsible operations. J Ind Inf Integr. 51, 101089 (2026).
  18. Tortora, A., Amaro, I., Della Greca, A., Barra, P. Exploring the role of generative artificial intelligence in virtual reality: Opportunities and future perspectives. , 125-142 (2025).
  19. Wang, S., Zhang, H. Generative AI in international hotel marketing: Impacts on employee creativity and performance. Int J Contemp Hosp Manag. 37 (8), 2601-2626 (2025).
  20. Christensen, N., Jones, M. A., Rivett, D. A., Jones, M. A., Rivett, D. A. Strategies to facilitate clinical reasoning development. Clinical reasoning in musculoskeletal practice. , 562-582 (2019).
  21. Cook, D. A., Sherbino, J., Durning, S. J. Management reasoning: Beyond the diagnosis. JAMA. 319 (22), 2267-2268 (2018).
  22. Croskerry, P. A universal model of diagnostic reasoning. Acad Med. 84 (8), 1022-1028 (2009).
  23. Norman, G. Research in clinical reasoning: Past history and current trends. Med Educ. 39 (4), 418-427 (2005).
  24. Topol, E. J. High-performance medicine: The convergence of human and artificial intelligence. Nat Med. 25 (1), 44-56 (2019).
  25. Durning, S. J., Artino, A. R., Pangaro, L. N., van der Vleuten, C., Schuwirth, L. Redefining context in the clinical encounter: Implications for research and training in medical education. Acad Med. 85 (5), 894-901 (2010).
  26. Bickmore, T. W., Giorgino, T. Health dialog systems for patients and consumers. J Biomed Inform. 39 (5), 556-571 (2006).
  27. Messick, S. Validity. educational measurement. , 13-103 (1989).
  28. Cook, D. A., Beckman, T. J. Current concepts in validity and reliability for psychometric instruments: Theory and application. Am J Med. 119 (2), 166.e7-166.e16 (2006).
  29. Downing, S. M. Validity: On meaningful interpretation of assessment data. Med Educ. 37 (9), 830-837 (2003).
  30. Artino, A. R., La Rochelle, J. S., Dezee, K. J., Gehlbach, H. Developing questionnaires for educational research: AMEE Guide No. 87. Med Teach. 36 (6), 463-474 (2014).
  31. Gwet, K. L. . Handbook of inter-rater reliability. , (2021).
  32. Norcini, J., Anderson, M. B., Bollela, V., Burch, V., Costa, M. J., et al. Criteria for good assessment: Consensus statement and recommendations from the Ottawa 2010 Conference. Med Teach. 33 (3), 206-214 (2011).
  33. Cook, D. A., Hatala, R., Brydges, R., Zendejas, B., Szostek, J. H., et al. Technology-enhanced simulation for health professions education: A systematic review and meta-analysis. JAMA. 306 (9), 978-988 (2011).
  34. Topol, E. J. . Deep medicine: How artificial intelligence can make healthcare human again. , (2019).
  35. Obermeyer, Z., Powers, B., Vogeli, C., Mullainathan, S. Dissecting racial bias in an algorithm used to manage the health of populations. Science. 366 (6464), 447-453 (2019).
  36. van der Vleuten, C. P. M., Schuwirth, L. W. T. Assessing professional competence: From methods to programmes. Med Educ. 39 (3), 309-317 (2005).
  37. Radianti, J., Majchrzak, T. A., Fromm, J., Wohlgenannt, I. A systematic review of immersive virtual reality applications for higher education: Design elements, lessons learned, and research agenda. Comput Educ. 147, 103778 (2020).
  38. Makransky, G., Petersen, G. B. Immersive virtual reality and learning: A meta-analysis. Educ Psychol Rev. 31 (4), 915-943 (2019).
  39. Luckin, R., Holmes, W., Griffiths, M., Forcier, L. B. . Intelligence unleashed: An argument for AI in education. , (2016).
  40. Sitzmann, T., Ely, K., Brown, K. G., Bauer, K. N. Self-assessment of knowledge: A cognitive learning or affective measure?. Acad Manag Learn Educ. 9 (2), 169-191 (2010).
  41. Reeves, S., Fletcher, S., Barr, H., Birch, I., Boet, S., et al. A BEME systematic review of the effects of interprofessional education. Med Teach. 38 (7), 656-668 (2016).
  42. Issenberg, S. B., McGaghie, W. C., Petrusa, E. R., Lee Gordon, D., Scalese, R. J. Features and uses of high-fidelity medical simulations that lead to effective learning: A meta-analysis. Med Teach. 27 (1), 10-28 (2005).
  43. Sweller, J. Cognitive load theory. Psychol Learn Motiv. 55, 37-76 (2011).
  44. Liaw, S. S., Huang, H. M. Perceived satisfaction, perceived usefulness and interactive learning environments as predictors to self-regulation in e-learning environments. Comput Educ. 60 (1), 14-24 (2013).
  45. Carr, S. AI gone wild: Implications of artificial intelligence for education. Educ Technol. 60 (2), 3-10 (2020).
  46. Wang, S., Sun, Z. Roles of artificial intelligence experience, information redundancy, and familiarity in shaping active learning. Educ Inf Technol. 30 (2), 2525-2546 (2025).
  47. European Parliament and Council. General Data Protection Regulation (GDPR). Off J Eur Union. L119, 1-88 (2016).
  48. Tornatzky, L. G., Fleischer, M. . The processes of technological innovation. , (1990).
  49. Rogers, E. M. . Diffusion of innovations. , (2003).
  50. Floridi, L., Cowls, J. A unified framework of five principles for AI in society. Harv Data Sci Rev. 1 (1), 1-15 (2019).

Access restricted. Please log in or start a trial to view this content.

Reprints and Permissions

Tags

Virtual Reality OncologyGenerative AI EducationSimulation Based AssessmentOncology Education ToolsCognitive LoadAlgorithmic BiasObjective Structured ExaminationsWorkplace Based AssessmentAssessment Frameworks