Search approach
This article was developed as a narrative review and perspective rather than a systematic review, scoping review, or meta-analysis. We searched PubMed, PsycINFO, Web of Science Core Collection, and ERIC from database inception to 31 December 2025. Search terms combined concepts related to sleep deprivation, sleep restriction, fatigue, decision fatigue, clinical performance, competency-based medical education (CBME), cognitive workload, electroencephalography (EEG), functional near-infrared spectroscopy (fNIRS), brain-computer interfaces (BCIs), and medical training. A targeted update was also performed to incorporate recent fNIRS literature relevant to sleep and fatigue monitoring.
Sources were selected and prioritized when they directly informed one of four areas: sleep- or fatigue-related changes in neurocognitive and clinical performance; EEG- or fNIRS-derived markers of fatigue, workload, vigilance, or sleep-related vulnerability; conceptual or applied work relevant to integrating physiological state information into CBME; and ethical or implementation issues relevant to educational use. Sources were not selected to provide exhaustive coverage of all sleep, BCI, or neurotechnology literature. We did not prioritize studies that were primarily basic neuroscience, device engineering, or therapeutic neuromodulation papers unless they directly informed fatigue-aware interpretation of trainee performance. No formal risk-of-bias assessment or meta-analytic pooling was performed; findings were synthesized descriptively and critically to support a conceptual, hypothesis-generating perspective framework.
Novel contribution and intended audience
The novelty of this review lies in connecting several related but often separately discussed bodies of literature—resident fatigue, sleep deprivation, decision fatigue, cognitive workload, neurophysiological monitoring, and CBME assessment—around a specific educational problem: how educators should interpret observed clinical performance when a trainee’s cognitive state may be temporarily altered by sleep loss, decision fatigue, or workload. The intended audience is medical educators, clinician-educators, program directors, and researchers interested in fatigue-aware interpretation of trainee performance. The proposed CBME-BCI-Fatigue framework is intended to organize future research questions and clarify validation requirements, not to validate an operational monitoring system or support current high-stakes assessment decisions.
Competence as dynamic performance in CBME
The original rationale for CBME was not simply to create new assessment forms, but to align medical education with the needs of patients and health systems. CBME emphasizes progression based on proficiency, individualized learning trajectories, and evidence that learners can perform meaningful clinical activities with appropriate supervision5,6. National and regional competency frameworks, including China-based residency competency frameworks and i-FIRST talent-development discussions, provide structured language for curriculum design, learner development, and educational accountability7,8. At the same time, literature on CBME implementation notes persistent concerns about feasibility, faculty development, assessment burden, and the risk of reducing complex professional practice to checklist behavior9,10.
EPAs and entrustment decisions have been particularly influential because they connect assessment to the clinical question of how much supervision a learner requires for a defined professional activity11,12,13,14. EPA implementation also highlights that entrustment is a judgment about a learner, a task, a context, and a level of supervision; it is not simply a fixed trait embedded in the trainee15. Similarly, learning curves and learning analytics can help educators identify trajectories over time rather than overinterpreting single data points16,17. Precision education, education passports, mastery-oriented learning environments, and coaching further reinforce the idea that assessment should support development, feedback, and adaptive support18,19,20,21.
A state-dependent perspective extends this logic. If a trainee performs poorly after prolonged wakefulness or a cognitively intense shift, the assessment question is not only whether the trainee lacks competence. It is also whether the observed behavior reflects fatigue-related vulnerability, excessive workload, poor supervision, system design, or a mismatch between task demands and cognitive state. A stable competence gap may require deliberate practice and coaching; a fatigue-related lapse may require recovery planning, workload redesign, supervision adjustment, or schedule-level intervention. State-aware assessment therefore does not redefine competence. Rather, it improves the context in which performance data are interpreted and reduces the risk of misclassifying transient impairment as durable inability.
Sleep deprivation, circadian disruption, and decision fatigue
Sleep deprivation and fatigue are pervasive yet underappreciated determinants of physician performance. In clinical training, extended-duration shifts, night work, unpredictable workload, and frequent handoffs expose trainees to conditions that can degrade alertness and decision-making. Studies of physician sleep and wellbeing have linked sleep and burnout-related distress with clinically significant medical errors23. Extended-duration shifts have also been associated with increased attentional failures, medical errors, and adverse events24. Resident fatigue and distress may further influence self-reported medical errors, suggesting that performance risk is partly behavioral, partly cognitive, and partly system-level25.
Observational and experimental studies in medical trainees and physicians reinforce this concern. Interns and residents working extended shifts show reduced sleep opportunity and altered alertness patterns26. Reducing intern work hours in intensive care units was associated with fewer serious medical errors in a landmark trial27. Surgical and procedural studies show that sleep deprivation may impair simulator dexterity and crisis-management performance28,29,30. Patient-level and procedure-level studies remain more heterogeneous and should not be oversimplified31,32,33. Differences across specialty, training level, case complexity, supervision, team structure, and outcome type may explain why fatigue-related decrements are clearer for vigilance, simulated performance, and process measures than for downstream patient outcomes.
Duty-hour trials illustrate why fatigue-aware education cannot rely only on shift length. The iCOMPARE trial found that flexible duty-hour policies were noninferior to standard policies for selected education outcomes and patient safety outcomes34,35. Companion analyses reported that sleep and alertness did not necessarily worsen under flexible duty-hour rules36. In contrast, a randomized trial of schedules without 24-hour shifts showed that eliminating extended shifts did not uniformly improve patient safety37. A related study of extended work shifts found neurobehavioral performance effects in resident physicians38. Taken together, these findings show that sleep opportunity, workload, continuity, handoffs, and supervision interact in complex ways.
The broader sleep science literature provides mechanistic context. Sleep disruption impairs sustained attention, working memory, response inhibition, emotion regulation, and complex task performance39,40,41,42. Irregular or unstable sleep schedules may also affect academic and cognitive outcomes, suggesting that sleep stability matters in addition to sleep duration43,44. Chronic sleep restriction with weekend recovery may leave residual effects on performance and wellbeing45. In medical education, these findings are important because training assessments are usually conducted during wakefulness but rarely document the learner state in which performance is expressed.
Decision fatigue adds a related but distinct mechanism. It refers to a deterioration in decision quality after repeated or effortful choices, particularly when decisions require self-control, uncertainty management, or value tradeoffs46. Clinical decisions often involve triage, prescribing, escalation of care, documentation, and handoffs under time pressure. Recent prescribing research suggests that time-on-task and accumulated decision demands can affect clinical choices47. Fatigue from overnight shifts can also alter search patterns and diagnostic performance in radiology48. However, decision fatigue should be interpreted cautiously because the literature is methodologically heterogeneous and effects may vary across tasks, environments, and clinicians.
EEG and fNIRS markers of fatigue and cognitive workload
EEG has been widely used to study vigilance, sleep pressure, fatigue, and mental workload. In physician studies, quantitative EEG after an overnight on-call duty has shown changes in theta, alpha, and beta activity consistent with mental fatigue49,50. These studies are important because they bring neurophysiological assessment closer to real clinical work, rather than relying solely on laboratory sleep-deprivation paradigms. However, they remain observational and cannot by themselves establish that EEG can determine whether a trainee is ready to perform a particular clinical task.
Feature-level EEG interpretation is broader than the commonly cited changes in theta or alpha/beta. Frontal EEG workload studies support the use of spectral and regional patterns for workload estimation51. Sleep-deprivation EEG studies have examined changes in spectral activity after sleep loss52. The theta band also has a dual interpretation: it may reflect sleep pressure, but can also participate in cognitive control, a point that complicates simple fatigue classification53. Real-time working-memory monitoring studies and driver-state classification work further show how EEG features can be embedded in BCI-style pipelines, but these pipelines remain sensitive to task design, preprocessing, and individual variability54,55. Reliability studies and neuroergonomic validation work emphasize that electrode configuration, preprocessing choices, subjective metrics, and task design influence workload indices56,57. More recent multilevel cognitive-load prediction work suggests that combining spatial, temporal, and spectral signatures can improve classification58.
Functional near-infrared spectroscopy fNIRS provides complementary information by measuring hemodynamic changes, including oxygenated hemoglobin, deoxygenated hemoglobin, and total hemoglobin. The fNIRS literature on fatigue and sleep deprivation supports its relevance for measuring cortical correlates of cognitive effort, reduced alertness, and social-cognitive changes59. Sleep-focused fNIRS reviews further show how optical hemodynamic signals may contribute to sleep research and fatigue detection60. A recent 2026 study on cardiac-band fNIRS biomarkers for non-rapid eye movement sleep differentiation with EEG validation illustrates the growing interest in frequency-domain and cardiac-band features beyond simple HbO amplitude61. Together, this literature supports a feature-level discussion that includes HbO, HbR, HbT, time-domain responses, connectivity, and frequency-domain biomarkers.
Prefrontal fNIRS studies are particularly relevant because executive control, response inhibition, and workload regulation depend heavily on frontal systems. Sleep deprivation has been associated with altered task-related functional connectivity in the frontal cortex62. Related work has found loss of frontal functional connectivity after sleep deprivation63. Studies in short-sleep young adults also suggest that prefrontal hemodynamic responses can vary with sleep behavior and physical activity64. Wearable fNIRS has been used for automatic cognitive-fatigue detection with machine learning65. Hybrid EEG-fNIRS systems can improve workload classification by integrating functional brain connectivity and complementary signal properties66,67. Collectively, these studies support the use of EEG and fNIRS as research tools and as candidate contextual indicators, but not yet as validated readiness measures in medical training.
Table 1 is therefore intended as a feature-level map rather than a hierarchy of validated biomarkers. It distinguishes EEG spectral, event-related, temporal, and connectivity features from fNIRS amplitude-based, time-domain, connectivity, and frequency-domain or cardiac-band features. This distinction is important because each feature family captures different dimensions of the fatigue workload construct and carries different artifact and interpretation risks. For example, a higher theta/alpha ratio, reduced P300 amplitude, or lower prefrontal HbO may be compatible with fatigue, but may also reflect task difficulty, compensatory effort, stress, sensor configuration, or individual baseline differences. The table is intended to support study design, construct alignment, and cautious interpretation, not to define operational thresholds for competence, entrustment, or clinical readiness.
BCI, multimodal fusion, and educational applications
Non-invasive BCIs and wearable neurophysiological technologies have expanded rapidly across medical, engineering, and educational domains. Reviews of wearable EEG-based BCI devices highlight progress in portability and medical applications68. Broader state-of-the-art reviews emphasize that non-invasive BCI systems are improving but still face limitations in signal quality, usability, calibration, and robustness69,70,71,72. These constraints matter for medical education because an educational tool must be not only accurate in a controlled setting, but also feasible, acceptable, interpretable, and aligned with educational constructs.
The most defensible educational applications are simulation, research, and low stakes feedback. VR laparoscopic training can synchronize performance data with workload related signals to study how cognitive demand relates to procedural learning73. Simulated neurosurgical tasks have used AI and EEG to distinguish expertise levels, suggesting a pathway for studying skill acquisition rather than replacing faculty judgment74. Serious-game and mixed-reality environments have used biosensor analytics to estimate affective engagement75. Passive BCI systems that adapt learning speed to cognitive load illustrate how state-sensitive feedback might operate in lower-risk educational contexts76. Ambulatory sleep detection work shows how multimodal time-series data can support longitudinal monitoring outside the laboratory77. These examples are best interpreted as supporting evidence or proof-of-concept, not as validation for high-stakes entrustment or remediation decisions.
AI-assisted sleep staging studies show how machine-learning methods can classify sleep-related physiological time series78,79,80. Targeted memory reactivation, wearable neuromodulation, fatigue interventions, and wireless fNIRS systems further illustrate how physiological signals can be tracked or used in adaptive feedback architectures81,82,83,84,85. These studies are technically informative, but their direct educational relevance is limited unless they are linked to daytime performance, feedback timing, supervision decisions, or learner support in training contexts. They should therefore be interpreted as methodological bridges rather than as evidence that fatigue-aware assessment within CBME is ready for routine use.
Workplace translation requires greater caution than simulation. Authentic clinical environments introduce motion artifacts, interruptions, emotional stress, teamwork demands, patient acuity, variable task goals, and implicit power dynamics between trainees and supervisors. Physiological data in workplace-based assessment should therefore be voluntary, low-stakes, and contextual. Reasonable applications include aggregate fatigue-risk mapping, schedule evaluation, simulation debriefing, or learner-controlled feedback. Unreasonable applications would include using EEG, fNIRS, or multimodal thresholds as standalone evidence for entrustment, promotion, remediation, or discipline.
Established evidence versus conceptual extensions
A central distinction in this review is between established evidence and conceptual extension. The established evidence is that sleep deprivation, circadian disruption, prolonged wakefulness, and fatigue impair neurocognitive domains that matter for clinical work23,24,25,26,27. Meta-analytic and experimental work further supports effects on attention, executive control, mood regulation, and complex task performance39,40,41,42. Medical and procedural studies indicate that fatigue can affect error risk, diagnostic search, crisis management, and technical performance28,29,30,48. However, procedure-level and outcome studies also show heterogeneity, which argues against simple one-to-one interpretation of fatigue and clinical harm31,32,33. Duty-hour trials demonstrate that fatigue-aware policy must consider work-hour design together with sleep opportunity, workload, handoffs, continuity, and educational consequences34,35,36,37,38. EEG studies in physicians and workload paradigms support neural sensitivity to fatigue and cognitive load49,50,51, while fNIRS, wearable, and hybrid EEG-fNIRS studies provide complementary hemodynamic and multimodal evidence59,60,61,62,63,64,65,66,67.
The conceptual extension is the proposal that physiological state indicators might help contextualize trainee performance within CBME. This extension is plausible but unvalidated. Evidence that sleep loss changes neural activity does not automatically imply that real-time monitoring can generate actionable supervisory or educational decisions. At present, no generalizable real-time pattern of neural or hemodynamic activity has been shown to reliably identify cognitive instability or clinical readiness across individuals, tasks, specialties, and operational settings. The same physiological pattern may reflect fatigue, cognitive effort, anxiety, motivation, task difficulty, compensatory engagement, sensor artifact, environmental stress, or individual baseline differences.
Accordingly, the CBME-BCI-Fatigue framework should be interpreted as a research model rather than as an implemented assessment system. Its purpose is to clarify how state-dependent modulators, physiological and behavioral monitoring, and competency-based interpretation of observed performance might be studied together. In simulation or other low-stakes research settings, EEG or fNIRS features could be collected alongside task errors, response time, workload ratings, sleep history, psychomotor vigilance testing, and supervisor feedback to test whether state indicators help explain performance variability. In workplace assessment contexts, physiological data require even greater caution. In the current evidence base, such data may be most appropriate for aggregate fatigue-risk research, schedule design, or voluntary learner support. They should not be used as a standalone basis for entrustment, promotion, remediation, or disciplinary decisions. Table 2 summarizes these distinctions by separating established evidence, conceptual extensions, validation requirements, and representative supporting references within the CBME-BCI-Fatigue framework.
Construct validity, incremental value, individual baselines, and governance
Construct validity is the central methodological challenge for fatigue-aware neurophysiological monitoring. EEG and fNIRS signals are not specific to fatigue or readiness; they may also reflect task difficulty, anxiety, motivation, cognitive effort, compensatory engagement, environmental stress, sensor artifacts, or individual baseline differences. Future studies should therefore predefine the educational construct being measured, specify the expected relation between physiological state and performance, and triangulate physiological signals with behavioral, contextual, and educational data rather than interpreting a single signal as a direct marker of competence.
Incremental value must also be demonstrated. EEG/fNIRS monitoring should be evaluated against simpler and more feasible comparators, including sleep logs, actigraphy, psychomotor vigilance testing, brief cognitive screening, workload ratings, performance telemetry, and expert observation. A neurophysiological approach would be justified only if it improves explanation, prediction, feedback timing, or learner support beyond these measures. Without such evidence, multimodal monitoring may add complexity without improving educational decision-making.
Individual baseline modeling is essential. Interindividual variation in neural activity, sleep vulnerability, chronotype, stress response, clinical experience, and sensor fit limits the value of universal thresholds. A trainee-specific baseline may help distinguish unusual fatigue-related deviation from normal physiological variation, but this requires repeated measurements, reliable sensor placement, standardized task anchors, and transparent analytic models. Machine learning may help integrate multimodal signals, but it cannot solve invalid constructs, biased data, poor signal quality, or unclear causal assumptions. Domain adaptation, privacy-preserving algorithms, and generative modeling may reduce some calibration burden, but they require broader validation before being used to support educational decisions86,87.
Technical and ecological barriers remain substantial. EEG is vulnerable to motion artifact, electrode impedance, environmental noise, montage dependence, and preprocessing choices. fNIRS is affected by optode placement, scalp blood flow, motion, hair, limited depth penetration, and lower temporal resolution. Wearable systems are improving, including compact wireless fNIRS devices for cognitive activity monitoring85, but comfort and feasibility do not establish educational validity. Deployment also requires workflow-compatible setup, interpretable reporting, data security, and end-user acceptance. Social acceptance studies remind us that user trust and perceived purpose are central to adoption88.
Ethical governance should be treated as a boundary condition of the framework rather than as an afterthought. Neurophysiological data can reveal or imply information about fatigue, attention, emotion, or health status. Responsible development of neurotechnology requires purpose limitation, data minimization, secure storage, transparency, access controls, and careful attention to autonomy and consent89,90. Networked BCI systems also raise cybersecurity concerns, especially if decoded outputs influence feedback, training opportunities, or safety-relevant decisions91. In education, learners may feel implicit pressure to participate, and physiological metrics could unintentionally amplify bias if models perform unevenly across populations or contexts. Future research should prioritize voluntary participation where possible, clear separation between supportive feedback and high stakes judgment, appeal mechanisms, and stakeholder-informed governance.