This study involving human participants was conducted in accordance with the Declaration of Helsinki and institutional requirements for behavioral and educational research. The study protocol was approved by the Human Research Ethics Committee of Shenyang City University (Approval No. SU1992837; Approval Date: January 18, 2025). All participants provided written informed consent before enrollment. Participants were informed that the study involved immersive virtual reality exposure, team-based task performance, questionnaire completion, behavioral observation, and short-term follow-up testing. Participation was voluntary, and withdrawal was permitted at any time without academic, employment, or institutional penalty. All study records were de-identified before analysis. Each participant was assigned a numerical participant code, and each team was assigned a separate team code. No names, student numbers, phone numbers, facial images, or directly identifiable audio recordings were included in the analytical dataset. Equipment categories, software versions, and associated replication resources are listed in the Table of Materials.
1. Study design and setting
This completed study used a single-center, three-arm, parallel-group experimental design to examine whether large-space co-located collaborative virtual reality was associated with stronger spatial presence, team coordination, and learning engagement during immersive training. The three training conditions were desktop collaborative simulation, room-scale distributed collaborative virtual reality, and large-space co-located collaborative virtual reality.
The conditions were designed as functional comparators rather than as a full factorial decomposition of immersion, collaboration, and walking space. The desktop condition represented collaboration without exposure to an immersive head-mounted display. The room-scale distributed collaborative virtual reality condition represented immersive team collaboration in separate small tracking areas. The large-space co-located collaborative virtual reality condition combined an immersive visual experience, shared physical movement, and real-time embodied team interaction within a single tracked training space. Group differences were therefore interpreted as differences among the complete training configurations rather than as effects attributable to any single component. This structure was selected because research on immersive learning treats spatial presence, agency, and engagement as mechanisms by which virtual environments may influence learning behavior11.
The study was conducted in the Immersive Learning and Human-Computer Interaction Laboratory at the participating institution from March 3, 2025–July 18, 2025. Recruitment and baseline testing were conducted from March 3–April 25, 2025. Intervention sessions were conducted from March 10–May 30, 2025. The 2-week retention assessments were completed between March 24–June 16, 2025. Data checking, observer scoring, and dataset locking were completed by July 18, 2025.
All sessions were conducted in the same laboratory complex to control lighting, temperature, internet connectivity, task instructions, and researcher supervision. The large-space co-located collaborative virtual reality condition used an unobstructed tracking area of 8 m × 8 m. The room-scale distributed collaborative virtual reality condition used three separated 3 m × 3 m marked areas. The desktop collaborative simulation condition used individual computer workstations separated by acoustic partitions so that participants could communicate only through the assigned communication channel.
Participants were assigned to three-person teams before randomization. The team was used as the training unit because the task required distributed role execution, verbal coordination, shared situational awareness, and joint problem-solving. Randomization was conducted at the team level to avoid contamination between participants within the same session. The overall study workflow and physical configuration of the three conditions are shown in Figure 1.
2. Participants and eligibility screening
Participants were recruited from undergraduate and postgraduate programs at the participating institution. Recruitment notices described the study as an immersive training study involving virtual reality, team coordination, and learning assessment. The notices did not state the expected superiority of any training condition. Interested participants completed an online screening form before laboratory scheduling.
Participants were eligible if they were 18–35 years old, had normal or corrected-to-normal vision, could understand the training language, had no prior participation in the same virtual training scenario, and could stand and move safely for at least 30 min. Participants were excluded if they reported severe motion sickness, uncontrolled epilepsy, recent musculoskeletal injury affecting walking or balance, uncorrected visual impairment, severe vertigo, severe uncorrected hearing impairment, current acute illness, or any medical condition that made immersive virtual reality exposure unsafe. Participants were also excluded if they had extensive prior experience with the exact training scenario or had previously participated in pilot testing of the same system.
Eligibility was verified in two steps. A research assistant first reviewed the online screening form. The session researcher then confirmed eligibility verbally on the day of testing before consent. Participants reporting mild prior motion discomfort were not automatically excluded but were monitored more closely and reminded that participation could be stopped immediately if discomfort occurred. The complete screening form, eligibility items, baseline questionnaire, response formats, and coding rules are provided in Supplementary File 1.
An a priori power calculation was conducted using G-Power software to determine the required sample size for the experimental design. Assuming a medium expected effect size of f = 0.25 based on previous immersive collaborative learning literature, an alpha level of 0.05, and statistical power of 0.80 for a three-group analysis, a minimum of 111 participants was required. To accommodate the three-person team structure and maintain balanced group allocation across conditions, the final sample was expanded to 126 participants comprising 42 independent teams. Because the intervention was delivered at the team level, team identifiers were retained for clustering adjustment, observer-rated coordination scoring, and sensitivity checks.
3. Randomization, allocation concealment, and blinding
After all members of a three-person team had completed eligibility confirmation, the team was randomly assigned to one of the three training conditions using a computer-generated block-randomization list with a fixed block size of six teams. The allocation sequence was generated before data collection using a randomization package in the statistical computing environment, with a reproducible seed of 20250301.
Allocation concealment was maintained using sequentially numbered, opaque, tamper-evident envelopes prepared by a researcher independent of participant recruitment and intervention delivery. The randomization file was stored in a password-protected spreadsheet and was not accessible to the outcome assessors.
Allocation was revealed only after completion of the baseline questionnaire and pre-training knowledge test. The session researcher opened the next sealed allocation envelope immediately before equipment setup. Blinding of participants to the physical training format was not possible because the three conditions were visibly different. Participants were not informed which condition was expected to perform better.
Two observers rated team coordination from standardized screen recordings. The observers were blinded to the study hypotheses, training condition labels, and questionnaire results and were not involved in training delivery. Video files were renamed using neutral team codes before scoring. Condition-identifying interface labels were cropped or masked where possible before observer scoring.
4. Training scenario and task structure
The training scenario was designed as a collaborative emergency logistics and spatial navigation task in a simulated industrial training environment. Each three-person team was required to locate virtual resources, interpret spatial cues, communicate route information, coordinate role-specific actions, and complete a sequence of mission steps within a fixed time window.
Each participant held one of three roles: navigator, resource manager, or safety checker. The navigator interpreted the route direction and destination sequence. The resource manager identified and collected task-relevant virtual items. The safety checker monitored hazards, confirmed procedural order, and prompted the team before final submission.
The scenario contained four zones: orientation zone, resource zone, hazard zone, and completion zone. Each team completed eight required subtasks: identifying the target location, selecting three correct resources from six possible resources, avoiding two incorrect route branches, responding to three hazard prompts, transferring information between team members, and submitting the final task sequence. The maximum task time was 20 min. Teams that did not complete all required subtasks within 20 min were stopped at the time limit, and the number of completed subtasks was recorded.
Before the main task, all participants completed a standardized 8 min orientation. The first 3 min introduced the study rules and safety procedures. The next 3 min introduced the interface controls. The final 2 min allowed participants to practice movement, selection, and communication in a neutral practice environment that did not contain objects, routes, hazards, or answers used in the main scenario. No performance data from the practice environment were analyzed.
The same learning content, task sequence, target answers, route structure, hazard prompts, and scoring criteria were used across all three conditions. The conditions differed in interaction modality, degree of immersion, physical movement, and spatial configuration. The immersive environments were developed using a commercial real-time three-dimensional development platform. Multiplayer networking and spatial state synchronization were implemented using a dedicated multiplayer synchronization framework operating over a local gigabit Ethernet architecture. Participants assigned to the immersive virtual reality conditions used high-resolution head-mounted displays with integrated headphones and microphones. Spatial tracking used an infrared laser-based tracking system with four base stations. Avatar implementation used inverse kinematics algorithms to map real-time head and hand controller telemetry to articulated worker avatars.
5. Pre-study procedure checking and measurement refinement
Before formal data collection, all questionnaire items, knowledge test items, observer-rating rules, and task procedures were reviewed by two researchers with experience in immersive learning and one researcher with experience in educational measurement. The review focused on item clarity, role-task alignment, scoring consistency, safety wording, and construct alignment.
A pilot procedure check was conducted from February 17–February 21, 2025, with six participants who met the same eligibility criteria as the formal sample. The pilot examined task timing, headset comfort, movement safety, instruction clarity, questionnaire completion time, data export accuracy, and observer-rating feasibility. Pilot data were used only to refine wording, setup timing, safety instructions, and log-export procedures. These participants were not included in the final sample, and their data were not included in formal analysis.
6. Desktop collaborative simulation condition
In the desktop collaborative simulation condition, each participant used an individual workstation connected to a 24-inch 1080p monitor. Input devices consisted of a computer mouse and keyboard, and audio communication was provided through a USB headset and microphone. Participants were seated at separate workstations and were instructed not to communicate outside the system.
The desktop simulation used the same virtual environment, task objects, route sequence, hazard prompts, and scoring logic as the virtual reality conditions. The camera view was first-person. As illustrated in Figure 2A, participants in the desktop group navigated using keyboard inputs and interacted with virtual resources using mouse clicks through a fixed two-dimensional monitor field of view. Field of view, movement speed, interaction distance, and audio settings were fixed before data collection and were not changed between sessions. One researcher monitored technical stability during each session but did not provide task advice after the main scenario began.
7. Room-scale distributed collaborative virtual reality condition
In the room-scale distributed collaborative virtual reality condition, the three participants entered the same virtual scenario from separate 3 m × 3 m marked tracking areas. Each participant used a head-mounted display and handheld controllers. Team members communicated through the assigned voice channel and completed the same role-based team task but could not physically co-navigate in a shared large tracking space.
Before the main task, headset fit, interpupillary distance, the boundary system, and controller straps were adjusted. Participants were instructed to stop immediately if dizziness, nausea, eye strain, headache, loss of balance, or anxiety occurred. A researcher remained within 2 m of each active tracking area throughout the session and monitored safety without providing task guidance.
If a participant crossed the marked boundary, the session was paused, and the participant was repositioned. If moderate discomfort occurred, defined as a self-rated discomfort score of 5 or higher on any 0–10 discomfort item, the session was paused and the participant was seated for recovery. If any discomfort item reached 7 or higher, the session was stopped.
8. Large-space co-located collaborative virtual reality condition
In the large-space co-located collaborative virtual reality condition, each three-person team entered the same 8 m × 8 m tracked training area. Each participant wore a head-mounted display, held two controllers, and was represented by a role-specific avatar visible to the other team members. Team members communicated verbally in real time while physically moving through the shared training space. The system synchronized participant location, hand movement, object selection, and task-state changes. As shown in Figure 2B and Figure 2C, participants in the immersive conditions executed tasks using three-dimensional spatial hand gestures with tracked controllers and physical head rotation for a 360° field of view.
The large-space condition used the same task content and 20 min time limit as the other conditions. Physical movement was calibrated so that the virtual route remained within the marked safety area. A safety boundary appeared in the headset when participants approached the edge of the tracking zone. Before the main task, floor clearance, cable-free movement, headset fit, tracking stability, audio input, and controller response were checked. Participants were instructed to use natural physical walking rather than artificial locomotion techniques to navigate the virtual environment. Physical contact between team members was prohibited. If two participants approached within 0.5 m of each other, a neutral safety reminder was provided without task information.
Each large-space session was supervised by two researchers. One researcher monitored the system dashboard and recording status. The second researcher stood outside the active tracking area and monitored participant safety. The task was paused if tracking was lost for more than 10 s, if a participant reported discomfort, or if a participant crossed the safety boundary. If the interruption lasted less than 2 min and the participant wished to continue, the task resumed from the last saved state. If the interruption exceeded 2 min, the session was recorded as incomplete and retained in the intention-to-treat dataset.
9. Outcome measures and assessment timing
Assessments were completed at baseline, immediately after training, and 2 weeks after training. Baseline measures included demographics, prior virtual reality and team-training experience, gaming familiarity, and a 30-point knowledge pretest. Immediate post-training measures included spatial presence, self-rated and observer-rated team coordination, learning engagement, cognitive load, simulator discomfort, post-training knowledge, and task-performance indicators. The 2-week follow-up used the same 30-point knowledge test with reordered items and minor wording variation. The assessment schedule is shown in Table 1.
Table 1: Summary of primary and secondary measures. This table summarizes the variables, measurement scales, assessment timing, and data sources for the primary outcomes, secondary learning outcomes, and safety and process measures evaluated throughout the experimental protocol. Please click here to download this Table.
Spatial presence was measured immediately after training using an 8-item, 7-point scale adapted from established presence dimensions12. Scores were averaged when at least 80% of items were completed; otherwise, the score was coded as missing. Full items and scoring rules are provided in Supplementary File 1.
Self-rated team coordination was measured using a customized six-item, 7-point scale covering communication clarity, role awareness, mutual monitoring, assistance timing, shared task comprehension, and error recovery. Pilot testing with 20 participants assessed item clarity and structure, and exploratory factor analysis identified a single dominant factor. Internal consistency in the main sample was Cronbach’s α = 0.88. Participant scores were calculated as the mean of completed items, and team-level scores as the mean of the three participant scores. Full items and scoring criteria are provided in Supplementary File 1.
Observer-rated team coordination was assessed using a customized rubric covering communication timing, role execution, mutual support, spatial coordination, and error recovery. Each domain was scored from 0–20, yielding a total score from 0–100. The rubric underwent expert review, and two observers completed 10 h of training before formal scoring. Inter-rater reliability was evaluated using a two-way random-effects intraclass correlation coefficient (ICC) for absolute agreement, with ICC ≥ 0.75 required before final scoring. Total-score differences greater than 12 points were resolved by consensus; otherwise, the mean of the two scores was used. The complete rubric and reconciliation procedure are provided in Supplementary File 1.
Learning engagement was measured immediately after training using a 12-item, 7-point short-form scale13. Scores were averaged across completed items, with higher values indicating stronger engagement. Internal consistency was assessed using Cronbach’s alpha. Cognitive load was assessed using five 1–7 ratings of mental demand, physical demand, time pressure, effort, and frustration14. Full items and scoring rules are provided in Supplementary File 1.
Knowledge acquisition was assessed using a 30-point test comprising 20 single-best-answer items and five short scenario judgment items. The same blueprint was used for pretest, post-test, and 2-week retention testing, with reordered items and equivalent wording changes. A gain of at least 3 points from pretest to post-test was defined as a practically meaningful descriptive improvement. The test blueprint, scoring rules, and judgment rubric are provided in Supplementary File 1.
Simulator discomfort was assessed before exposure, after orientation, and after the main task using five 0–10 symptom ratings. The mean of the five items formed the discomfort index. Sessions were stopped if any item reached 7 or higher. Post-task indices above 6 required at least 15 min of monitoring, and indices of 5 or higher required 24 h follow-up. Full safety-scoring rules are provided in Supplementary File 1.
Task-performance indicators included completion status, completion time, correct resources selected, hazard prompts correctly handled, route errors, and safety interruptions. Incomplete teams were assigned 1,200 s for completion time. Outcome definitions, timing, score ranges, data sources, and analysis levels are summarized in Table 1 and detailed in Supplementary File 1.
10. Data collection procedure
Each session followed the same sequence. Participants arrived 15 min before the scheduled training time. Identity was confirmed using the appointment list, eligibility was checked again, the consent form was explained, and procedural questions were answered. Participants then signed the consent form and received participant codes. The expected superiority of any condition was not discussed.
Participants completed the baseline questionnaire and pre-training knowledge test on a laboratory tablet. The baseline phase lasted approximately 12 min. After baseline completion, team allocation was revealed, and the assigned training setup was prepared. Equipment preparation lasted approximately 5 min for the desktop condition, 8 min for the room-scale distributed collaborative virtual reality condition, and 10 min for the large-space co-located collaborative virtual reality condition. Equipment preparation was not counted as training exposure.
All participants then completed the standardized 8 min orientation. The main task started immediately after orientation and lasted up to 20 min. No teaching, hints, or corrective feedback were provided during the main task. Technical assistance was limited to restoring system function, adjusting headset fit, repositioning participants within the safety boundary, or repeating previously stated safety instructions. After the task, participants removed the equipment and completed the post-training questionnaire and post-training knowledge test at individual stations. The post-training assessment lasted approximately 15 min.
The 2-week retention test was administered online 14 days after the training session, with an acceptable completion window from day 13–day 16. Participants received one reminder 24 h before the due time and one reminder on day 15 if the test had not been completed. Responses submitted after day 16 were retained in the raw dataset but excluded from the primary retention analysis. The actual number of days between training and follow-up was recorded for sensitivity analysis.
11. Data quality control
Electronic questionnaires included range checks for key outcomes, while demographic items could be skipped. Submission times were recorded, and questionnaires completed in less than one-third of the median completion time were flagged for review. System logs were exported after each session and matched against session records within 48 h; discrepancies were resolved using timestamped recordings.
Observer-rated coordination scores were entered independently by two observers. Total-score differences greater than 12 points were flagged for consensus review, and final scores were locked before group-level analysis. Multi-item scale scores were calculated when at least 80% of items were completed; otherwise, scores were coded as missing. Unanswered single-best-answer knowledge items were scored as 0. Short scenario judgment items were scored independently by two raters, with disagreements greater than 1 point reviewed by a third rater.
The final dataset was locked after verification of eligibility records, participant and team codes, questionnaire ranges, knowledge scores, observer ratings, system logs, retention-test validity, data merges, and missingness. Detailed rules for scoring, cleaning, and dataset locking are provided in Supplementary File 1.
12. Safety monitoring and stopping criteria
Safety monitoring was conducted throughout all training sessions. Before exposure, participants were reminded that discomfort could occur during virtual reality training and that withdrawal was allowed at any time. During the task, posture, balance, movement speed, and verbal signs of discomfort were monitored. The task was stopped immediately if a participant reported severe dizziness, nausea, visual discomfort, headache, anxiety, loss of balance, or a wish to stop.
Predefined stopping criteria were applied consistently. The session was stopped if any discomfort item reached 7 or higher on the 0–10 scale, if a participant crossed the safety boundary twice in one session, if tracking loss lasted more than 2 min, if equipment malfunction prevented normal interaction, or if continued participation was judged to create a safety risk. Stopped sessions were documented using a session interruption form. Participants were seated after stopping and monitored until symptoms returned to a mild level, defined as all discomfort items below 3. No participant was allowed to leave the laboratory while reporting moderate or severe discomfort.
13. Data management and confidentiality
All data were stored in a de-identified project folder on an encrypted institutional drive. The linkage file connecting participant identities to participant codes was stored separately and was accessible only to the principal investigator. Questionnaire data, task logs, observer ratings, and knowledge test scores were merged using participant and team codes only, and the analytical dataset contained no personal identifiers.
Data were organized at participant and team levels. Participant-level files contained identifiers, condition assignment, demographic variables, questionnaire outcomes, knowledge outcomes, and task-related variables. Team-level files contained team identifiers, condition assignment, observer-rated coordination, task completion, completion time, resource selection, hazard handling, route errors, and safety interruptions. Raw observer scores were retained separately before averaging or consensus resolution. Variable definitions, coding ranges, and dataset structure are provided in Supplementary File 1.
14. Statistical analysis
Primary outcomes were spatial presence, self-rated team coordination, observer-rated team coordination, and learning engagement. Secondary outcomes included post-training knowledge, 2-week retention, cognitive load, simulator discomfort, task completion, completion time, and task-process indicators. Training conditions comprised desktop collaborative simulation, room-scale distributed collaborative virtual reality, and large-space co-located collaborative virtual reality.
Participant- and team-level variables were summarized using appropriate descriptive statistics. Baseline characteristics included age, gender, prior virtual reality experience, gaming familiarity, prior team-based training experience, and pre-training knowledge. Because randomization occurred at the team level, intraclass correlation coefficients were calculated to assess within-team dependence. Participant-level primary outcomes were analyzed using linear mixed-effects models with training condition as a fixed effect and team identifier as a random intercept. Models used restricted maximum likelihood estimation and included available data under the missing-at-random assumption. No imputation was applied to primary questionnaire outcomes. Estimated marginal means were compared using Tukey-adjusted pairwise contrasts. Observer-rated coordination was analyzed at the team level.
The Benjamini-Hochberg false discovery rate procedure was applied across the four primary outcomes as a sensitivity analysis. All randomized teams and participants were analyzed according to original allocation. Multi-item scale scores were calculated when at least 80% of items were completed; otherwise, scores were coded as missing. Residual normality was assessed using Q-Q plots and Shapiro-Wilk tests, and homogeneity of variance was assessed using Levene’s test. Strongly non-normal distributions were examined using the Kruskal-Wallis test as a sensitivity analysis.
Effect sizes were reported with p values. Analysis of variance results included eta squared and partial eta squared; pairwise comparisons included Cohen’s d with 95% confidence intervals; and mixed-effects models reported estimated mean differences with 95% confidence intervals. Statistical significance was defined as two-sided p < 0.05. Associations among spatial presence, team coordination, learning engagement, and knowledge outcomes were examined using Pearson or Spearman correlations, as appropriate. Statistical syntax for descriptive analyses, mixed-effects models, team-level coordination analysis, multiple-outcome adjustment, assumption checks, knowledge analyses, and correlation analyses is provided in Supplementary File 1.