The study protocol was reviewed by Henan Agricultural University and was determined to be exempt from formal ethics review on December 20, 2025. No separate ethics approval or exemption reference number was issued by the reviewing body. Written informed consent was obtained from all participants before enrollment.
Participant preparation
Adult participants aged 18–60 years with normal or corrected-to-normal vision were recruited. Participants who reported severe vestibular disorders, uncontrolled epilepsy, severe visual impairment, recent major surgery, or previous serious discomfort during virtual-reality exposure were excluded. Each participant completed a screening form before the experiment. Age group, gender, education level, prior virtual-reality experience, and prior museum or exhibition-visiting frequency were recorded as control variables.
Participants were assigned to one of three layout conditions using a computer-generated randomization list: manual layout, standard simulated annealing layout, or adaptive simulated annealing with reheating layout. A between-subject design was used to avoid learning effects across layouts. The allocation ratio was set to 1:1:1, with 60 participants assigned to each condition and 180 participants in total. Age group, gender, education level, previous virtual-reality experience, and museum/exhibition-visiting frequency were recorded before allocation for descriptive assessment of participant characteristics; these variables were not used as post-randomization exclusion criteria.
Exhibition-space model preparation
A rectangular immersive exhibition hall measuring 36 m x 24 m with a ceiling height of 4.5 m was constructed. The total usable floor area was set to 864 m2. One entrance and one exit were placed on opposite short sides of the hall. Both the entrance and exit widths were set to 3.0 m and were kept fixed across all layout conditions.
The exhibition hall was divided into six thematic exhibition zones, two transition corridors, one rest area, and one main visitor route. The allowable area for each thematic zone was set to 80–130 m2, and the rest-area allocation was set to 40 m2. The allowable main-pathway width was set to 1.8–3.5 m. Eight to twelve interaction nodes were placed in each layout. Each interaction node was defined as a point where participants could stop, activate digital content, view an exhibit, or change direction.
The floor plan was discretized into 0.5 m x 0.5 m grid cells. The grid was used for visibility analysis, route-continuity checking, heatmap coverage calculation, and local density estimation. Each thematic zone was represented by its center coordinate, boundary polygon, area, and route-sequence position. Each interaction node was represented by its coordinate, associated zone, trigger radius, and clearance area.
Spatial configuration was treated as the core design variable because exhibition layout, visibility, and route structure affect visitor movement, wayfinding, and spatial cognition18. The complete workflow, from spatial model construction to layout optimization, virtual-scene implementation, participant testing, and statistical analysis, is shown in Figure 1.
Layout parameters and feasibility constraints
The same feasibility constraints were applied to all layout conditions. First, all thematic zones were required to remain inside the 36 m x 24 m boundary. Second, thematic zones were not allowed to overlap. Third, the main visitor route was required to connect the entrance, six thematic zones, a rest area, and the exit without interruption. Fourth, interaction nodes were not allowed to block the main path. Fifth, each interaction node was required to retain a clearance radius of at least 1.2 m. Sixth, the predicted local density on the main route was required to be no more than 1.50 persons/m2. Seventh, the final route was required to contain no dead ends or disconnected subpaths.
The same scene scale, route rules, environmental targets, and interaction settings were used across the three layout conditions. Before the optimization algorithms were run, the hall dimensions, functional-zone parameters, pathway-width range, interaction-node settings, feasibility thresholds, virtual-reality settings, objective-function weights, algorithm parameters, and outcome definitions were recorded. These parameters defined the reproducible boundary of the protocol and are summarized in Table 1.
Spatial-cost function
Each candidate layout was evaluated using the weighted spatial-cost function:

where F(x) is the total spatial cost of layout x. Cdistance(x) is the walking-distance cost, Cvisibility(x) is the visual-accessibility cost, Ccrowding(x) is the crowding cost, Ccomfort(x) is the environmental-comfort penalty, and Cinteraction(x) is the interaction-balance penalty. A lower F(x) indicates a more favorable layout.
The walking-distance cost was calculated as:

where Lx is the total route length of layout x in meters. A reference maximum route length of 120 m was used.
The visual-accessibility cost was calculated as:

where Vx is the thematic-zone frontage visible from the main route, and Vtotal is the total frontage length of all thematic zones. Line-of-sight rays were emitted every 0.5 m along the visitor route. Exhibition partitions, walls, and zone boundaries were treated as occluding objects.
The crowding cost was calculated as:
Ccrowding(x) = min(Dmax / 1.50, 1)
where Dmax is the maximum predicted local density in persons/m2. An upper density threshold of 1.50 persons/m2 was used for the main route.
The environmental-comfort penalty was calculated as:

The illuminance penalty was set as:

The sound-pressure penalty was set as:

The carbon dioxide penalty was set as:

The temperature penalty was set as:

The interaction-balance penalty was calculated as:

where sk is the route distance between two adjacent interaction nodes. A lower value indicates a more even distribution of interaction opportunities along the visitor route.
Environmental comfort is included because lighting, sound, air quality, and temperature influence fatigue, comfort, and willingness to remain in exhibition environments19. The environmental ranges are protocol control targets rather than universal regulatory limits. The 300–500 lx illuminance band is consistent with current Chinese building and museum-lighting guidance; the carbon-dioxide target of <900 ppm is deliberately more conservative than the 1,000 ppm daily-average limit in GB/T 18883-202220,21; the 22–25 °C temperature band is within commonly accepted thermal-comfort conditions for lightly active occupants22; and 45–60 dB(A) is used as a controlled ambient-sound band rather than as a statutory indoor-noise limit23,24.
The objective-function weights were selected as explicit, scenario-specific design priorities rather than as universal coefficients or preference estimates fitted from participant outcomes. Walking distance and crowding each received a weight of 0.25 because route efficiency and congestion control were treated as the two primary operational constraints; visual accessibility received 0.20 because visibility affects orientation and exhibits exposure; environmental comfort and interaction each received 0.15 so that these factors influence the search without dominating circulation-related terms. Because exhibition priorities can differ by context, the protocol tests local weight robustness by perturbing each weight by ±10% individually and proportionally renormalizing the remaining weights so that the total remains 1.00. The resulting coefficients should therefore be interpreted as protocol-specific design settings that can be recalibrated for artistic, historical, scientific, or commercial exhibitions.
Manual layout generation
The manual layout was created as the baseline condition. The six thematic zones were arranged in a conventional narrative sequence from entrance to exit. The rest area was placed near the middle section of the route. The initial main-pathway width was set to 2.2 m. Interaction nodes were placed near the entrance of each thematic zone and at major route-turning points. The manual layout was checked against all feasibility constraints. Zone boundaries were adjusted only when a constraint was violated. After the layout became feasible, total route length, visible frontage ratio, maximum local density, environmental-comfort penalty, interaction-spacing coefficient, and total spatial cost were recorded.
Standard simulated annealing
The feasible manual layout was used as the initial state. The initial temperature was set to 100, the cooling coefficient to 0.95, the maximum number of iterations to 1,500, and the stagnation threshold to 200 iterations. The algorithm was run with 30 independent random seeds. These values were operational protocol settings used to define a matched computational budget; they were not presented as literature-derived universal constants or outcome-tuned optima. T₀ = 100 permitted broad early exploration, α = 0.95 provided gradual cooling, 1,500 iterations provided each run with the same upper search budget, and the 200-iteration stagnation threshold prevented prolonged continuation after no improvement. The same temperature and maximum-iteration settings were used for the adaptive method so that differences were attributable to proposal scheduling and reheating rather than to a larger nominal search budget.
At each iteration, one neighboring layout was generated by applying one move: moving one thematic zone by 0.5–2.0 m, swapping the sequence positions of two thematic zones, adjusting one pathway segment by 0.1–0.3 m, or moving one interaction node by 0.5–1.5 m. A proposed layout was rejected immediately if it violated any feasibility constraint.
If the proposed layout was feasible and had a lower spatial cost, it was accepted. If it had a higher spatial cost, it was accepted according to the probability equation:

where ΔF is the increase in spatial cost, and T is the current temperature. The temperature after each iteration was updated using:
Tnew = 0.95Told
Each run was stopped when the maximum iteration count was reached or when the best spatial cost did not improve for 200 consecutive iterations. Simulated annealing was considered suitable for this task because exhibition layout optimization is a combinatorial spatial problem that cannot be solved efficiently by exhaustive search.
Standard simulated annealing was used as the primary numerical algorithmic benchmark because it shared the same layout encoding, feasibility rules, objective function, initialization, temperature schedule, and computational budget as the adaptive condition. Standard simulated annealing was the sole numerical metaheuristic comparator in this study and therefore functioned as a matched ablation-like baseline for the added proposal scheduling and reheating at the workflow level. Genetic algorithms, particle swarm optimization, ant colony optimization, and learning-based optimizers were not part of the experiment; accordingly, performance claims were restricted to the matched standard-SA baseline. For each run, the random seed, initial spatial cost, final spatial cost, improvement percentage, number of iterations to the best solution, runtime, acceptance rate, and feasibility status were stored. Improvement percentage was calculated using the equation:

Adaptive simulated annealing with reheating
The same initial layout, feasibility constraints, initial temperature, cooling coefficient, maximum iteration count, and 30 random seeds as the standard simulated annealing condition were used. The adaptive algorithm differed only in its stage-dependent move-class probabilities and reheating rule, while retaining the standard Metropolis acceptance rule for feasible candidate layouts. Four proposal move classes were defined: local displacement (L), zone-sequence swap (S), pathway-width adjustment (W), and interaction-node relocation (N). During the first 40% of iterations, the move-class probabilities were set to pL = 0.20, pS = 0.35, pW = 0.30, and pN = 0.15. During the final 60%, they were set to pL = 0.35, pS = 0.15, pW = 0.20, and pN = 0.30. This change in move-class frequency shifted the search from broader sequence/pathway exploration toward local spatial refinement and adjustment of interaction nodes.
At iteration t, the full proposal density was expressed as q(x′|x,t) = pm(t)qm(x′|x), where m ∈ {L,S,W,N} represented the selected move class and pm(t) represented its phase-specific probability. For local displacement, one of six thematic zones was selected uniformly, a displacement direction was drawn uniformly on [0,2π), and a displacement magnitude was drawn uniformly on [0.5,2.0] m; therefore, qL was proportional to (1/6)(1/2π)(1/1.5). For a zone-sequence swap, one of the 15 unordered zone pairs was selected uniformly, giving qS = 1/15. For pathway-width adjustment, one of J adjustable pathway segments was selected uniformly, and a signed change was drawn uniformly from [−0.3,−0.1] ∪ [0.1,0.3] m, giving qW = (1/J)(1/0.4). For interaction-node relocation, one of K current interaction nodes was selected uniformly, a direction was drawn uniformly on [0,2π), and a displacement magnitude was drawn uniformly on [0.5,1.5] m, giving qN proportional to (1/K)(1/2π)(1/1.0). Constraint-violating proposals were rejected and contributed a self-transition.

For every reversible feasible move in this implementation, the conditional proposal kernel was symmetric: qm(x′|x) = qm(x|x′). Because the same phase-specific move-class probability pm(t) was applied to the forward and reverse transition at a given iteration, q(x|x′,t)/q(x′|x,t) = 1. The acceptance probability, therefore, is reduced to the standard simulated-annealing Metropolis rule, A(x→x′) = min{1, exp[−(C(x′)−C(x))/T]}. No non-trivial Hastings correction was applied; the method is therefore termed adaptive simulated annealing with reheating rather than a Metropolis-Hastings sampler25.
Reheating was applied when the best spatial cost did not improve for 150 consecutive iterations, thereby intentionally intervening before the 200-iteration stagnation stop used in standard simulated annealing. The current temperature was increased by 10%, and up to 3 reheating events were allowed per run. The same algorithm-performance variables as in the standard simulated annealing condition were stored. The experiment evaluated the combined stage-dependent proposal schedule plus reheating configuration and therefore did not estimate the independent causal contribution of each component. Stepwise pseudocode for the exact proposal, feasibility, Metropolis acceptance, temperature update, reheating, and output storage sequence was provided in Supplementary File 3.
Treheated = 1.10Tcurrent
Final layouts
One final layout was selected for each condition. The feasible manual layout was used for the manual condition. Under the standard simulated annealing conditions, the run with the lowest final spatial cost among the 30 runs was selected. For the adaptive simulated annealing with reheating condition, the feasible run with the lowest final spatial cost among the 30 runs was selected. Before the selected layouts were exported, a final feasibility check confirmed that all zones were inside the boundary, no zones overlapped, the route was continuous, the entrance and exit were unchanged, interaction-node clearance was at least 1.2 m, maximum density remained below 1.50 persons/m2, and all thematic zones were reachable from the main route.
Virtual exhibition scenes
Three virtual exhibition scenes were built in Unity 2022.3.22f1 LTS (Unity Technologies) using XR Interaction Toolkit 2.5.4. The scenes were presented through a Meta Quest 2 head-mounted display (Meta Platforms) using six-degree-of-freedom head and controller tracking. Exhibition content, visual style, object models, signage style, lighting assets, ambient sound, interaction mechanism, and navigation instructions were kept identical across the three scenes. Only spatial arrangement, route sequence, pathway width, and interaction-node placement varied.
The virtual camera height was set to 1.65 m, and the walking speed was set to 1.2 m/s. Teleportation was disabled. The same starting trigger at the entrance and the same completion trigger at the exit were used. The interaction-node trigger radius was set to 1.0 m. Whether each interaction node was visited was recorded.
The same six thematic contents were used in all scenes. Each theme was assigned to its corresponding zone, regardless of layout conditions. Wall height, partition style, media display size, text density, and exhibit object scale were kept constant. Virtual-reality environments were considered appropriate for this protocol because they enabled the examination of spatial presence, route clarity, and visitor experience under controlled layout conditions26.
The hardware, software, environmental-monitoring tools, questionnaire/codebook files, and analysis resources required to reproduce the protocol were listed in the Table of Materials and organized as Supplementary Files 1–3. Research resources with registered Research Resource Identifiers (RRIDs) were listed by RRID in the Table of Materials.
Participant session
One participant was tested at a time. The participant sat for 2 min before the session. The head-mounted display was fitted, and the interpupillary distance was adjusted according to the participant's comfort. A 2 min practice scene that was not part of the exhibition was provided. The practice scene was used only to familiarize the participant with movement and interaction.
Because one participant was tested at a time, crowding in the protocol was represented by model-based density estimates rather than emergent simultaneous group movement. Multi-user co-presence, interpersonal avoidance, and dynamic crowd interactions were therefore not directly simulated and were considered when interpreting crowd-related outcomes.
After practice, the assigned exhibition scene was started. The participant was instructed to enter the exhibition, follow the route naturally, interact with marked nodes, and exit after completing all six thematic zones. Additional route guidance was provided only if the participant could not proceed for more than 30 s. Completion time, dwell time, walking distance, hesitation count, interaction-node visits, and trajectory coordinates were recorded automatically.
The session ended when the participant reached the exit trigger. The head-mounted display was removed, and the participant was allowed to rest for 3 min. The participant completed the post-experience questionnaire immediately after the rest period.
Behavioral, environmental, and questionnaire outcomes
Completion time was defined as the time from entrance-trigger activation to exit-trigger activation. Dwell time was defined as the total time spent inside thematic zones and interaction-node areas. Walking distance was defined as cumulative path length from trajectory coordinates. Hesitation count was defined as the number of pauses longer than 3 s outside interaction-node areas. Heatmap coverage was calculated using:

Average illuminance was recorded using a Testo 540 illuminance meter (Testo SE & Co. KGaA), sound pressure level using a Testo 816-1 Class 2 sound level meter (Testo SE & Co. KGaA), carbon-dioxide concentration using a Testo 535 CO₂ measuring instrument (Testo SE & Co. KGaA), and ambient temperature using a Testo 605i thermohygrometer (Testo SE & Co. KGaA). The same monitoring positions were used across the three scene conditions. Illuminance was recorded in lx, sound pressure level in dB(A), carbon dioxide in ppm, and temperature in °C.
Presence and spatial-presence domains were measured using the 7-point scoring structure specified in the study protocol. These constructs were interpreted with reference to the validated Presence Questionnaire (PQ) framework of Witmer and Singer and the Igroup Presence Questionnaire (IPQ) framework of Schubert et al.27. The study variables were reported as domain-level 7-point scores rather than as standard PQ or IPQ total scores. For simulator sickness, the Simulator Sickness Questionnaire symptom framework was used together with the prespecified 0–30 study summary variable; higher values indicated greater discomfort. This 0–30 study summary was reported separately from the standard weighted SSQ Total Severity score, consistent with the distinction between simulator-sickness measurement conventions and cybersickness reporting28. Interaction quality, wayfinding clarity, visual comfort, acoustic comfort, thermal comfort, perceived crowding, cognitive load, and overall satisfaction were measured using the study-specific, multi-item, 7-point rating scales. Supplementary File 2 documented the construct source, score range, and reporting convention for each domain.
Study outcomes
The primary algorithmic outcome was defined as the final spatial cost. The primary visitor-experience outcome was defined as overall satisfaction. Secondary algorithmic outcomes were defined as the improvement percentage, the number of iterations to the best solution, the runtime, and the acceptance rate. Secondary visitor-experience outcomes were defined as presence, spatial presence, interaction quality, wayfinding clarity, visual comfort, acoustic comfort, thermal comfort, perceived crowding, cognitive load, simulator sickness, completion time, dwell time, walking distance, hesitation count, and heatmap coverage. Completion time and walking distance were used to evaluate route efficiency. Hesitation count and wayfinding clarity were used to evaluate navigation quality. Presence, spatial presence, interaction quality, and dwell time were used to evaluate immersive engagement. Perceived crowding, cognitive load, comfort ratings, and simulator sickness were used to evaluate experiential burden.
Data cleaning and quality control
All participant records were checked before analysis. It was confirmed that each record had a single participant identifier, a single condition label, complete behavioral logs and questionnaire scores, and no duplicate identifiers. Presence, spatial-presence, and study-specific experience ratings were restricted to their defined 1–7 ranges. The simulator-sickness study summary variable was restricted to its prespecified 0–30 range. The questionnaire/codebook file was checked to confirm that it documented the construct source, response range, domain construction, and derived-score definitions used in the analysis.
A participant record was excluded if the session was not completed, completion time was less than 3 min or greater than 25 min, heatmap coverage was less than 10%, the trajectory log was incomplete or technically invalid, walking distance was less than 20 m in conjunction with an incomplete/invalid trajectory record, or the participant requested early termination. No upper walking-distance cutoff was applied; walking distance was retained as a continuous study outcome across its valid observed range. An algorithm run was excluded from layout selection if it violated any feasibility constraint. Random seeds, layout files, algorithm outputs, questionnaire/codebook files, and analysis scripts were preserved for reproducibility. Random-seed preservation was necessary because stochastic optimization outputs could vary with initialization and proposal behavior29.
Algorithm performance
The 30 standard simulated annealing runs and 30 adaptive simulated annealing with reheating runs were analyzed. Initial spatial cost, final spatial cost, improvement percentage, iterations to best solution, runtime, and acceptance rate were summarized as mean ± standard deviation (SD), and 95% confidence intervals were reported for the group means and between-algorithm differences.
Normality was assessed using quantile-quantile (Q-Q) plots and the Shapiro-Wilk test. Homogeneity of variance was assessed using Levene's test. The two algorithmic conditions were compared using independent-samples t tests when assumptions were met. Mann-Whitney U tests were used when assumptions were not met. Two-sided p values, 95% confidence intervals, and effect sizes were reported. Cohen's d was calculated using the equation

where:

Rank-biserial correlation was used for nonparametric comparisons. Sensitivity analysis was conducted by changing each objective-function weight by ±10% individually and proportionally renormalizing the remaining four weights so that the total remained 1.00. The final spatial cost and improvement percentage were recalculated for each alternative weight vector. The algorithmic comparison was considered stable when the adaptive simulated annealing with reheating condition retained a lower final spatial cost across the alternative weight settings.
Participant outcome analysis
Participant outcomes across the three layout conditions were analyzed using IBM SPSS Statistics 27.0 (IBM Corp.). Continuous variables were reported as mean ± SD, and categorical variables were reported as counts and percentages. Normality was assessed using Q-Q plots and the Shapiro-Wilk test. Homogeneity of variance was assessed using Levene's test. One-way analysis of variance (ANOVA) was used for normally distributed outcomes with acceptable homogeneity of variance. Welch's ANOVA was used when the assumption of homogeneity of variance was violated. Kruskal-Wallis tests were used for non-normal outcomes.
Pairwise comparisons were conducted among the manual layout, standard simulated annealing layout, and adaptive simulated annealing with reheating layout. The three pairwise comparisons within each outcome were adjusted using the Holm method. The 95% confidence intervals were reported together with adjusted p-values. Eta-squared was reported for standard ANOVA, omega-squared for Welch's ANOVA, epsilon-squared for Kruskal-Wallis tests, and standardized mean differences for pairwise comparisons. As a supplementary robustness analysis, pairwise Welch contrasts were calculated from the group mean, SD, and n = 60 per group. Welch degrees of freedom, two-sided p values, Holm-adjusted p values, 95% CIs for mean differences, and Hedges g were calculated. This supplementary analysis used the group-level statistics and was presented in Supplementary Table 1.
The randomized 1:1:1 three-group comparison was treated as the primary analysis. Age group, gender, education level, previous virtual-reality experience, and museum/exhibition-visiting frequency were recorded before allocation to characterize the sample and identify obvious imbalance. The trial was not designed or powered for subgroup interaction tests, and no prespecified moderator model was included; these characteristics were therefore treated as contextual variables rather than demonstrated moderators. Random allocation provided design-based control over baseline characteristics, while the residual influence of prior VR familiarity was explicitly retained as a limitation. Eta-squared was calculated using

Epsilon-squared for Kruskal-Wallis tests was calculated as:

where H is the Kruskal-Wallis test statistic, k is the number of groups, and n is the total sample size.
Overall satisfaction, presence, wayfinding clarity, perceived crowding, cognitive load, completion time, walking distance, hesitation count, and heatmap coverage were prioritized in the interpretation because these variables directly reflected spatial efficiency and immersive visitor experience.