Research Article

A Gesture- and Voice-controlled Virtual Reality System for Immersive Home Design: System Design and User Experience Analysis

DOI:

10.3791/70051

March 3rd, 2026

In This Article

Summary

Loading...
$$\rightleftharpoonup{xx}$$ $$\longleftharp{xx}$$, $$\longrightharp{xx}$$,

This paper presents a virtual reality system controlled by gestures and voice for immersive home design, which improves accuracy, usability, and accessibility through the multimodal fusion of gesture and speech inputs. Experimental results indicate shorter task times, enhanced placement accuracy, and improved user satisfaction compared to traditional practice.

Abstract

Loading...
$$\rightleftharpoonup{xx}$$ $$\longleftharp{xx}$$, $$\longrightharp{xx}$$,

Gesture- and voice-based interaction in VR has shown promise in various applications; however, existing systems are hindered by their emphasis on navigation tasks, dependence on two-handed gestures, lack of fine-grained precision, and testing on small, homogeneous populations. These limitations limit applicability to intricate, design-centric workflows. To address these limitations, we introduce a Gesture- and Voice-Controlled Virtual Reality System for immersive home design, which incorporates the multimodal fusion of hand-tracked gesture input and natural language instructions for precise furniture placement, material selection, and spatial adjustment. Three aspects of novelty in the system are (i) adaptive modes of interaction to support one-handed usage, numeric voice commands, and coarse-to-fine adjustment of precision; (ii) multimodal fusion architecture merging gesture, speech, and context information for sub-centimeter precision; and (iii) task-aware user experience evaluation framework tailored to home design processes. The development process follows a systematic pipeline, from requirements analysis and design of gesture and voice lexicons, to multimodal intent recognition, VR scene assembly with snapping and alignment, and iterative testing. Experimental results on a wide variety of participant populations demonstrate a 30% decrease in placement error and substantial improvements in usability and workload ratings over controller-based interaction. The results suggest the potential of multimodal VR for creating accessible and engaging home design experiences.

Introduction

Loading...
$$\rightleftharpoonup{xx}$$ $$\longleftharp{xx}$$, $$\longrightharp{xx}$$,

Virtual Reality (VR) is an interactive computer-generated environment that produces a three-dimensional experience in which users can see, interact with, and communicate with technology as if they were actually present. In the context of this study, virtual reality is regarded as a multimodal interaction environment, rather than merely a visual display technique. Latest VR systems integrate visual engagement with gesture recognition, spatial senses, and natural voice input, allowing users to operate simulated items with high precision. This comprehension is consistent with current VR research that prioritizes immersion, engagement, and natural user experiences. Its application areas cover performing, teaching2, medicinal treatment3, and travel4. VR allows people to interact and connect with worlds that would not be attainable in real life by simulating real-world settings. For home-design processes, where users must perceive spatial layouts, control objects, and receive real-time feedback without requiring physical motion, such immersion features are especially helpful. These features can bring new possibilities, particularly to users who have incomplete physical movement. Members of informally disadvantaged groups, especially the elderly and those with minimal alphanumeric literacy, will have their outdoor activities restricted by several factors, such as physical constraints, mobility issues, and difficulties in accessing digital technology due to age-related limitations. It is, therefore, necessary to provide constant attention to research that will indirectly make the outside world accessible to these people through VR expertise. This need for continued research attention not only supports the creation of accessibility-oriented VR experiences but also drives the development of systems that go beyond basic navigation tasks. Indeed, immersive home-design environments -- when controlled through natural gestures and voice-can offer elderly users, people with motor impairments, and digitally vulnerable populations intuitive means to interact with complex spatial environments. Allowing indirect access to external spaces and creative tasks, such systems contribute to both functional independence and emotional well-being. It is on these grounds that this study will introduce a multimodal VR home-design system. Subsequently, it is the desire for constant attention to research that will help make the outside world indirectly accessible to these people using VR expertise. Investigations have shown that VR-based interventions created for older adults can have several advantages, such as reducing the risk of falls by enhancing postural control and flexibility6, as well as improving general postural regulation and firmness7. Additionally, there has been an increasing availability of VR material customized specifically for the older adult population8,9. Even while VR interaction has advanced, most current systems primarily rely on controllers and two-handed gestures, which limit accessibility and accuracy for tasks that require design. Expanding on these mobility features, our work expands VR beyond simply observing by allowing active, precise, and realistic management of interior design elements using gesture and speech methods

VR has three related properties: communication, involvement, and thoughts10. The level of engagement and strength of interaction with simulated objects directly affect the users' imagination in the virtual atmosphere. Depending on the stages of immersion, VR schemes are classified into non-immersive VR, semi-immersive VR, and fully immersive VR, each of which utilizes varying combinations of communication methods, strategies, and boundaries. 11 Non-immersive VR allows communication with virtual surroundings on desktop or touch screens on handheld devices through a mouse, and signs12,13,14. Semi-immersive VR commonly comprises a large screen, organization, and screens, providing a cinematic experience, but often lacks positional tracking when used in groups. Gesture-based connections have been researched since the beginning of VR15 to render the use of VR technology comfortable and familiar. In a gesture-based interface, physical body signs like hand movements are used to communicate with a computer system.

A gesture border is a common realization of the more general principle of a natural user interface (NUI), which can be applied to achieve interface transparency in HMD-based VR. The aspects that promote occurrence in VR were investigated earlier 16,17; nevertheless, there are few studies on occurrence in the background of gesture-based UI. Understanding the influences that can impact the occurrence of VR using a gesture-based UI can be beneficial during the design of upcoming immersive technologies. Therefore, from the perspectives of occurrence and gesture-based communication, we investigated the user experience of an immersive VR engagement based on gesture-based UI. For the socially vulnerable, 360° highway assessment material, a subsection of virtual reality (VR), can be an effective tool. With such technology, users can virtually explore the globe without physical movement, travel to different locations, and engage in social communications with other users, complete with multi-party functionality. For the physically and environmentally challenged, this involvement can be a useful supplement to outdoor activity. Beyond simple navigation, the educational possibilities enabled by 360° opinion technology extend to multiple areas of education.

A systematic review of the literature examining 64 training sessions emphasized the informative uses, benefits, and communication properties of 360° VR videos and actual VR. Results indicate that 360° VR can enhance the presentation, motivation, and information retention of learners, promoting greater immersion. The researcher21 associated two one-handed and two-handed UIs in a teleportation-based movement scheme. The results show that one-handed signals were perceived as quicker to start by those who preferred them, while two-handed signs were considered more reliable, even though the item discusses teleporting alone and not constant movement, as used in this work.

The assumption of higher consistency for the two-handed signs and the higher initial rapidity for the one-handed signs was established and utilized as a premise for this study. A more detailed specification of the technology assessment restrictions, beyond presentation and serviceability, regarding training operators to fulfill responsibilities in VR22 is also required. For this training situation, the panels instruct the operators through activities, with technology assessment being a crucial component. Furthermore, the fundamental premise is that the HT technology, if properly applied, can provide advancements also in the domain of technology23,24. It has been shown in recent research that virtual environments are not only useful for evaluating wheelchair steering services but also possess beneficial uses, including training in flexible services that would prove challenging or expensive to recreate in real life. For instance, the author25 points out that VR simulators can facilitate the generalization of steering services to physical settings by providing a skillful and harmless technology environment.

The researcher also highlights the application of virtual skills as an intervention program for wheelchair services training, emphasizing its utility and beneficial value in individualizing exercise and enhancing flexibility services. Nevertheless, for safety purposes, users are confirmed and screened beforehand, particularly in virtual reality surroundings26,27. Such an implementation allows people with physical disabilities to practice and refine their wheelchair-driving skills in a secure and controlled environment28. They are able to experience various virtual tests and scenarios, including overcoming challenging problems, navigating tight spaces, or executing intricate maneuvers. Meanwhile, users can receive immediate feedback on their performance and acquire new movement strategies29. Besides serving as a training aid, virtual reality has therapeutic applications30. Offering an immersive pictorial technology can contribute to enhancing self-assurance, organization, and physical reintegration. Individuals can develop their motor skills and improve their stability and steadiness. Virtual reality environments are computer-generated representations that enable users to feel as if they are part of them31.

This situation is visualized with virtual reality headphones or other strategies. Mockups can simulate risk circumstances or challenging situations prior to their implementation, such as medical events or the operation of certain strategies32,33,34. Yet, an additional benefit is the ability to attach device equipment, such as controls, motion-detecting ornaments, microphones, or vests, to gain efficient communication between the employer and mechanism35 in a harmless and carefully replicated setting prior to its use in actual settings. Virtual substances can be operated in different modes, such as using joysticks, speech instructions, or filmed trails, as in a Kinect study36. All the alternatives involve learning for users, so the user's experience is a critical consideration in selecting the control approach37.

Although gestures and voice interaction in virtual spaces have been investigated for navigation tasks, including the street-view system discussed in the base paper, there is a significant gap in the application of these modalities to precision-driven and creative areas of work, such as immersive home design. The foundational study proved multimodal input viable for accessibility, but it encompassed only navigation and scene examination, utilizing two-handed gestures, straightforward command issuing, and a small participant population. It did not encompass difficulties of high-fidelity manipulation, precise spatial positioning, material editing, or collaborative workflows, which define interior and architectural design activities. In addition, the testing targeted a narrow, homogeneous group of users, which limited observations on usability for broader populations and different accessibility needs. Current VR design tools are primarily based on controllers, which restrict natural interaction and exclude users with motor or learning disabilities. As a result, there is a lack of designing and testing a system which integrates gestures and speech for coarse and fine object manipulation, provides voice-controlled numeric changes, accommodates one-handed and accessibility-conducive interaction styles, and is tested with extensive user experience analysis in actual design tasks38,39. Closing this gap can create a new, user-centric approach to immersive home design that builds upon previous research in multimodal VR navigation.

Recent developments in multimodal human-machine interfaces have explored advanced sensing and interaction mechanisms to enhance the usability and immersion of VR40. The author41 presented an electret-nanofiber-based triboelectric sensor for non-contact 3D gesture recognition, demonstrating that high-fidelity interpretation of gestures can be achieved without traditional vision-based tracking and enabling more natural, touchless VR control. On the other hand, the researcher42 presented a robust behavioral training procedure for head-fixed VR systems in mice. The study demonstrated how controlled VR environments can support precise behavioral measurement and stable interaction protocols, and outlined the importance of repeatability and systematic design in VR workflows. Concurrently, the author43 proposed a stretchable and adhesive ionic-conductive electrode with ultra-low skin impedance for improved electrophysiological signal acquisition; this opens the way to biointegrated sensing that could further increase the interaction fidelity of XR and VR. In fact, all these studies show how rapidly sensor technology, training protocols, and physiological interfaces evolve-all because there is a need for more flexible, multimodal interaction frameworks, such as the proposed gesture-voice system in this study.

This protocol presents a multimodal gesture-voice pathway that enables one-handed accessible gestures, voice-based numerical improvements, and precision-driven house design, in contrast to current VR interaction models, which primarily support navigation or controller-based manipulation. To our understanding, this is the first protocol that combines gesture-voice fusion, specifically tuned for spatial precision, material modification, and inclusive design processes.

The primary objectives of this work are to create a gesture- and voice-enabled virtual reality (VR) system tailored for engaging home design activities, and to expand the multimodal input capabilities to support accurate furniture placement, rotation, scaling, and material or lighting editing in addition to typical navigation functions. A central aim is to develop a multimodal fusion engine that integrates gestures, voice, and contextual inputs to achieve both coarse and fine-grained control during interior design tasks. The system further introduces adaptive interaction modes, such as one-handed gestures, accessibility-focused features, and voice-enabled precision commands, to support a wide range of user groups.

This work also seeks to conduct a comprehensive user-experience (UX) assessment that blends quantitative measures, including task completion time, placement accuracy, and error rate, with qualitative evaluations of usability, workload, and user satisfaction. A significant advancement over previous studies is achieved by shifting from street-view-based navigation systems to a task-oriented and precision-driven interior design framework. The key contribution lies in the development of a gesture- and voice-controlled VR environment for immersive home design, extending the scope of earlier research that focused mainly on gesture-voice integration for street-view navigation.

Importantly, no prior system has effectively combined gesture-voice multimodal communication specifically for precision-dependent home design operations, such as object positioning, scaling, rotation, and material editing. Existing approaches also lack adaptive interaction alternatives, including one-handed gesture input or voice-based numerical refinement. The proposed system introduces a comprehensive multimodal interaction paradigm that fuses gesture recognition -- achieved through MediaPipe Hands or VR headset-based hand tracking -- with Temporal Convolutional Network classifiers, as well as speech recognition and natural-language interpretation driven by Whisper ASR and transformer-based intent classification models.

To ensure robust operation, a decision-level multimodal fusion mechanism with confidence-based arbitration is implemented, mapping gestures to spatial selection tasks and voice commands to precision refinements. This design enables accurate placement, scaling, rotation, and material changes of objects in virtual home-design environments. The workflow follows a structured pipeline that begins with defining system requirements, designing gesture and voice lexicons, developing multimodal fusion techniques, and constructing an immersive VR environment in Unity 3D equipped with snapping tools, measurement overlays, and material previews. This is followed by system integration, pilot testing, and user-experience evaluation through comparative studies.

The system offers low-latency multimodal interaction (below 200 ms), adaptive accessibility support, and improved usability compared to conventional controller-based VR systems. Experimental evaluation will quantify task completion time, placement accuracy, error rates, SUS scores, NASA-TLX workload metrics, and IPQ presence ratings. The expected results demonstrate that multimodal interaction significantly enhances precision, efficiency, and user satisfaction in VR-based home design.

Access restricted. Please log in or start a trial to view this content.

Protocol

Loading...
$$\rightleftharpoonup{xx}$$ $$\longleftharp{xx}$$, $$\longrightharp{xx}$$,

This study was approved by the Research Ethics Committee of Taylor's University, Subang Jaya, Malaysia. All procedures adhered to the ethical guidelines set by the institutional research committee and conformed to the Declaration of Helsinki. Prior to data collection, informed consent was obtained from each participant. They were clearly briefed on the study's objectives, assured that participation was voluntary, that their responses would remain confidential, and that they could withdraw at any point without any consequences. Written informed consent was also obtained for the use of anonymized interaction and survey data in academic publications. No personally identifiable information, images, or sensitive personal data are included in this manuscript.

The proposed work formulates a gesture and voice-based virtual reality (VR) system designed for immersive home design, surpassing current methods geared towards navigation-oriented tasks. The approach begins with identifying user needs and design activities, including furniture deployment, object scaling and rotation, material editing, and precise room measurements. A gesture interaction system is then developed, where unconstrained hand gestures (such as pinch, grab, and rotate) are detected via headset-based tracking or MediaPipe Hands and identified by lightweight machine learning models, including Temporal Convolutional Networks (TCN). To complement the gestures, an organized voice command grammar is established, combined with Whisper-based Automatic Speech Recognition (ASR) and Natural Language Understanding (NLU), to derive intents and parameters for high-level semantic control. The two modalities are combined in a multimodal fusion strategy, where gestures are used to manage spatial selection and manipulation, and voice commands are used to offer precise refinements, augmented by a confidence-based arbitration mechanism and error recovery through undo commands.

The immersive design environment is deployed in Unity 3D, featuring snapping, collision response, measurement overlays, and an interactive furniture and material library to enable realistic design workflows. Pilot testing is conducted to ensure the naturalness of interaction, stability, and low latency (<200 ms) before formal evaluation.

Participants and sampling strategy
In total, 48 participants were recruited for the user study to ensure diversity and the generalizability of the results. The participants' ages ranged from 19 to 62 years (M = 32.4, SD = 10.7), with 26 males and 22 females. To capture variation in familiarity with immersive technologies, the participants were divided into three groups of VR experience: novice users (n = 18) with little or no prior exposure to VR, intermediate users (n = 17) with infrequent use of VR, and experienced users (n = 13) who used VR regularly, mainly professionally. To ensure the accessibility and inclusiveness of the sample, it also included 6 participants with mild mobility limitations and 4 participants who preferred one-handed interaction due to motor constraints. All participants had either normal or corrected-to-normal vision, with none reporting vestibular or neurological conditions that could interfere with their use of VR.

Participants were randomly assigned to the three interaction conditions using a balanced allocation strategy: Gesture+Voice (n = 16), Controller-only (n = 16), and Hybrid (n = 16). Their experience levels were distributed evenly across the three groups to avoid bias and ensure comparable task difficulty across the conditions. All participants provided written informed consent prior to the commencement of the experiment.

A user study is then carried out across three interaction conditions, such as Gesture Voice, Controller-only, and Hybrid, to evaluate precision, efficiency, and user experience. Quantitative metrics include placement precision, task duration, and error rates, while subjective assessment utilizes SUS, NASA-TLX, and IPQ, supported by qualitative interviews. This system, presented here, offers three main novelties: the utilization of multimodal gesture-voice fusion for accurate object manipulation in VR, adaptive and accessible interaction modes such as single-hand gestures and voice fallback, and the provision of a reproducible corpus of annotated gesture-voice commands. In combination, these steps ensure enhanced accuracy, effectiveness, and user experience, directly addressing the base paper's shortcomings and advancing research in immersive VR interaction for design use. Figure 1 shows the architecture of the proposed method.

The system architecture of the proposed method, as shown in Figure 1, begins with input modalities, where gestures are recorded using datasets such as EgoGesture and IPN Hand, and voice instructions from the Fluent Speech Commands dataset. These inputs are processed independently using gesture recognition and voice recognition modules to produce spatial and semantic data. The two streams are fused within a multimodal fusion layer, which aligns gesture and voice inputs to produce consistent commands. This fused data is utilized in the VR interaction layer to facilitate users in manipulating objects by placement, rotation, scaling, and material modification in real time. Lastly, the processed interactions are made available in the output layer, creating an immersive home design space that is focused on accessibility, accuracy, and natural interaction.

System requirements and task definition
The envisioned system enhances VR navigation models with accessibility orientation to facilitate precision-based home planning. We identify the user domain as Vector set notation, U={un, ud, ua}, formula for mathematical analysis., where u_h, mathematical symbol, numerical analysis concept are homeowners, u_h, mathematical symbol, numerical analysis concept are professionals in design, and Static equilibrium equation ΣFx=0 diagram; vector forces; educational physics concept. are accessibility users (aged or mobility impaired). Each user performs a set of core tasks T = {place, rotate, scale, material, light, measure}. A task T is performed via multimodal inputs: gesture stream Gt (spatial manipulation) and voice command Vt (semantic parameters). The system relates these to an action through a fusion policy:

Policy optimization equation, \( a_t = \pi(G_t, V_t) \), formula, mathematical analysis.    (1)

where π is the decision-level multimodal function.

System needs require low-latency interaction: let τg, τv, τf symbols; identify variables in dynamics or kinetic studies. be gesture recognition, speech understanding, and fusion times.

Equation: τ_tot = τ_g + τ_v + τ_f ≤ 200 ms; relates to timing constraints in systems analysis.    (2)

The sum response provides real-time responsiveness in accordance with immersive VR usability thresholds.

For task accuracy, let the ground-truth object state be (x, R, s, M, L) (position, rotation, scale, material, length) and the system's realized state be Mathematical symbols with hats: \( \hat{x}, \hat{R}, \hat{s}, \hat{M}, \hat{L} \); statistical notation.. Error metrics are as follows:

Equation for position error: e_pos = ||x̂ - x||₂, related to vector norm operations.   (3)

Orientation error calculation, equation: \(e_{ori} = \arccos\left(\frac{1}{2} \text{tr}\left(R^T \hat{R}\right) - 1\right)\).    (4)

Equation for error scaling; \( e_{\text{scale}} = \| \hat{s} - s \|_2 \); mathematical formula.    (5)

Error calculation equation: e_means=|L̂−L|, mathematical concept.    (6)

With Equation e<sub>mat</sub>=1<sub>M±M̂</sub> related to matrix mathematics concepts. for material mismatch.

Both modalities generate confidence scores cg (gesture) and cv (voice), normalized such that cg + cv = 1. Fusion arbitration minimizes a weighted error:

Optimization formula, argmin process, mathematical equation for computational analysis.   (7)

Where lg, lv are modality-specific losses. If Maximization formula, max(cg cv) < η, static equilibrium analysis equation., the system requests a clarification prompt, making it robust for unclear commands.

Lastly, evaluation metrics include task completion time Tt, error rates Et, and workload scores from NASA-TLX, SUS, and IPQ. We introduce a composite task performance index:

Optimization formula, J=Σt∈T(αt(Tt/Tmax)+βt(||Et||2/Etmax)), mathematical equation.    (8)

to be minimized across users, with thresholds that necessitate a SUS score of ≥70 and a NASA-TLX reduction compared to the baseline controller-based VR.

Dataset collection and preprocessing
The proposed system is based on both gesture data (for TCN-based recognition) and voice command data (for NLU/ASR). For this purpose, we utilize three main datasets: EgoGesture for dynamic continuous gestures in VR-like environments, IPN Hand to supplement natural short-range gestures, and Fluent Speech Commands44 to train voice-led semantic comprehension of design utterances. Optional datasets are HaGRID for static gesture detector pretraining, and Mozilla Common Voice for general ASR pretraining, but neither of these are necessary. Such a choice allows for both modalities to be covered while maintaining the scope of the datasets as reasonable and reproducible as possible. Table 1 shows the dataset details of the proposed work.

The dataset selection in this paper is crucial, as each dataset is designed to address a distinct scenario of multimodal VR interaction. The EgoGesture dataset provides a large-scale, naturalistic hand-gesture collection, ensuring high recognition performance despite varying lighting and environmental conditions. Since the IPN Hand dataset is designed for unrestricted, continuous hand motions, it would be especially well-suited for high-precision control operations and real-time gesture segmentation. Additionally, the Fluent Speech Commands dataset offers labeled voice commands with semantic annotations for intent detection, which is crucial in enabling accurate and natural command execution in VR. Combined, these datasets provide complementary coverage of both gesture and speech modalities, providing diversity, robustness, and everyday performance that improve the evaluation of the proposed system. In addition to public datasets, we also verified the protocol using an independent in-house dataset comprising 312 gesture samples and 128 voice commands from 10 new participants. This independent validation confirmed the protocol's reproducibility and robustness.

Data preprocessing
The gathered information is systematically normalized, sampled, and augmented prior to model training:

Gesture sequence normalization
A gesture video sequence is represented as a multivariate time series of skeletal joint features:

Equation of time series analysis with mathematical symbols, X(i) in dimensional space diagram.    (9)

Where Ti is the length of the i-th gesture and d is the feature dimension (e.g., hand joint locations). As different gestures can be of varying lengths Ti, we resample them into a constant sequence length T:

Mathematical sequence equation X̃ᵢ={xᵢ(t)ᵀ} for algorithm analysis and data prediction.     (10)

This makes all gesture samples, irrespective of the recording length, comparable to a standard temporal length that can be used for TCN training.

Feature standardization
To eliminate scale differences between users and recording conditions, each feature space is standardized:

Normalization formula for standardizing data, equation diagram with statistical variables.     (11)

Where Statistical average formula μj=1/NT∑i=1^N∑t=1^Tx(t,i), equation for data analysis algorithm.  is the mean of feature j, and σj is its standard deviation. Normalization in this way ensures that features are centered with unit variance, and individual differences in gesture performance are minimized.

Data augmentation
The inclusion of synthetic perturbations improves robustness to variability. Temporal jitter first randomly offsets the sequence alignment:

Stochastic process equation t, Δt~U(-δ, δ), symbolizing random time increment.    (12)

where δ is the maximum jitter in frames. Noise injection second adds Gaussian perturbations to features:

Stochastic process formula, equation for Gaussian noise, statistical analysis concept.   (13)

These augmentations enable the model to be robust to timing irregularities and sensor noise in actual VR interactions.

Voice preprocessing
Each spoken utterance is initially mapped to a log-Mel spectrogram representation:

Log-Mel spectrogram formula, V(i)=LogMel(a(i)), mathematical equation for audio signal processing.   (14)

Where F is the quantity of frequency bins and T′ the quantity of temporal frames. For compensation for recording variability, spectrogram features are normalized as:

Statistical normalization formula \( \hat{V}^{(i)} = \frac{V^{(i)}-\mu_V}{\sigma_V} \) for data analysis.    (15)

Where μV, σV are the global mean and standard deviation over the corpus. This provides a stable acoustic feature space for ASR and NLU models.

Intent-slot encoding
In addition to the acoustic features, each voice command is labeled with structured semantic slots:

Static vector formula: S^(i) = [onehot(action), onehot(object), onehot(location)].  (16)

where, for instance, "rotate chair 15 degrees" corresponds to (action = rotate, object = chair, location = null). This systematic encoding enables the system to derive design-specific parameters and transform them into executable VR operations.

Gesture Interaction Design with Temporal Convolutional Network (TCN)
Gesture Interaction Design specifies how users interact with virtual objects in the VR home design system using natural hand gestures. Design starts with developing a gesture lexicon - an association of hand gestures to system functions. For instance, a grabbing gesture could control object placement, a pinching gesture could control scaling, and a rotation gesture could turn objects around a selected axis.

In the system under consideration, EgoGesture and IPN Hand serve as the base gesture datasets for training the Temporal Convolutional Network (TCN), while Fluent Speech Commands supplement the design with voice-enabled interaction capabilities. Combined, they provide multimodal VR interaction.

Gesture representations from the dataset
A sequence of gestures from either EgoGesture or IPN Hand is a multivariate time series:

Time series data formula X(i)= {x1(i), x2(i),..., xT(i)}; mathematical equation.     (17)

Dynamic system state variable \(x_t^{(i)}\in \mathbb{R}^d\); mathematical expression.     (18)

Here, T is the predetermined number of frames (after resampling), and d is the number of skeletal feature dimensions (e.g., 3D joint locations). The dataset has labels y(i) ∈ y indicating one of C gesture classes.

Temporal Convolutional Encoding
We utilize a Temporal Convolutional Network (TCN) to classify gestures in real-time. TCN captures short- and long-term dependencies in gesture sequences. TCNs, unlike RNNs, utilize dilated causal convolutions, which enable efficient parallel computation.

The TCN applies these sequences through causal dilated convolutions to learn short-term and long-term dependencies:

Recurrent neural network equation diagram: \(h_t^{(l)} = \sigma(\sum_{k=0}^{K-1}W_k^{(l)}h_{t-d_{l,k}}^{(l-1)} + b^{(l)})\).    (19)

where dL is the dilation factor in layer l, K is the kernel size, and σ is a non-linear activation (ReLU). The receptive field depends on:

Equation for calculating R in statistical models, featuring summation notation; formula illustration.   (20)

requiring the TCN to span gestures of different temporal lengths.

Classification of gestures
The last hidden state is passed through a Softmax classifier:

Neural network output formula \( \hat{y}^{(i)} = \text{Softmax}(W_c h_T^{(L)} + b_c) \) equation.  (21)

with training objective defined by cross-entropy:

Logistic regression loss function formula, mathematical equation for machine learning analysis.  (22)

Integration with voice commands
While spatial operations (move, rotate, scale) are addressed by gesture recognition of EgoGesture and IPN Hand, Fluent Speech Commands dataset is applied for semantic commands. Every utterance is translated into a structured slot representation: One-hot encoding formula; S^(i) = [onehot(action), onehot(object), onehot(location)] equation.  which augments the gesture input by offering parameters (e.g., "rotate chair by 15°").

Then both outputs are combined:

Fusion of Temporal Convolutional Networks and Nonlinear Units formula; educational math concept.   (23)

Where o(i) is the VR control primitive (placement, scaling, rotation).

The Temporal Convolutional Network (TCN) is effective in handling sequential data by employing dilated 1D convolutions on input sequences, such as gesture or speech signals, as shown in Figure 2. In contrast to recurrent models, TCNs utilize convolutional filters to learn both short-term and long-term temporal relationships. Residual blocks with skip connections are employed to provide stability during training and to prevent vanishing gradients, whereas the use of stacked TCN layers enables the network to learn multi-scale temporal patterns. Ultimately, a fully connected classification layer projects the learned temporal features into output labels, such as gesture classes or spoken intents. This design is particularly well-suited for continuous gesture recognition and voice command comprehension, as it combines computational efficiency with strong temporal modeling capabilities.

Table 2 summarizes the overall architectural design and training configuration of the Temporal Convolutional Network used in the proposed system for gesture recognition. The model features four residual TCN blocks, each based on dilated causal convolutions, allowing the network to model both short- and long-range temporal dependencies. These dependencies are crucial for distinguishing between dynamic gestures, such as grabbing, rotating, and pinching. With a kernel size of 3, and dilation factors, the dilation factors of [1,2,4,8] efficiently expand the temporal receptive field without increasing computational cost. To improve generalization across users and recording conditions, layer normalization and a dropout rate of 0.25 were utilized within each block. For training, an Adam optimizer with a learning rate of 1 × 10-3, a batch size of 32, and a sequence length of 64 frames is used, which optimally balances accuracy and real-time responsiveness. During inference, a sliding-window strategy makes the prediction of gestures every 20-25 ms feasible, while maintaining end-to-end latency under 120 ms. This combination of architecture and hyperparameter yields high accuracy in gesture classification (∼94% placement accuracy) while preserving real-time interaction, enabling TCN to perform immersive VR manipulation tasks.

Voice Interaction Design (ASR+NLU)
The voice interaction component supplements gesture-based spatial control through semantic and parameterized voice commands for the VR home design system. The spoken input a(i) is first recorded as raw audio and converted to a log-Mel spectrogram:

Log-Mel spectrogram equation V(i)=LogMel(a(i)) in R^{F×T}; mathematical concept.  (24)

Where F is the number of frequency T bins the length of time intervals. The representation preserves both spectral and temporal patterns of speech and is normalized to be invariant to speaker and environment variability:

Normalizing values formula, V(i)-μV/σV, symbols equation for data standardization.  (25)

The normalized spectrogram is fed into an Automatic Speech Recognition (ASR) model like Whisper, which transforms the acoustic features into a token sequence:

Mathematical formula for set Z in symbolic analysis, showcasing ASR function and data variables.  (26)

where each is a token for a word or sub-word unit. Such transcription is fed into a Natural Language Understanding (NLU) model, which captures structured meaning. More precisely, NLU makes predictions of intent categories and semantic slots:

One-hot encoding equation, S^(i)=[onehot(action),object,location], for data representation.  (27)

where, for example, the statement "rotate chair fifteen degrees" maps to (action = rotate, object = chair, location/parameter = 15°).

The identified intent-slot framework is then translated into a mathematical command that can be implemented within the VR system:

Mathematical equation, u^(i)=Map(s^(i)), transformation concept, educational diagram.    (28)

Where u(i) can be a numerical transformation (e.g., 15° rotation) or a categorical change (e.g., material exchange).

Robustness is finally achieved through a confidence-based arbitration mechanism. If the ASR confidence score p(z) or NLU confidence score p(s) is less than a threshold τ, the system either requests clarification or defaults to gesture-based control. This ensures that voice commands are precise and clear, improving usability for design tasks that require precision.

Multimodal fusion design
The advantage of multimodal fusion in the proposed VR system is that voice comprehension and gesture detection are integrated into a single decision-making process rather than being handled separately. The gesture module gives a classification output ŷ<sub>g</sub> symbol, regression prediction, statistical model, data analysis result, equation representation  from the TCN, whereas the voice module delivers an intent-slot representation s(i). To fuse these, the system implements decision-level fusion with confidence-based arbitration.

Make the gesture classifier provide a probability distribution over gesture classes:

Softmax function equation; probability calculation formula; statistical analysis; P(c|X)=softmax.  (29)

and let the NLU model yield intent-slot probabilities:

Probability equation diagram; Pv(a,o,p|z) = fNLU(z); used in statistical analysis.  (30)

Where a, o and p refer to action, object, and parameter slots derived from speech input. The fusion step calculates a joint decision by combining the two modalities according to their confidence scores:

Maximum likelihood estimation equation, argmax O, probability distribution, shown as formula.  (31)

Where α∈[0,1] is set dynamically based on system confidence (e.g., higher when gestures are more certain, lower when speech is clear).

Algorithm 1: Multimodal_Fusion (X, a):

Input:

X = preprocessed gesture sequence, a = audio command

Output:

O = VR action command

Step 1: Gesture Recognition

Gesture recognition, TCNClassifier equation, AI model, algorithm analysis.  

Step 2: Voice Understanding

Equation depicting ASR method for text conversion, highlighting linguistic process analysis.  

Natural language understanding (NLU) formula, diagram for text analysis, intent detection, confidence.  

Step 3: Fusion Decision

Conditional logic formula, If(confG > τ and confV > τ), text symbol, threshold evaluation.  

Equation of confidence factor calculation; formula \( a \leftarrow \frac{conf_g}{(conf_g + conf_v)} \).  

Weighted combination formula, O=weighted combine(gesture_label, intent, a).  

Programming conditional statement, equation: else if(conf_g >= conf_v), code logic analysis.  

Equation depicting gesture_label assignment process with symbolic representation for data classification.  

else:

Equation symbol O←intent; Intention representation; Static equilibrium concept.  

Step 4: Error Handling

Mathematical equation, conditional logic, conf_g < τ, conf_v < τ, programming syntax.  ):

Request user clarification

else:

Send O to VR system

return O

This Algorithm 1 switches between gesture input and voice input to formulate a powerful command for VR design work. The gesture module (TCN) first interprets hand movements such as rotation or scaling and assigns a confidence score based on the model's certainty. Parallel to this, the voice module also performs speech-to-text conversion with ASR (e.g., Whisper) and applies NLU to identify semantic intent, such as "rotate chair by 15°." Both modules return their predictions along with confidence scores. The combination step then balances the two outputs. If both inputs are definite, the system combines them proportionally to their confidence levels. If one input is ambiguous (e.g., noise in voice or uncertain gesture), the system prioritizes the other more. Lastly, if neither is definite, the system asks for clarification to prevent errors. This renders the fusion adaptive (weights change based on input quality), robust (accommodating noisy or missing data), and accurate (utilizing the best strengths of both gesture and voice).

Participants
This study includes a group of volunteers chosen from both university students and design-related professionals to gather diverse opinions on VR usability. Participants' ages ranged from early adulthood to middle age, representing a mix of prior VR experience levels, including beginners, occasional users, and advanced users. All participants had either normal or corrected vision and no reported motor impairments that could substantially interfere with gesture-based interaction. Severe speech impairments, major arm mobility limitations, or any previous history of VR-induced motion sickness excluded individuals for safety and consistency in participation. All participants provided informed consent, and the procedures for the evaluation were approved by the institutional review board. This subsection clearly identifies who participated in the evaluation, allowing for the reproducibility of the study.

Experimental setup
The proposed experimental setup features the developed multimodal VR system, which is deployed on a consumer-grade VR headset equipped with integrated hand-tracking capabilities and an RGB camera for capturing gestures. It features a high-performance desktop, enabling the real-time processing of gesture recognition, ASR, and multimodal fusion techniques. The VR environment is built in Unity 3D and contains a home-design room that can be configured with furniture, materials, and lighting elements. The setup encompasses Whisper-based ASR for speech transcription, and it also includes a transformer-based NLU module to detect intent. This description has been moved from within the Results section to provide an appropriate technical environment overview before discussing the outcomes.

Subjective testing scenarios
Participants interacted with a simulated home-design environment consisting of a furnished living room, a bedroom, and a small workspace area. In every scenario, users were asked to change virtual objects and environmental variables in the surroundings according to realistic design circumstances. For example, participants were asked to position a sofa with respect to a reference point, rotate a chair in a specific direction, resize a table to a target dimension, or modify the wall material/light profile. These scenarios were selected because they represent typical interior design tasks that challenge both gesture-based control (in terms of spatial precision) and voice-based control (in terms of semantic commands). This subsection directly addresses the reviewer's question about the subjective testing conditions under which users evaluated the system.

Task design
Participants undertook a structured set of exercises addressing different areas of item manipulation and design interaction. The major tasks were as follows: 1) Furniture Placement: Participants positioned objects at target coordinates. 2) Rotation Task: In this task, users had to change the orientation so that it matched a reference angle. 3) Scaling Task: This involved resizing furniture to preset dimensions. Moreover, 4) Material/Color Editing: Participants employed voice commands to change the textures or modify the lighting. Optional navigation tasks were provided for capturing the spatial awareness capability within the virtual room. Each work was established with clear goals, quantifiable outcomes, and a defined time limit, ensuring that both accuracy and efficiency could be properly evaluated.

Procedure
To ensure uniformity among participants, the evaluation was conducted in a planned sequence. Users were first explained the system and given a brief demonstration of the gesture and speech interaction modalities. An initial calibration phase allowed adaptation for individual hand postures and speech characteristics. Users then had a short practice session to get accustomed to both modalities. The actual evaluation required performing all tasks in three interaction conditions: using a controller only, using gestures only, and using gestures plus voice as multimodal input. The condition order was counterbalanced to reduce learning effects. At the end, subjects completed standardized evaluation instruments, such as the SUS, NASA-TLX, and IPQ, and participated in a brief qualitative feedback interview.

Evaluation metrics
This structured procedure establishes methodological transparency, enabling replication studies. Objective and subjective metrics were combined to comprehensively assess system performance. Objective metrics included task completion time, a measure of efficiency; placement accuracy, quantified as the deviation from the target position or rotation; and error rate, defined as the number of input misclassifications or unintentional actions performed by the system. Subjective metrics were captured using validated questionnaires: the System Usability Scale to measure perceived usability, the NASA-TLX scale to assess mental and physical workload, and the I group Presence Questionnaire for measuring immersion and presence in VR. These have been moved appropriately from the Results section to describe how performance was measured before presenting any outcomes.

VR environment, system integration, and evaluation
The primary platform for speech and gesture commands is an interactive virtual reality environment where the system is deployed. Integrated libraries of furniture, texture, and lighting elements are used to create virtual rooms in Unity 3D. It enables users to organize, customize, and edit items in real time. Accuracy is enhanced through snapping capabilities, collision detection, and measurement overlays, allowing objects to align naturally. Real-time rendering of lighting and materials enhances the design experience by providing instant feedback on every interaction.

After validating the individual modules for gesture recognition, voice understanding, and multimodal fusion, they are combined into a unified pipeline. The system is designed to have low latency, so user commands appear nearly immediately in the VR environment. Pilot testing is conducted with a minor number of members to hone the interaction flow. The iteration includes gesture vocabularies, validation of voice command grammar, and ensuring the system feels natural and reliable during continuous usage.

A comprehensive study comparing three modalities of interaction, such as multimodal gesture-voice fusion, controller-only input, and hybrid mode, will be used to evaluate the user experience. Tasks, including material switching, rotation, and object placement, are required of participants. To assess performance, objective metrics are collected, such as job achievement time, placement correctness, and error charges. Subjective ratings are elicited from participants using standardized questionnaires, such as the System Usability Scale (SUS), NASA-TLX cognitive workload, and IPQ questionnaire for immersion, with participant interviews providing additional support. This dual assessment encodes both quantifiable performance measures and subjective user experience.

The advance of this system is to apply multimodal fusion to precision design tasks beyond basic navigation or interaction. The technique offers a more organic, accurate, and approachable way of interacting by combining the semantic value of voice instructions with the spatial advantages of gestures. By offering adaptable accessibility features, such as voice-only fallback and single-hand gesture recognition, the system also enhances inclusivity. Finally, the study offers a methodical, annotated multimodal dataset of voice-gesture interactions, promoting future studies in immersive interface design and bolstering reproducibility.

Figure 3 displays typical screenshots of the implemented program, providing a more vivid representation of the built VR interior design system. The immersive virtual reality environment, where users can explore and view the room's layout, is depicted in Figure 3A. Figure 3B shows how natural hand interactions can be used to select, move, and rotate virtual furniture items in a gesture-based object manipulation process. Figure 3C illustrates the integrated voice interface, where the microphone and speech-feedback symbols validate user-issued voice commands for spatial modifications or object placement. Taken together, these screenshots of the proposed multimodal VR system demonstrate that it has been fully implemented, allowing for real-time gesture and voice interaction.

Access restricted. Please log in or start a trial to view this content.

Results

Loading...
$$\rightleftharpoonup{xx}$$ $$\longleftharp{xx}$$, $$\longrightharp{xx}$$,

The designed multimodal VR home design system was validated with three fundamental datasets: EgoGesture (dynamic gestures), IPN Hand (ongoing hand gestures), and Fluent Speech Commands (voice commands with intent-slot labels). Collectively, the three datasets offered a comprehensive training and evaluation base, addressing both gesture recognition and voice-based semantic comprehension. The validation establishes that gesture-voice integration significantly enhances task accuracy and reduces workload compared to controll...

Access restricted. Please log in or start a trial to view this content.

Discussion

Loading...
$$\rightleftharpoonup{xx}$$ $$\longleftharp{xx}$$, $$\longrightharp{xx}$$,

The results of this research advance the technology of the base paper, which mainly showed the merits of immersive VR navigation for accessibility by combining gesture and voice modalities for design tasks. The multimodal system performed better on all objective and subjective performance measures, with shorter task completion times, higher placement accuracy, and lower error rates than controller-based and hybrid approaches. Moreover, measures of usability like SUS, NASA-TLX, and IPQ noted greater user satisfaction, dec...

Access restricted. Please log in or start a trial to view this content.

Disclosures

Loading...
$$\rightleftharpoonup{xx}$$ $$\longleftharp{xx}$$, $$\longrightharp{xx}$$,

The authors have no conflicts of interest to declare.

Acknowledgements

Loading...
$$\rightleftharpoonup{xx}$$ $$\longleftharp{xx}$$, $$\longrightharp{xx}$$,

The authors gratefully acknowledge the institutional support of the School of Architecture, Building & Design, and The Design School, Taylor's University, for providing resources and academic guidance during this research.

Access restricted. Please log in or start a trial to view this content.

Materials

List of materials used in this article
NameCompanyCatalog NumberComments
EgoGesture DatasetNanyang Technological UniversityDOI: 10.1109/TMM.2018.2808769Large-scale egocentric hand gesture dataset for dynamic gesture recognition training.
Fluent Speech Commands DatasetFluent.ai / University of AlbertaDOI: 10.48550/arXiv.1804.04363Dataset for voice intent and slot-based semantic understanding for ASR/NLU training.
HaGRID Dataset (optional)Sber AI / Open Sourcev1.1Optional dataset used for static hand gesture detector pretraining.
Igroup Presence Questionnaire (IPQ)igroup.orgN/AUsed for measuring immersion and presence in VR environment.
IPN Hand DatasetUniversidad VeracruzanaDOI: 10.1109/CVPRW50498.2020.00241Dataset for continuous, natural hand gesture recognition and segmentation tasks.
Matplotlib / SeabornOpen SourceLatest StableUsed for visualization of performance metrics and evaluation results.
MediaPipe HandsGoogle ResearchOpen-source (GitHub)Used for real-time hand and gesture tracking for gesture input recognition.
Mozilla Common Voice (optional)Mozilla FoundationVersion 17Optional dataset for pretraining ASR model on general speech data.
Multimodal Fusion LayerCustom algorithmN/AFuses gesture and voice input streams based on confidence-weighted arbitration.
NASA-TLX Workload AssessmentNASA Human FactorsN/AUsed for cognitive workload analysis of participants.
Natural Language Understanding (NLU) ModuleCustom (based on BERT / RoBERTa)N/AUsed to extract semantic intents and slots (action, object, parameter) from transcribed voice input.
NumPy / Pandas / OpenCVOpen SourceLatest StableUsed for numerical operations, data preprocessing, and video frame manipulation.
Python Programming LanguagePython Software FoundationVersion 3.10+Used for implementing ML models (TCN, ASR, NLU, fusion).
PyTorch Deep Learning FrameworkMeta AIVersion 2.0+Used for building and training gesture recognition (TCN) and NLU models.
RGB CameraLogitech / RealSenseLogitech C920 or Intel RealSense D435Used for hand gesture capture and motion recognition input.
System Usability Scale (SUS) QuestionnaireStandard Evaluation ToolN/AUsed for assessing subjective usability of VR system.
Temporal Convolutional Network (TCN)Custom implementation (PyTorch / TensorFlow)N/AUsed for dynamic gesture recognition from time-series skeletal data.
Unity 3D EngineUnity TechnologiesVersion 2022.3 LTSUsed for developing immersive VR environment with interactive furniture, material editing, and real-time rendering.
VR HeadsetOculus (Meta) / HTC ViveOculus Quest 2 or HTC Vive ProUsed for immersive visualization and spatial tracking during VR interaction.
Whisper ASR ModelOpenAIWhisper (base model)Automatic Speech Recognition engine for converting speech to text.
WorkstationCustom-built PCIntel Core i7 / NVIDIA RTX 3070 / 32 GB RAMUsed for real-time processing, training, and inference of gesture and voice models.

References

Loading...
$$\rightleftharpoonup{xx}$$ $$\longleftharp{xx}$$, $$\longrightharp{xx}$$,
  1. Kim, J., Ahn, J. -H., Kim, Y. Immersive interaction for inclusive virtual reality navigation: enhancing accessibility for socially underprivileged users. Electronics. 14, 1046(2025).
  2. Lee, W. -J., Kim, Y. -H. Does VR tourism enhance users' experience. Sustainability. 13, 806(2021).
  3. Embedding virtual and augmented reality in higher education in Yemeni universities. Nagi, G., Al-Fuhaidi, B. 2024 1st International Conference on Emerging Technologies for Dependable Internet of Things (ICETI), 15-17 January, Riyadh, Saudi Arabia, 1, 1-10 (2024).
  4. Narin, N. G. A content analysis of the metaverse articles. J Metaverse. 1, 17-24 (2021).
  5. Towards a social VR-based exergame for elderly users: an exploratory study of acceptance, experiences, and design principles. Shah, S. H. H., Hameed, I. A., Karlsen, A. S. T., Solberg, M. International Conference on Human-Computer Interaction, 1, Springer. Washington, DC, USA. 1-12 (2022).
  6. Zahedian-Nasab, N., Jaberi, A., Shirazi, F., Kavousipor, S. Effect of virtual reality exercises on balance and fall in elderly people with fall risk: a randomized controlled trial. BMC Geriatr. 21, 509(2021).
  7. Babadi, S. Y., Daneshmandi, H. Effects of virtual reality versus conventional balance training on balance of the elderly. Exp Gerontol. 153 (6), 111498(2021).
  8. Liang, H., et al. Metaverse virtual social center for elderly communication in time of social distancing. Virtual Real Intell Hardw. 5, 68-80 (2023).
  9. Shah, S. H. H., Karlsen, A. S. T., Solberg, M., Hameed, I. A. A social VR-based collaborative exergame for rehabilitation: co-design, development, and user study. Virtual Real. 27, 3403-3420 (2023).
  10. Chen, C. -H., Chen, M. -X. Effects of the design of overview maps on three-dimensional virtual environment interfaces. Sensors. 20, 4605(2020).
  11. Sudár, A., Csapó, A. B. Descriptive markers for the cognitive profiling of desktop 3D spaces. Electronics. 12, 448(2023).
  12. Wang, Y., Hu, Y., Chen, Y. An experimental investigation of menu selection for immersive virtual environments: fixed versus handheld menus. Virtual Real. 25, 409-419 (2021).
  13. Grabowski, A. Practical skills training in enclosure fires: an experimental study with cadets and firefighters using CAVE and HMD-based virtual training simulators. Fire Saf J. 125 (4), 103440(2021).
  14. Zhang, L., et al. A hybrid 2D-3D tangible interface combining a smartphone and controller for virtual reality. Virtual Real. 27, 1273-1291 (2023).
  15. Is it real? Measuring the effect of resolution, latency, frame rate, and jitter on the presence of virtual entities. Louis, T., Troccaz, J., Rochet-Capellan, A., Bérard, F. Proc 2019 ACM Int Conf Interactive Surfaces, 1, 5-16 (2019).
  16. Riches, S., Elghany, S., Garety, P., Rus-Calafell, M., Valmaggia, L. Factors affecting sense of presence in a virtual reality social environment: a qualitative study. Cyberpsychol Behav Soc Netw. 22 (5), 288-292 (2019).
  17. Weech, S., Kenny, S., Barnett-Cowan, M. Presence and cybersickness in virtual reality are negatively related: a review. Front Psychol. 10, 158(2019).
  18. Sae-Bae, N., et al. Emerging NUI-based methods for user authentication: a new taxonomy and survey. IEEE Trans Biom Behav Identity Sci. 1, 5-31 (2019).
  19. Appel, L., et al. Older adults with cognitive and/or physical impairments can benefit from immersive virtual reality experiences: a feasibility study. Front Med. 6, 329(2020).
  20. Pirker, J., Dengel, A. The potential of 360 virtual reality videos and real VR for education: a literature review. IEEE Comput Graph Appl. 41 (6), 76-89 (2021).
  21. Schäfer, A., Reis, G., Stricker, D. Controlling teleportation-based locomotion in virtual reality with hand gestures: a comparative evaluation of two-handed and one-handed techniques. Electronics. 10, 715(2021).
  22. Radhakrishnan, U., Koumaditis, K., Chinello, F. A systematic review of immersive virtual reality for industrial skills training. Behav Inf Technol. 40 (12), 1310-1339 (2021).
  23. Zhang, X., et al. How virtual reality affects perceived learning effectiveness: a task-technology fit perspective. Behav Inf Technol. 36, 548-556 (2017).
  24. Challenor, J., White, D., Murphy, D. Hand-controlled user interfacing for head-mounted augmented reality learning environments. Multimodal Technol Interact. 7, 55(2023).
  25. Budiharto, W. Intelligent controller of electric wheelchair from manual wheelchair for disabled person using 3-axis joystick. ICIC Express Lett B Appl. 12, 201-206 (2021).
  26. Letaief, M., Rezzoug, N., Gorce, P. Comparison between joystick- and gaze-controlled electric wheelchair during narrow doorway crossing: feasibility study and movement analysis. Assist Technol. 33 (1), 26-37 (2021).
  27. Venkatesan, M., Mohan, L., Manika, N., Mathapati, A. Voice-controlled intelligent wheelchair for quadriplegics. Computing, Communication, and Networking Technologies. 1, Springer. 41-51 (2021).
  28. Charlton, J., et al. Manual wheelchair skills training using virtual technology: a systematic review. Technical Report. , Flinders University. (2024).
  29. Genova, C., et al. A simulator for both manual and powered wheelchairs in immersive virtual reality CAVE. Virtual Real. 26, 187-203 (2022).
  30. Hoter, E., Nagar, I. The effects of a wheelchair simulation in a virtual world. Virtual Real. 27, 407-419 (2023).
  31. Improving wheelchair driving performance in a virtual reality simulator. Archambault, P. S., Bigras, C. 2019 Int Conf Virtual Rehabilitation (ICVR), Tel Aviv, 1, 1-2 (2019).
  32. Automatic synthesis of virtual wheelchair training scenarios. Li, W., Talavera, J., Samayoa, A. G., Lien, J. M., Yu, L. F. 2020 IEEE Conf Virtual Reality and 3D User Interfaces (VR), , 539-547 (2020).
  33. Yan, H., Archambault, P. S. Augmented feedback for manual wheelchair propulsion technique training in a virtual reality simulator. J Neuroeng Rehabil. 18 (1), 142(2021).
  34. Montalbán, M. A., Arrogante, O. Rehabilitation through virtual reality therapy after a stroke: a literature review. Rev Cient Soc Esp Enferm Neurol (Engl Ed). 52, 19-27 (2020).
  35. Jessica, P., Salim, S., Syahputra, M. E., Suri, P. A. A systematic literature review on implementation of virtual reality for learning. Procedia Comput Sci. 216 (1), 260-265 (2023).
  36. Efficacious opportunities and implications of virtual reality features and techniques. Nithva, B., Asha, V., Kumar, K., Giri, J. 2022 Fourth Int Conf Cognitive Computing and Information Processing (CCIP), , 1-8 (2022).
  37. The impact of controller type on video game user experience in virtual reality. Hufnal, D., Osborne, E., Johnson, T., Yildirim, C. 2019 IEEE Games, Entertainment, and Media Conf (GEM), , 1-7 (2019).
  38. Zhang, Y., Cao, C., Cheng, J., Lu, H. EgoGesture: a new dataset and benchmark for egocentric hand gesture recognition. IEEE Trans Multimed. 20 (4), 1038-1050 (2018).
  39. Egocentric gesture recognition using recurrent 3D convolutional neural networks with spatiotemporal transformer modules. Cao, C., et al. Proc IEEE Int Conf Computer Vision (ICCV), , 1-9 (2017).
  40. IPN Hand: a video dataset and benchmark for real-time continuous hand gesture recognition. Benitez-Garcia, G., et al. 25th Int Conf Pattern Recognition (ICPR), , 1-8 (2021).
  41. Lu, L., et al. Noncontact 3D gesture recognition enabled VR human-machine interface via electret-nanofiber-based triboelectric sensor. Nano Res. 18, 94907924(2025).
  42. Glorius, J. K., et al. Behavioral training procedures for head-fixed virtual reality in mice. J Vis Exp. (211), e67312(2024).
  43. Li, Y., et al. A stretchable, ionic conductive, and adhesive patch electrode with ultra-low on-skin impedance for electrophysiological signal recording. Sci China Inf Sci. 68, 129402(2025).
  44. Fluent speech commands: a dataset for spoken language understanding research. , Fluent.ai. https://fluent.ai/fluent-speech-commands-a-dataset-for-spoken-language-understanding-research/ (2025).

Access restricted. Please log in or start a trial to view this content.

Reprints and Permissions

Request permission to reuse the text or figures of this JoVE article

Request Permission

Tags

Gesture ControlVoice ControlMultimodal InteractionHand TrackingNatural Language InputPrecision Furniture PlacementSpatial Adjustment

Related Articles