Review Article

Seamless Multimodal Human-Robot Communication: Integration Techniques in Human–Computer Interaction

DOI:

10.3791/70218

June 9th, 2026

In This Article

Summary

Loading...
$$\rightleftharpoonup{xx}$$ $$\longleftharp{xx}$$, $$\longrightharp{xx}$$,

This review synthesizes recent advances in deep learning, multimodal sensing, and integration strategies that enable seamless, adaptive, and human-centered communication between humans and robots.

Abstract

Loading...
$$\rightleftharpoonup{xx}$$ $$\longleftharp{xx}$$, $$\longrightharp{xx}$$,

Seamless multimodal human-robot communication has become essential as social, assistive, and educational robots move into everyday settings like homes, healthcare facilities, and classrooms, where they need to integrate speech, gesture, gaze, facial expressions, and tactile cues with high precision, low latency, and real-world robustness. This review systematically examines how current techniques achieve real-time, reciprocal multimodal human-robot interaction (HRI), focusing on fusion strategies, system architectures, and applications in specific domains. We searched the Web of Science Core Collection for English-language empirical studies from 2015 to 2025, selecting 148 papers that address communicative integration. The most important insights include a clear shift toward hybrid and attention-based fusion since 2022 (about 43% of approaches), which better handles temporal asynchrony and embodied interactions than earlier methods; widespread inconsistency in evaluation metrics that hinders cross-study comparisons; and a persistent gap between strong laboratory performance and weaker real-world robustness under noise, occlusion, or user variability. Key challenges include real-time processing, semantic alignment, data efficiency, deployment robustness, and the under-explored integration of tactile signals with affect. Looking ahead, the review suggests prioritizing adaptive fusion policies, standardized benchmarks for synchronization and fluency, and continual learning to enable user-personalized adaptation. Ultimately, it aims to guide the development of more human-centered robotic agents that can engage collaboratively, meaningfully, and are socially acceptable in daily life.

Introduction

Loading...
$$\rightleftharpoonup{xx}$$ $$\longleftharp{xx}$$, $$\longrightharp{xx}$$,

Robots have evolved from industrial tools in structured settings to socially embedded agents in homes, hospitals, and classrooms. They now assist in education, therapy, and companionship, reflecting a shift from task-oriented automation to user-centered, adaptive interaction. At the core of this transition lies multimodal communication, enabling robots to perceive and respond to human cues—speech, gaze, gestures, facial expressions, and touch—with the temporal precision needed for natural exchanges in dynamic environments. This emphasis on coordination and responsiveness aligns with central concerns in human–computer interaction and cognitive science. In this review, seamless multimodal communication refers to interactions that feel fluid and intuitive to users, characterized by sub-second response times (typically less than 300 ms), minimal perceptible delays or mismatches, and robust adaptation to real-world variability. This distinguishes it from traditional command-based or unimodal systems, where timing errors, modality failures, or rigid responses disrupt natural flow.

Within human-robot interaction (HRI), this review focuses on human-robot communication, defined as the reciprocal, real-time exchange of multimodal signals between humans and robots. While HRI also encompasses planning, control, and safety, communication concerns the synchronization of perception and action across sensory channels. Early command-based systems prioritized efficiency over engagement, often producing rigid or delayed behavior. Contemporary work pursues adaptive, context-aware models that accommodate user state and emotion1,2. This evolution underscores the need for precise, intuitive, and personalized interaction frameworks.

Initial social and assistive robots commonly relied on scripted routines or single-modality inputs such as voice or touch, limiting flexibility3. Rule-based logic and slow processing frequently led to timing mismatches between gestures and speech, or to poor adaptation to user behavior. To address these issues, researchers employ low-latency fusion, multimodal sensory integration, and user-adaptive learning4,5. These strategies can be understood through three foundational communication principles widely recognized in HRI: complementarity (modalities compensate for each other), redundancy (overlap improves robustness), and synergy (convergence enhances naturalness).

Advances in sensing and learning have yielded socially responsive robots that integrate visual, auditory, tactile, and physiological inputs in real-time5,6. Rather than fixed scripts, these systems continuously adapt to user cues. Accessible programming interfaces and learning-from-demonstration methods allow educators and therapists to customize behaviors for caregiving and instruction, where precision, responsiveness, and adaptability are essential.

Communication patterns in social HRI can be grouped into parallel, complementary, and interactive modes (Figure 1). In parallel interaction, humans and robots act independently with minimal exchange. Complementary interaction introduces structured prompts or event-based support. The most advanced, interactive mode involves continuous bidirectional adjustment across channels such as speech, movement, and gaze7.

Literature selection and analysis

Systematic review methodology: We conducted a systematic literature search in the Web of Science Core Collection, targeting peer-reviewed journal articles and full conference proceedings from major publishers (IEEE, ACM, Springer, Elsevier, Wiley, Taylor & Francis). The Boolean query was: (("human-robot interaction" OR HRI) AND (multimodal OR "sensor fusion" OR gesture OR gaze OR speech OR tactile) AND ("real-time" OR "low-latency" OR adaptive OR robust)). Inclusion criteria were: English-language publications from January 2015 to 2025—a period capturing the shift from rule-based, single-modality systems to deep learning architectures emphasizing low-latency, adaptive, and robust interaction—with a focus on multimodal integration techniques for communicative HRI and empirical validation (quantitative or qualitative). We excluded studies limited to unimodal interaction, those addressing robot control without communicative intent, and non-empirical works (e.g., purely theoretical or simulation-only). The initial search returned a substantial corpus. After removing duplicates, we screened titles and abstracts against the inclusion criteria, then conducted full-text reviews to assess relevance and contribution to reliable multimodal fusion. This multi-stage process, summarized in the PRISMA-style flowchart (Figure 2), yielded a final corpus of 148 papers that form the empirical foundation of this review.

Overview of selected literature: Figure 3 shows publication trends and disciplinary distribution, with Computer Science and Automation & Robotics dominating, alongside growing contributions from Behavioral Sciences and Neuroscience, which underscores the field's interdisciplinary trajectory. We thematically clustered the literature by technical contribution to multimodal integration (Table 1). This synthesis covered modalities, fusion methods, latency handling, robustness strategies, adaptability, and metrics. Studies from 2019 to 2021 predominantly used early or late fusion, often emphasizing gestures and tactile inputs. Since 2022, hybrid and attention-based architectures have gained favor, particularly for affective and embodied interaction. In this corpus, hybrid fusion accounts for approximately 43% of approaches, early fusion 29%, and late fusion 28%—a distribution suggesting improved tolerance for temporal asynchrony.

Although real-time capability is commonly claimed in the broader literature screened initially, only a limited number of studies detail specific mitigation strategies (e.g., buffered processing, predictive models, or sensor-level optimization). Robustness typically relies on filtering, redundancy, or rule-based fallbacks, whereas anticipatory adaptation remains rare. Evaluation practices are inconsistent, and standardized assessments of temporal synchronization or interaction fluency remain infrequent. The metrics vary so widely across studies that direct comparisons are often difficult. Accuracy, F1 scores, and latency figures are common, but they stem from different datasets, tasks, and environmental conditions; few papers adopt shared benchmarks for temporal alignment or noise robustness. Consequently, strong performance in clean lab setups frequently overstates what methods can achieve outside controlled settings. Developing consistent evaluation frameworks—perhaps including standardized latency tests under realistic variability—would make it far easier to judge which approaches truly advance the field.

Key gaps remain in tactile-affective integration (only six studies8,9,10,11,12,13 address this) and dynamic latency adjustment under user-driven temporal uncertainty. Several interesting tensions emerge from the literature. Hybrid fusion is often praised for handling temporal misalignment well14, yet truly anticipatory or adaptive behaviors remain rare, calling into question just how seamless that claimed real-time performance really is. At the same time, impressive progress in vision and speech-based affect recognition stands in stark contrast to the minimal work integrating tactile cues with emotion understanding. High accuracy in controlled experiments also tends to degrade in real-world conditions, underscoring ongoing challenges in generalization across environments and users. Table 2 summarizes representative architectural trends and evaluation approaches, informing the challenges discussed in subsequent sections.

This review offers a distinct contribution to multimodal HRI literature. Recent surveys, such as Zhao et al.14, provide broad overviews of perception-driven decision-making, while Wang et al.15 focus on 5G-enabled visual-tactile transmission. Both emphasize general frameworks or communication protocols, with limited discussion of real-time fusion trade-offs, metric comparability, or deployment challenges. In contrast, we offer a more targeted analysis: comparing fusion strategies under specific constraints, critiquing cross-study metric inconsistencies, and highlighting practical gaps—particularly the lab-to-field performance divide that prior surveys have largely overlooked.

Access restricted. Please log in or start a trial to view this content.

Review and Perspective

Loading...
$$\rightleftharpoonup{xx}$$ $$\longleftharp{xx}$$, $$\longrightharp{xx}$$,

As robots become fixtures in homes, classrooms, and healthcare settings, natural and adaptive communication is critical to their success. Increasingly viewed not as tools but as social partners, they must perceive, interpret, and respond to human signals across modalities. Systems that accurately capture user intent and deliver timely feedback enhance task performance, trust, and emotional comfort—particularly when engaging children, older adults, or individuals with disabilities. Achieving this requires continuous, intuitive interaction that aligns robot responses with ongoing human behavior rather than interrupting it. Such fluency demands multimodal inputs—voice, gaze, gesture, facial expression—integrated through mechanisms that mirror human communication itself14,16.

Perception and communication channels
Vision remains the foundation of multimodal human–robot interaction. Visual cues such as facial expression, gaze, posture, and gesture underpin intent recognition, emotion inference, and engagement tracking17,18,19,20,21. Early handcrafted methods using Haar detection22 or background subtraction23 often failed under occlusion or lighting changes16,21. Deep learning systems now dominate, employing CNNs and recurrent models with multi-camera inputs. Figure 4 illustrates the HI-ROS tracking pipeline, which reduces motion error by 27% in real-time gesture monitoring.

Integration of gaze, arm, and head cues improves prediction of user actions17, while proprioceptive and depth fusion enhances precision24,25. Mutual gaze19 and biomimetic eye designs18 further demonstrate the role of subtle signals in trust and coordination. Figure 5 depicts the architecture from 26, where dual Leap Motion Controllers capture hand motion, fused through LSTM-RNN for robust surgical teleoperation. Recent multi-camera systems, such as Hi-ROS20 and RPF21,27, enhance person-following in unstructured settings. Figure 6 shows the adaptive ReID lifecycle that maintains tracking under occlusion by continuous classifier refinement. Complementary work on active labeling28 and semantic SLAM26 improves context awareness, extending perception from object recognition to behavioral understanding.

Language provides a parallel communicative channel. Early rule-based dialogue systems struggled with ambiguity, prompting the adoption of data-driven and transformer-based models29. Emotional and social responsiveness are now key: multimodal systems align facial and vocal signals for adaptive dialogue2,30. Reinforcement learning improves coordination between affective and acoustic cues5,31,32. Large language models such as GPT-333 and ChatGPT34 enable contextual reasoning and instruction following across tasks35,36,37. Their integration with vision and reward modeling38,39,40,41,42,43,44 supports socially aware, generalized interaction.

Auditory perception complements vision and language. Prosody, pitch, and rhythm communicate intent when visual cues are unreliable. Reinforcement-based feature extraction improves synchronization between speech and expression. Datasets such as AVA ActiveSpeaker45, AFFECT-HRI5, SAVEn-Vid46, HumorHRI47, and CHiME-5 enrich model robustness in noisy, affective environments. Pan et al.32 advance speech–lip synchronization for multi-speaker scenes, while multimodal fusion with tactile or physiological inputs informs adaptive behavior.

Embodied and physiological sensing
Tactile and physiological perception enhance social and physical awareness. Force, pressure, and biosignals enable inference of intent, stress, and engagement. Vision-based tactile simulators such as TACTO4 and neural mappers like FOTS48 generate high-resolution contact data at 300 fps, supporting low-latency control. FOTS models illumination and deformation via a learned reflectance function R:

Edge detection equation in image processing, formula: I(x,y)=R(∂H/∂x, ∂H/∂y).

and shadow projection:

Quaternion transformation matrix equation in static equilibrium context; mathematical diagram.

where Ms encodes projection geometry. High-resolution tactile skins have been realized using magnetorheological elastomers49, EtherCAT sensor arrays50, and auxetic Hall-effect materials51. Systems such as DiGeTac13 unify gesture, tactile, and distance signals through modular pipelines. Filtering of time-of-flight distance data is modeled as:

ARMA model equation: yt = a0xt + a1xt-1 + a2xt-2 + b1yt-1 + b2yt-2; time series analysis.

followed by neural gesture classification and CNN-based contact localization. Safety control is enforced by scaling the robot velocity with separation d. These modules enable robust human–robot collaboration. Multimodal transformers further integrate tactile, visual, and textual signals. The TVT-Transformer6 aligns modality-specific representations (qi, ki, vi) into joint matrices:

Static equilibrium equations: Q, K, V matrices; used in force balance calculations; mathematical concept.

enabling cross-modal self-attention for unified understanding (Figure 7).

Physiological sensing advances emotional alignment. Heinisch et al.5 synchronize physiological and audiovisual signals for affect modeling. Electromyography-based systems52,53 decode gestures and silent speech, while MetaSonic54 improves spatial perception through acoustic localization. Large-scale datasets such as AgiBot World Colosseo55 support multimodal training with over one million tactile-visual trajectories. Together, these advances mark a transition from unimodal perception to integrated understanding of social, physical, and internal states. Improved sensing fidelity enables robots not only to measure contact but to interpret hesitation, tension, or discomfort, responding with sensitivity and adaptive control, which are key foundations for the seamless multimodal frameworks discussed in later sections.

Multimodal fusion and representation learning
Fusion enables the transformation of distributed sensory streams into coherent, context-aware control. Approaches are typically grouped as early, late, or hybrid fusion. Early fusion integrates features at the input stage, supporting synchronized cues such as speech–gesture or facial–prosodic alignment. The Multimodal Transformer by Tsai et al.14 exemplifies attention-based early integration of visual, acoustic, and textual signals. Xue et al.56 extended this with residual coupling of facial landmarks and speech features, improving robustness under lighting variation, while Hu et al. Proposed MMSERNet57, a multiscale fusion network using dilated convolution and residual fusion to merge MFCC, spectrogram, and lexical features for emotion recognition. The fused output is expressed as:

Neural network equation: Yi calculation using ReLU, batch normalization, G2, and G4 transformations.

Late fusion, in contrast, combines decisions after independent processing, offering higher tolerance to asynchronous or noisy inputs. Examples include decision-level combination of gaze, inertial, and motion signals in underwater HRI58,59. Hybrid fusion integrates both approaches—mapping heterogeneous modalities to shared embedding spaces while preserving temporal specificity. The Multimodal Brain–Computer Fusion Network60 aligns EEG and visual features before classification, while cross-modal designs for 3D pose estimation merge inertial and monocular data61. Recent frameworks combine convolutional, recurrent, and transformer layers with attention alignment62,63, dynamically weighting modalities by context or task relevance64,65,66. The choice between early, late, and hybrid fusion is not a matter of one-size-fits-all. Early fusion tends to work well when modalities are tightly synchronized and noise levels are low, allowing the model to capture rich cross-modal relationships from the start. Late fusion, by contrast, is generally more robust in messy, asynchronous real-world conditions because each modality can be processed independently before decisions are combined. Hybrid approaches, which now appear in roughly 43% of recent studies, strike a useful middle ground, especially for affective and embodied interactions, by using attention to dynamically balanced inputs, though they do come with added computational demands. Ultimately, the best strategy depends on the specific task, noise profile, and hardware constraints, suggesting that future systems might benefit from switching fusion modes on the fly.

The complementarity between modalities also varies strongly with context. Vision paired with haptics is particularly effective in close-range physical tasks, such as grasping or object handover, where tactile feedback compensates for visual occlusion, lighting changes, or distance limitations and supplies precise force and contact information that audio cannot provide. Conversely, vision combined with audio performs better in distant or social settings, such as emotion recognition or dialogue in noisy environments, where prosody and tone carry intent and affect when visual cues are unavailable or unreliable.

Beyond the classic early, late, and hybrid framework, deep learning has enabled more granular categorizations. Attention-based fusion allows models to learn dynamic modality importance per input, effectively creating task- and context-specific integration paths without rigid stage boundaries. Learned fusion strategies go further by treating the entire fusion process as an end-to-end differentiable module, letting the network discover optimal combination patterns directly from data. These developments move the paradigm from predefined rules toward fully data-driven, adaptive architectures that better address the complexity of real-time multimodal HRI.

Vision-language grounding
The rapid growth of large language models has expanded the reach of multimodal reasoning, particularly through vision–language models (VLMs). These architectures align linguistic and visual information to enable robots to interpret commands, perceive scenes, and generate context-sensitive actions. Foundational frameworks such as LXMERT64 and CLIP66 established joint embeddings for image–text alignment. Flamingo65 introduced few-shot multimodal reasoning, while Maggio’s Clio39 dynamically links instructions to scene graphs via CLIP semantics. Gao et al.41 developed LAMARS for anticipatory planning based on multimodal prediction, and Xie et al.67 applied spatiotemporal convolutional models for efficient gesture comprehension. Arenas et al. introduced prompt-based customization for personalized interaction styles68. These systems allow commands like pick up the red cup to be grounded directly in scene semantics.

Integration with spatial reasoning further enhances robotic navigation and interpretation. Wilson’s LatentBKI38 grounds instructions in voxel-based spatial maps, while Nwankwo et al.40 fuse large language models with VLM reasoning to interpret dialogue acts under visual uncertainty. Event-based segmentation and asynchronous perception methods add efficiency to dynamic tasks. Collectively, VLMs act as bridges between symbolic instruction and embodied understanding, providing the cognitive substrate for flexible interaction.

Domain-specific applications
In healthcare and eldercare, multimodal communication enables safe, intuitive, and empathetic assistance. Core functions—fall detection, daily activity monitoring, emergency response—rely on posture, voice, and facial cues3. Hierarchical control models improve motion precision in arm-assistive robots1, while proprioceptive feedback guides cooperative dressing assistants69. Object handover, depicted in Figure 8, exemplifies synchronized perception and motor coordination70. Emotional responsiveness, studied through systems like LOVOT, fosters user trust but raises ethical concerns regarding overdependence.

Social robots require fluid multimodal responsiveness to sustain empathy and trust in domestic and public settings. Figure 9 outlines a typical architecture integrating sensing, planning, and socio-emotional reasoning layers. Large language models integrated with perception allow open-ended dialogue modulated by gaze and tone40. Pediatric studies show that expressive movement and prosody reduce anxiety in children71, while surveys highlight nonverbal behaviors as determinants of engagement8. Proactive models such as RoboAgent37 and Sub-Prior-guided Ant Colony Optimization72 enable joint exploration of shared goals. Systems capable of expressing internal states (ready, confused)73 enhance transparency and coordination, while adaptive social agents mediate group collaboration74,75.

Educational robots: In education, multimodal HRI promotes motivation and learning by leveraging embodiment and social signaling76. Combining gestures, facial, and verbal cues enhances accessibility and engagement6,9,77,78. Emotional adaptivity supports younger learners: systems that sense affective state and modulate speech and motion increase attention and vocabulary7,79. Integration with virtual reality improves immersion but challenges scalability80. Collectively, these systems shift the robot’s role from instructor to emotionally attuned collaborator.

Personalization ensures relevance and trust in multimodal systems71,81,82. Ethical personalization balances autonomy and empathy, tailoring responses through participatory feedback. Maaz et al.83 demonstrated affective tutoring guided by facial and cognitive cues, while Gomez-Izquierdo et al.84 modeled user preferences for synchronized behavior. Personalized adaptation can be formalized as:

Utility function equation: U(u,c)=Σwi·fi(u,c), mathematical formula, theoretical model.

where wi are learned preference weights. Implementations such as RFIS achieve real-time person re-identification and gesture recognition at the edge85, while proprioceptive touch interfaces10 enable intuitive physical communication. Social perception effects further modulate personalization outcomes, where narrator identity shifts user response timing and utterance length86. Personalization thus transforms users from passive recipients into co-designers, shaping interaction through mutual adaptation.

While much of the reviewed work remains in research prototypes, several multimodal HRI techniques have moved from research prototypes to commercial or industrial applications. SoftBank's Pepper robot, for instance, integrates speech, gesture, and facial expression recognition for customer service and education in retail and educational settings87. Toyota's Human Support Robot employs vision-gesture fusion for assistive tasks in homes and healthcare facilities88. These commercial implementations often simplify fusion strategies, typically using late or hybrid approaches with rule-based fallbacks to ensure reliability and low cost, underscoring the persistent gap between advanced academic models and scalable industrial solutions. Bridging this gap will require greater emphasis on edge deployment and rigorous real-world validation.

Challenges to seamless communication in HRI
Despite advances in multimodal sensing and fusion, sustaining seamless human–robot communication outside controlled settings remains difficult. Models that excel in labs often de-grade amid noise, delay, occlusion, shifting intent, or resource limits. The literature converges on six intertwined barriers: vision–language integration, real-time synchronization, multimodal robustness, semantic precision, dataset coverage, and deployment constraints.

Vision language integration
Visual understanding underpins grounded language and action. Foundational steps improved spatial consistency and abstraction through visual–inertial SLAM without manual initialization89, re-identification plus sensor fusion for occlusion-robust tracking16, lane graph prediction90, semi-supervised segmentation that aligns labels with navigation utility, and egocentric video with SLAM and large-scale segmentation for traversability26. Uncertainty-aware mapping remains central: uPLAM fuses probabilistic localization with panoptic segmentation for dynamic scenes91, while Clio adaptively builds CLIP-based hierarchical 3D scene graphs for task-sensitive semantics30. Grounding approaches progressed from structured descriptions and perception APIs92,93 to multimodal encoders and decoders94,95.

Instructions following in open environments integrate perception and reasoning. RobotGPT35 and GroundingGPT36 interpret commands via coupled visual–linguistic inference; SayPlan, LFG, and LLM-Planner align navigation and planning to grounded language34,44. To reduce hallucination and enforce semantic consistency, GLAM and Vtimellm introduce alignment objectives96,97, while adaptive frameworks validate LLM outputs against perception98,99. Timing critically shapes legibility: optimized gesture timing improves predictability9; reinforcement learning reduces audiovisual lag in emotion pipelines2; gesture entrainment increases perceived fluency100; and reviews consolidate LSTM, DTW, and GAN tools for alignment101. Overall, seamlessness demands a latency-aware loop across perception, fusion, and response, where mimicry30 and cerebellar-inspired prediction102 coordinate timing, as also illustrated by the system-level timing loops in Figure 10.

Real-time perception across modalities
Fluency depends on bounded end-to-end delay across heterogeneous streams. TACTO delivers high-resolution tactile feedback with low latency for responsive manipulation. Tendon-driven hands achieve fast interactive play through event vision and lightweight inference103, as shown in Figure 11. For vision–language–action, coupling CLIP with GraspNet accelerates grasping in clutter104. Transformers support rapid social haptic gesture classification11. Edge-optimized audiovisual models sustain sub-second inference24. Precise head-movement timing improves engagement105. Together, these results argue for synchronized fusion, buffer policies tailored by modality, and lightweight attention to preserve temporal co-adaptation. Acceptable latency varies significantly by task domain. In social dialogue or emotion recognition, delays under 200–300 ms maintain perceived fluency and natural interaction. Assistive tasks like object handover or fall detection demand tighter constraints (50–150 ms) to ensure safety and responsiveness. Teleoperation or collaborative manipulation often requires sub-100 ms latency to prevent operator disorientation or motion sickness, while navigation or monitoring applications can tolerate 300–500 ms without critical impact.

Temporal and semantic processing challenges
Misalignment and latency erode trust and task efficiency, especially in distributed teleoperator–robot–human systems106. Perceptual and interface strategies help users tolerate delay: bodily gestures and non-lexical fillers107, visual overlays and adaptive cues108,109,110,111. Control-level prediction shortens response gaps102. Parallel dialogue pipelines overlap filler generation, speech synthesis, and prompt editing to reduce perceived lag, as in112 and Figure 12. System studies decompose cumulative latency across capture, encode/decode, networks, decision, and actuation113. As shown in Figure 13, latency is introduced through sensor capture, video encoding and decoding, network transmission, operator processing, and vehicle actuation, forming a complex, bidirectional feedback loop. Syntalos provides millisecond synchronization for closed-loop experiments114. In affective exchanges, real-time facial mimicry enhances synchrony23. Delay-sensitive domains such as remote surgery show that anchoring haptics against delayed vision supports precision and responsiveness115. As shown in Figure 14, anchoring haptic feedback led to better performance and a stronger sense of responsiveness.

Fine-grained effects and intention recognition remain challenging. Microexpressions, prosody, gaze, and anticipatory motion often occur briefly and asynchronously. Emo produces anticipatory facial expressions aligned to user state116. Body language alone conveys intent where face or voice is unavailable117. Pepper-based multimodal displays combine verbal classification with expressive gestures and emoji for emotional synchrony118. Online RL strengthens fusion of facial and vocal streams2, while lightweight mimicry supports edge deployment30. Open issues include group effect, cultural variation, and temporal effect drift.

Generalization in open domains is also limited. Reinforcement learning from human feedback (RLHF) exhibits sparse rewards and brittleness under out-of-distribution inputs119. RoboAgent blends spatial reasoning with language grounding for cross-task transfer29. Reasoning through Action-free Data (RAD) leverages passive video and language for zero-shot capabilities without manual labels120. Perceiver-Actor grounds object properties to manipulate unfamiliar items121. Continual alignment between perception and semantics in RobotGPT and GroundingGPT improves adaptive reasoning but still faces real-time feedback challenges35,36.

Multimodal robustness and environmental variability
Robustness must span signal integrity and behavioral consistency. SimuMuHRI injects modality-specific noise and dropout to test resilience122. Physically grounded structured-light simulation improves RGB-D reliability under lighting and occlusion123,124. Latency–haptics coupling exposes sensitivity in remote manipulation125. Principles of redundancy and thermal resilience from safety-critical energy systems inform fault-tolerant HRI126,127.

Behavioral coherence is equally vital. Voice–appearance mismatch reduces trust128, and erratic but correct motion undermines cooperation129. Observer-based adaptive force control and robust interaction controllers sustain stability under structured and unstructured uncertainties130,131. Dataset coverage also constrains progress. Multiobot perception datasets—OPV2V132, CoPeD133, CSE134, and AgiBot World Colosseo55—expand collaborative sensing across simulation and field. AV-HRI resources—CHiME-5135 with the subsequent CHiME-8 challenge136, AVA-ActiveSpeaker45, AFFECT-HRI5, and SAVEn-Vid46—enable ASR, speaker activity, affect, and long-context instruction following. Large-scale social benchmarks—HumorHRI47, THÖR-MAGNI137, RW4T138, and InViG139—target humor, navigation, teaming, and interactive grounding. Despite breadth, many datasets are domain-specific or scripted, with limited temporal depth for evaluating seamlessness, as summarized in Table 3.

Evaluation of seamlessness
Conventional metrics such as accuracy and latency do not fully capture co-adaptation, mutual prediction, or social fluency. New evaluators rate contextual coherence and collaboration140, integrate user feedback and situational signals141, and add predictability, coordination, and comfort to physical HRI84. Team frameworks quantify shared anticipation74, while digital twins track trust calibration and team fluency142. A unified benchmark that fuses temporal synchronization, affective alignment, and user-centered fluency remains an open need.

Efforts toward standardization in multimodal HRI evaluation remain limited and fragmented. Notable initiatives such as the CHiME challenges135,136 have advanced benchmarks for audio-visual speech recognition, particularly in terms of noise robustness. Datasets like THÖR-MAGNI137 and RW4T138 provide shared resources for motion capture and teaming behaviors. However, no widely adopted unified protocol yet exists for cross-modal temporal synchronization, interaction fluency, or affective alignment across diverse modalities and domains. This absence of standardized metrics continues to hinder reproducibility and comparability, underscoring the need for community-driven benchmarks.

From lab to real-world deployment
Performance often degrades in unstructured settings due to sensory irregularities, shifting intent, and computing bottlenecks143. This performance degradation highlights a persistent divide between laboratory and real-world settings. In controlled environments, hybrid and attention-based fusion excel, delivering high accuracy and low latency for tasks like object manipulation or scene understanding. Real-world deployments, however, tell a different story. Assistive robots in eldercare, companion systems in homes, and educational tools in classrooms routinely contend with noise, occlusion, lighting changes, and unpredictable user behaviors that erode effectiveness. Late fusion approaches tend to prove more reliable under such variable, asynchronous conditions, while hybrid models, despite their strengths in affective and contextual adaptation, can be constrained by computational demands on edge devices. Many successful field implementations still incorporate simpler redundancy or rule-based safeguards for reliability, underscoring that moving from promising lab demonstrations to robust everyday performance remains a formidable challenge. Ecologically valid testing is essential to identify which techniques truly succeed beyond controlled experiments. Adaptive robust controllers maintain force accuracy amid disturbance131. LLM-driven collaborative planning updates goals as user intent evolves144. Visual overlays and latency-aware feedback help sustain operator awareness under degraded links107,109. Lightweight supervision with LoRa and digital shadows enables monitoring in low-bandwidth fields145. Experience from space robotics highlights that autonomy must co-evolve with human cognition under uncertainty and delay146.

Resource constraints shape feasibility. Sound-based affect and touch recognition run under 1 MB and 0.7 GFLOPs for low-power platforms12. Adding social features may improve presence, but risks overload when delays rise147. Scheduling and feature selection must balance cognitive demand with continuity148. Latency mitigation spans perception, control, and networks as organized in 113. Users frequently prefer transparent, stable interaction to maximal task efficiency, reinforcing that seamless HRI must prioritize technical capability alongside human comfort149,151. Across domains, real-world deployments show distinct preferences: social and educational applications often rely on hybrid vision-audio fusion for affective dialogue, while assistive and collaborative tasks favor vision-haptics combinations with late or early fusion for precise physical interaction. These patterns align with the summarized strategies—prioritizing late fusion for robustness in variable settings and hybrid/attention-based for richer contextual adaptation—though simplification for reliability remains common.

Access restricted. Please log in or start a trial to view this content.

Conclusions

Loading...
$$\rightleftharpoonup{xx}$$ $$\longleftharp{xx}$$, $$\longrightharp{xx}$$,

This review has several limitations. Restricting the search to English-language publications in the Web of Science Core Collection may have introduced language bias and excluded relevant non-English work. Focusing on peer-reviewed articles from 2015 to 2025 could also reflect publication bias, potentially overlooking preprints, grey literature, or emerging unpublished results. Manual screening of the 148 papers, though guided by explicit criteria, inevitably involves some subjectivity. Finally, the thematic interpretatio...

Access restricted. Please log in or start a trial to view this content.

Disclosures

Loading...
$$\rightleftharpoonup{xx}$$ $$\longleftharp{xx}$$, $$\longrightharp{xx}$$,

The authors declare that they have no conflicts of interest.

Acknowledgements

Loading...
$$\rightleftharpoonup{xx}$$ $$\longleftharp{xx}$$, $$\longrightharp{xx}$$,

This work was supported by Universiti Sains Malaysia, Bridging Grant with Project No: R501-LR-RND003-0000001342-0000.

Access restricted. Please log in or start a trial to view this content.

References

Loading...
$$\rightleftharpoonup{xx}$$ $$\longleftharp{xx}$$, $$\longrightharp{xx}$$,
  1. Frennert, S., et al. Towards a hierarchical user requirement structure for upper limb assistive robotics. ACM Trans Hum-Robot Interact. 14 (1), 1-26 (2024).
  2. Kansizoglou, I., Bampis, L., Gasteratos, A. An active learning paradigm for online audio- visual emotion recognition. IEEE Trans Affect Comput. 13 (2), 756-768 (2019).
  3. Keroglou, C., et al. A survey on technical challenges of assistive robotics for elder people in domestic environments: The aspida concept. IEEE Trans Med Robot Bionics. 5 (2), 196-205 (2023).
  4. Wang, S., Lambeta, M., Chou, P. W., Calandra, R. Tacto: A fast, flexible, and open-source simulator for high-resolution vision-based tactile sensors. IEEE Robot Autom Lett. 7 (2), 3930-3937 (2022).
  5. Heinisch, J., et al. Physiological data for affective computing in HRI with anthropomorphic service robots: The AFFECT-HRI data set. Sci Data. 11 (1), 333 (2024).
  6. Li, B., et al. TVT-Transformer: A tactile-visual-textual fusion network for object recognition. Inf Fusion. 118, 102943 (2025).
  7. Kim, H., et al. Technological applications of social robots to create healthy and comfortable smart home environment. Build Environ. , 112269 (2024).
  8. Leoste, J., Marmor, K., Heidmets, M. Nonverbal behavior of service robots in social interactions: A survey on recent studies. Interact Des Archit. 61, 164-192 (2024).
  9. Liu, N., et al. A lightweight network-based sign language robot with facial mirroring and speech system. Expert Syst Appl. 262, 125492 (2024).
  10. Iskandar, M., Albu-Schäffer, A., Dietrich, A. Intrinsic sense of touch for intuitive physical human-robot interaction. Sci Robot. 9 (93), eadn4008 (2024).
  11. Ren, Q., Hou, Y., Belpaeme, T. Low-latency classification of social haptic gestures using transformers. , (2023).
  12. Hou, Y., Ren, Q., Wang, W., Botteldooren, D. Sound-based recognition of touch gestures and emotions for enhanced human-robot interaction. , (2025).
  13. Al, G., Martinez-Hernandez, U. DiGeTac unit for multimodal communication in human-robot interaction. IEEE Sens Lett. 8 (5), 6001804 (2024).
  14. Tsai, Y. H., et al. Multimodal transformer for unaligned multimodal language sequences. , (2019).
  15. Wang, Z., Chen, M., Liu, Q. A review on multimodal communications for human-robot collaboration in 5G: From visual to tactile. Intell Robot. 5 (3), 579-606 (2025).
  16. Antonucci, A., et al. Humans as path-finders for mobile robots using teach-by-showing navigation. Auton Robots. 47 (8), 1255-1273 (2023).
  17. Duarte, N. F., Rakovic, Action anticipation: Reading the intentions of humans and robots. IEEE Robot Autom Lett. 3 (4), 4132-4139 (2018).
  18. Hayashi, K. Investigation of joint action in go/no-go tasks: Development of a human-like eye robot and verification of action space. Int J Soc Robot. , 1-14 (2024).
  19. Belkaid, M., et al. Mutual gaze with a robot affects human neural activity and delays decision-making processes. Sci Robot. 6 (58), eabc5044 (2021).
  20. Guidolin, M., Tagliapietra, L., Menegatti, E., Reggiani, M. Hi-ros: Open-source multi-camera sensor fusion for real-time people tracking. Comput Vis Image Underst. 232, 103694 (2023).
  21. Ovur, S. E., Demiris, Y. Naturalistic robot-to-human bimanual handover in complex environments through multi-sensor fusion. IEEE Trans Autom Sci Eng. 21 (3), 3730-3741 (2023).
  22. Viola, P., Jones, M. Rapid object detection using a boosted cascade of simple features. , (2001).
  23. Stauffer, C., Grimson, W. E. L. Adaptive background mixture models for real-time tracking. , (1999).
  24. Qi, W., et al. Multi-sensor guided hand gesture recognition for a teleoperated robot using a recurrent neural network. IEEE Robot Autom Lett. 6 (3), 6039-6045 (2021).
  25. Sun, J., Shen, Y., Rosen, J. Sensor reduction, estimation, and control of an upper-limb exoskeleton. IEEE Robot Autom Lett. 6 (2), 1012-1019 (2021).
  26. Kim, Y., et al. Learning semantic traversability with egocentric video and automated annotation strategy. IEEE Robot Autom Lett. 9 (3), 2662-2669 (2024).
  27. Ye, H., et al. Person re-identification for robot person following with online continual learning. IEEE Robot Autom Lett. 9 (11), 9151-9158 (2024).
  28. Vaswani, A., et al. Attention is all you need. , (2017).
  29. Schneider, S., Baevski, A., Collobert, R., Auli, M. Wav2vec: Unsupervised pre-training for speech recognition. , (2024).
  30. Liu, X., Chen, Y., Li, J., Cangelosi, A. Real-time robotic mirrored behavior of facial expressions and head motions based on lightweight networks. IEEE Internet Things J. 10 (2), 6039-6048 (2022).
  31. Tsai, C. Y., Su, Y. K. MobileNet-JDE: A lightweight multi-object tracking model for embedded systems. Multimed Tools Appl. 81 (7), 9915-9937 (2022).
  32. Pan, Z., Tao, R., Xu, C., Li, H. Selective listening by synchronizing speech with lips. IEEE/ACM Trans Audio Speech Lang Process. 30, 1650-1664 (2022).
  33. Brown, T. B., et al. Language models are few-shot learners. , (2020).
  34. Jin, Y., et al. RobotGPT: Robot manipulation learning from ChatGPT. IEEE Robot Autom Lett. 9 (3), 2543-2550 (2024).
  35. Li, Z., et al. GroundingGPT: Language Enhanced Multi-modal Grounding Model. , (2024).
  36. Bharadhwaj, H., et al. Roboagent: Generalization and efficiency in robot manipulation via semantic augmentations and action chunking. , (2024).
  37. Wilson, J., et al. LatentBKI: Open-dictionary continuous mapping in visual-language latent spaces with quantifiable uncertainty. IEEE Robot Autom Lett. , (2025).
  38. Maggio, D., et al. Clio: Real-time task-driven open-set 3D scene graphs. IEEE Robot Autom Lett. 9 (10), 5123-5130 (2024).
  39. Nwankwo, L., Rueckert, E. The Conversation is the Command: Interacting with Real-World Autonomous Robots Through Natural Language. , (2024).
  40. Gao, Y., et al. LAMARS: Large language model-based anticipation mechanism acceleration in real-time robotic systems. IEEE Access. 13, 3864-3880 (2024).
  41. Rana, K., et al. Sayplan: Grounding large language models using 3d scene graphs for scalable robot task planning. , (2023).
  42. Shah, D., et al. Navigation with large language models: Semantic guesswork as a heuristic for planning. , (2023).
  43. Song, C. H., et al. LLM-Planner: Few-shot grounded planning for embodied agents with large language models. , (2023).
  44. Roth, J., et al. AVA-ActiveSpeaker: An audio-visual dataset for active speaker detection. , (2020).
  45. Li, J., et al. SAVen-Vid: Synergistic audio-visual integration for enhanced understanding in long video context. arXiv. , (2024).
  46. Zhang, H., Hei, X., Cardenas, J., Miao, X., Tapus, A. Toward a multi-dimensional humor dataset for social robots. , (2024).
  47. Zhao, Y., Qian, K., Duan, B., Luo, S. FOTS: A fast optical tactile simulator for sim2real learning of tactile-motor robot manipulation skills. IEEE Robot Autom Lett. 9 (6), 2150-2157 (2024).
  48. Chen, D., et al. A high spatial resolution magnetorheological elastomer tactile sensor for texture recognition. IEEE Sens J. 25 (5), 1234-1241 (2025).
  49. Giovinazzo, F., et al. Design, implementation and testing of an EtherCAT-based network for multi-modal distributed sensing architectures. , (2024).
  50. Yun, Y., et al. Enhancing sensitivity across scales with highly sensitive Hall effect-based auxetic tactile sensors. Adv Intell Syst. 7 (3), 2300456 (2024).
  51. Zafar, M., Moosavi, S. K. R., Sanfilippo, F. Federated learning-enhanced edge deep learning model for EMG-based gesture recognition in real-time human-robot interaction. IEEE Sens J. 25 (5), 1234-1241 (2025).
  52. Dong, P., et al. Decoding silent speech cues from muscular biopotential signals for efficient human-robot collaborations. Adv Mater Technol. 10 (4), 2301123 (2024).
  53. Wang, J., An, Z., Guo, Y. MetaSonic: Advancing robot localization with directional embedded acoustic signals. IEEE Robot Autom Lett. 10 (2), 3150-3157 (2025).
  54. AgiBot-World-Contributors. AgiBot World Colosseo: A large-scale manipulation platform for scalable and intelligent embodied systems. , (2025).
  55. Xue, F., Li, Y., Liu, D., Xie, Y., Wu, L., et al. LipFormer: Learning to lipread unseen speakers based on visual-landmark transformers. IEEE Trans Circuits Syst Video Technol. 33 (9), 4507-4517 (2023).
  56. Hu, H., et al. Speech emotion recognition based on multimodal and multiscale feature f usion. Signal Image Video Process. 19 (2), 123 (2024).
  57. Lee, D., et al. Intuitive six-degree-of-freedom human interface device for human-robot interaction. IEEE Trans Instrum Meas. 73, 1-10 (2024).
  58. Kvasic, I., et al. Diver-robot communication dataset for underwater hand gesture recognition. Comput Netw. 245, 110392 (2024).
  59. Li, Z., et al. MBCFNet: A multimodal brain-computer fusion network for human intention recognition. Knowl-Based Syst. 296, 111826 (2024).
  60. Popescu, M., et al. Fusion of inertial sensor suit and monocular camera for 3D human pelvis pose estimation. , (2024).
  61. Zhou, S., Yang, B., Yuan, M., Jiang, R., Yan, R., et al. Enhancing SNN-based spatiotemporal learning: A benchmark dataset and cross-modal attention model. IEEE Trans Neural Netw Learn Syst. 35 (8), 11234-11248 (2024).
  62. Biswas, S., Saw, R., Nandy, A. Attention-enabled hybrid convolutional neural network for enhancing human-robot collaboration through hand gesture recognition. Comput Electr Eng. 123, 110020 (2025).
  63. Tan, H., Bansal, M. LXMERT: Learning cross-modality encoder representations from transformers. , (2019).
  64. Alayrac, J. B., et al. Flamingo: a visual language model for few-shot learning. , (2022).
  65. Radford, A., et al. Learning transferable visual models from natural language supervision. , (2021).
  66. Xie, J., Zhang, B., Lu, Q., Borisov, O. A dynamic head gesture recognition method for real- time intention inference and its application to visual human-robot interaction. Int J Control Autom Syst. 22, 252-264 (2024).
  67. Arenas, M., et al. How to prompt your robot: A promptbook for manipulation skills with code as policies. , (2024).
  68. Zhu, J., Gienger, M., Franzese, G., Kober, J. Do you need a hand?—A bimanual robotic dressing assistance scheme. IEEE Trans Robot. 40, 1906-1919 (2024).
  69. Ortenzi, V., et al. Object handovers: A review for robotics. IEEE Trans Robot. 37 (6), 2150-2168 (2022).
  70. Kabacińska, K., Teng, K., Robillard, J. Social robot interactions in a pediatric hospital setting: Perspectives of children, parents, and healthcare providers. Multimodal Technol Interact. 9 (2), 14 (2025).
  71. Viyuela, &. #. 2. 1. 1. ;., Sanfeliu, A. Human-robot collaborative minimum time search through sub- priors in ant colony optimization. IEEE Robot Autom Lett. 9 (11), 1123-1130 (2024).
  72. Roy, L., Croft, E., Kulić, D. Learning to communicate functional states with nonverbal expressions for improved human-robot collaboration. IEEE Robot Autom Lett. 9 (6), 7123-7130 (2024).
  73. San Martin, A., Kildal, J., Lazkano, E. An analysis of the role of different levels of exchange of explicit information in human-robot cooperation. Front Robot AI. 12, 1511619 (2025).
  74. Gillet, S., et al. Interaction-shaping robotics: Robots that influence interactions between other agents. ACM Trans Hum-Robot Interact. 13 (1), 25 (2024).
  75. Belpaeme, T., et al. Social robots for education: A review. Sci Robot. 3 (21), eaat5954 (2018).
  76. Zhang, R., et al. Predicting emotion reactions for human-computer conversation: A variational approach. IEEE Trans Hum-Mach Syst. 51 (4), 712-721 (2022).
  77. Huang, C., Fu, L., Hung, S. C., Yang, S. Effect of visual programming instruction on students' flow experience, programming self-efficacy, and sustained willingness to learn. J Comput Assist Learn. 41 (1), 123-135 (2025).
  78. Berrezueta-Guzman, S., Dolón-Poza, M. Enhancing preschool language acquisition through robotic assistants: An evaluation of effectiveness, engagement, and acceptance. IEEE Access. , (2025).
  79. Lei, Y., Su, Z., Cheng, C. Virtual reality in human-robot interaction: Challenges and benefits. Electron Res Arch. 31 (4), 2374-2408 (2023).
  80. Pareto, J., Coeckelbergh, M. Social assistive robotics: An ethical and political inquiry through the lens of freedom. Int J Soc Robot. 16 (8), 1797-1808 (2024).
  81. Hung, L., et al. “It’s always happy to see me”: Exploring LOVOT robots as companions for older adults. J Rehabil Assist Technol Eng. 12, 20556683251320669 (2025).
  82. Maaz, N., Mounsef, J., Maalouf, N. CARE: Towards customized assistive robot-based education. Front Robot AI. 12, 1474741 (2025).
  83. Gómez-Izquierdo, G., Laplaza, J., Sanfeliu, A., Garrell, A. Enhancing context-aware human motion prediction for efficient robot handovers. , (2025).
  84. Lee, S., Lee, S., Park, H. Integration of tracking, re-identification, and gesture recognition for facilitating human-robot interaction. Sensors. 24 (15), 4850 (2024).
  85. Bakhoda, I., et al. Exploring the impact of narrator type on response latency and utterance length during interactive storytelling. , (2024).
  86. Pandey, A. K., Gelin, R. A mass-produced sociable humanoid robot: Pepper: The first machine of its kind. IEEE Robot Autom Mag. 25 (3), 40-48 (2018).
  87. Yamamoto, T., et al. Toyota Human Support Robot (HSR): Development and future prospects. 2019 IEEE/SICE International Symposium on System Integration (SII). , 634-639 (2019).
  88. Lupton, T., Sukkarieh, S. Visual-inertial-aided navigation for high-dynamic motion in built environments without initial conditions. IEEE Trans Robot. 28 (1), 61-76 (2011).
  89. Zürn, J., Posner, I., Burgard, W. Autograph: Predicting lane graphs from traffic observations. IEEE Robot Autom Lett. 9 (1), 73-80 (2024).
  90. Sirohi, K., Büscher, D., Burgard, W. UPLAM: Robust panoptic localization and mapping leveraging perception uncertainties. IEEE Robot Autom Lett. 9 (8), 5123-5130 (2024).
  91. Wang, L., et al. A survey on large language model based autonomous agents. Front Comput Sci. 18 (6), 186345 (2024).
  92. Zang, Y., et al. Contextual object detection with multimodal large language models. Int J Comput Vis. 133 (2), 1-19 (2024).
  93. Ivanovic, B., et al. Multimodal deep generative models for trajectory prediction: A conditional variational autoencoder approach. IEEE Robot Autom Lett. 6 (2), 295-302 (2020).
  94. Floridi, L., Chiriatti, M. GPT-3: Its nature, scope, limits, and consequences. Minds Mach. 30, 681-694 (2020).
  95. Carta, T., et al. Grounding large language models in interactive environments with online reinforcement learning. , (2023).
  96. Huang, B., et al. VTimeLLM: Empower LLM to grasp video moments. , (2024).
  97. Gupta, M., et al. From ChatGPT to ThreatGPT: Impact of generative AI in cybersecurity and privacy. IEEE Access. 11, 80218-80245 (2023).
  98. Li, X., et al. Chain-of-knowledge: Grounding large language models via dynamic knowledge adapting over heterogeneous sources. , (2024).
  99. Kimoto, M., Iio, T., Shiomi, M., Shimohara, K. Coordinating entrainment phenomena: Robot conversation strategy for object recognition. Appl Sci. 11 (5), 2358 (2021).
  100. Nyatsanga, S., et al. A comprehensive review of data-driven co-speech gesture generation. Comput Graph Forum. 42 (2), 569-596 (2023).
  101. Abadía, I., et al. A cerebellar-based solution to the nondeterministic time delay problem in robotic control. Sci Robot. 6 (58), eabf2756 (2021).
  102. Deng, X., Weirich, S., Katzschmann, R. K., Delbruck, T. A rapid and robust tendon-driven robotic hand for human-robot interactions playing rock-paper-scissors. , (2024).
  103. Xu, K., et al. A joint modeling of vision-language-action for target-oriented grasping in clutter. , (2023).
  104. Lee, H., Hahn, S. Effect of robot head movement and its timing on human-robot interaction. Int J Soc Robot. 17 (1), 3-14 (2024).
  105. Kim, S., Hernandez, I., Nussbaum, M., Lim, S. Teleoperator-robot-human interaction in manufacturing: Perspectives from industry, robot manufacturers, and researchers. IISE Trans Occup Ergon Hum Factors. 12 (1-2), 28-40 (2024).
  106. Pelikan, H., Hofstetter, E. Managing delays in human-robot interaction. ACM Trans Comput-Hum Interact. 30 (4), 1-24 (2022).
  107. Kang, D., Nam, C., Kwak, S. Robot feedback design for response delay. Int J Soc Robot. 16 (2), 1341-1361 (2024).
  108. Akita, E., Zaidner, G., Pryor, M. Improved situational awareness and performance with dynamic task-based overlays for teleoperation. , (2024).
  109. Scholz, C., et al. Improving robot-to-human communication using flexible display technology as a robotic-skin-interface: A co-design study. Int J Intell Robot Appl. 9 (1), 146-163 (2024).
  110. Reardon, C., et al. Augmented reality visualization of autonomous mobile robot change detection in uninstrumented environments. ACM Trans Hum-Robot Interact. 13 (3), 1-30 (2024).
  111. Asaka, S., Itoyama, K., Nakadai, K. Improving impressions of response delay in AI-based spoken dialogue systems. , (2024).
  112. Kamtam, S., et al. Network latency in teleoperation of connected and autonomous vehicles: A review of trends, challenges, and mitigation strategies. Sensors. 24 (12), 3957 (2024).
  113. Klumpp, M., et al. Syntalos: A software for precise synchronization of simultaneous multi- modal data acquisition and closed-loop interventions. Nat Commun. 16 (1), 708 (2025).
  114. Du, J., et al. Sensory manipulation as a countermeasure to robot teleoperation delays: System and evidence. Sci Rep. 14, (2024).
  115. Hu, Y., et al. Human-robot facial coexpression. Sci Robot. 9 (88), eadi4724 (2024).
  116. Gao, W., Shen, S., Ji, Y., Tian, Y. Human perception of the emotional expressions of humanoid robot body movements: Evidence from survey and eye-tracking measurements. Biomimetics. 9 (11), 684 (2024).
  117. Cárdenas, P., et al. Evaluation of robot emotion expressions for human-robot interaction. Int J Soc Robot. 16 (9), 2019-2041 (2024).
  118. Casper, S., et al. Open problems and fundamental limitations of reinforcement learning from human feedback. , (2025).
  119. Clark, J., Mirchandani, S., Sadigh, D., Belkhale, S. Action-free reasoning for policy generalization. arXiv. , (2025).
  120. Shridhar, M., et al. Perceiver-Actor: A multi-task transformer for robotic manipulation. , (2022).
  121. Wang, X. Mobile robot environment perception system based on multimodal sensor fusion. Appl Comput Eng. 127, 42-49 (2025).
  122. Bai, K., Zhang, L., Chen, Z., Wan, F., Zhang, J. Close the sim2real gap via physically-based structured light synthetic data simulation. , (2024).
  123. Fatehi, K., Torres, M., Kucukyilmaz, A. An overview of high-resource automatic speech recognition methods and their empirical evaluation in low-resource environments. Speech Commun. 167, 103151 (2024).
  124. Louca, J., Eder, K., Vrublevskis, J., Tzemanaki, A. Impact of haptic feedback in high latency teleoperation for space applications. ACM Trans Hum-Robot Interact. 13 (2), 1-21 (2024).
  125. Luo, W., Zhang, S., Gao, Y., Shen, C. Review of mechanisms and detection methods of internal short circuits in lithium-ion batteries. Ionics. 31 (5), 3945-3964 (2025).
  126. Luo, W., et al. Study on the comprehensive multisource thermal runaway failure characteristics and cooling effects of extinguishing agents in 280 Ah batteries. Phys Fluids. 37 (3), (2025).
  127. Alimardani, M., Roode, R., Vaitonyte, J., Louwerse, M. Effect of a virtual agent's appearance and voice on uncanny valley and trust in human-agent collaboration. , (2024).
  128. Moradinezhad, R., Solovey, E. Investigating trust in interaction with inconsistent embodied virtual agents. Int J Soc Robot. 13 (8), 2103-2118 (2021).
  129. Zixuan, H., et al. Observer-based adaptive robust force control of a robotic manipulator integrated with external force/torque sensor. Actuators. 14 (3), 116 (2025).
  130. Huang, J., et al. Adaptive robust interaction force control of a robotic manipulator in uncertain environments. IEEE Trans Ind Electron. 72 (8), 8251-8260 (2025).
  131. Xu, R., et al. OPV2V: An open benchmark dataset and fusion pipeline for perception with vehicle-to-vehicle communication. , (2022).
  132. Zhou, Y., Quang, L., Nieto-Granda, C., Loianno, G. CoPeD: Advancing multi-robot collaborative perception—A comprehensive dataset in real-world environments. IEEE Robot Autom Lett. 9 (7), 6416-6423 (2024).
  133. Park, H., et al. A benchmark dataset for collaborative SLAM in service environments. IEEE Robot Autom Lett. 9 (12), 11337-11344 (2024).
  134. Barker, J., Watanabe, S., Vincent, E., Trmal, J. The fifth CHiME speech separation and recognition challenge: Dataset, task and baselines. , (2018).
  135. Cornell, S., et al. The CHiME-8 DASR challenge for generalizable and array agnostic distant automatic speech recognition and diarization. , (2024).
  136. Schreiter, T., et al. THÖR-MAGNI: A large-scale indoor motion capture recording of human movement and robot interaction. Int J Robot Res. 44 (4), 568-591 (2024).
  137. Savko, L., Qian, Z., Gremillion, G., Neubauer, C., Canady, J., et al. RW4T dataset: Data of human-robot behavior and cognitive states in simulated disaster response tasks. , (2024).
  138. Zatsarynna, O., Farha, Y., Gall, J. Multi-modal temporal convolutional network for anticipating actions in egocentric videos. , (2021).
  139. Gkournelos, C., Konstantinou, C., Makris, S. An LLM-based approach for enabling seamless human-robot collaboration in assembly. CIRP Ann. 73 (1), 9-12 (2024).
  140. Rahimi, H., et al. User-VLM 360: Personalized vision language models with user-aware tuning for social human-robot interactions. , (2025).
  141. Langås, E., Zafar, M., Sanfilippo, F. Exploring the synergy of human-robot teaming, digital twins, and machine learning in industry 5.0: A step towards sustainable manufacturing. J Intell Manuf. , 1-24 (2025).
  142. Nizamuddin, M., Kamisetty, A., Gummadi, J. C., Talla, R. R. Integrating neural networks with robotics: Towards smarter autonomous systems and human-robot interaction. Robotics Xplore: USA Tech Digest. 1 (1), 157-169 (2024).
  143. Asuzu, K., Singh, H., Idrissi, M. Human-robot interaction through joint robot planning with large language models. Intell Serv Robot. 18 (2), 261-277 (2025).
  144. Shamshiri, R., et al. Internet of robotic things with a local LoRa network for teleoperation of an agricultural mobile robot using a digital shadow. Discov Appl Sci. 6, 414 (2024).
  145. Smith, C., Mott, T., Williams, T. Perspectives on level of autonomy decisions in space robotics. , (2024).
  146. Kim, T., Song, Y., Kim, D., Song, H. Too much is as bad as too little: The impact of implementing multiple social interaction features on trust and acceptance of automated vehicle agents. Int J Hum-Comput Interact. 41 (21), 13875-13890 (2025).
  147. Elsayyad, S., et al. An effective robot selection and recharge scheduling approach for improving robotic networks performance. Sci Rep. 14 (1), 28439 (2024).
  148. Mayer, L., et al. Human-AI collaboration: Trade-offs between performance and preferences. Cogn Res Princ Implic. 11, 18 (2026).
  149. Yang, B., et al. Gaze and environmental context-guided deep neural network and sequential decision fusion for grasp intention recognition. IEEE Trans Neural Syst Rehabil Eng. 31, 3687-3689 (2023).
  150. Rückin, J., Magistri, F., Stachniss, C., Popović, M. Semi-supervised active learning for semantic segmentation in unknown environments using informative path planning. IEEE Robot Autom Lett. 9 (3), 2662-2669 (2024).

Access restricted. Please log in or start a trial to view this content.

Reprints and Permissions

Request permission to reuse the text or figures of this JoVE article

Request Permission

Tags

Fusion StrategiesSystem ArchitecturesReal Time ProcessingSemantic AlignmentAdaptive FusionContinual LearningSocially Assistive Robots

Related Articles