Research Article

Designing Adaptive Multimodal Interaction Frameworks to Enhance Human-AI Collaboration

DOI:

10.3791/69830

February 24th, 2026

In This Article

Summary

Loading...
$$\rightleftharpoonup{xx}$$ $$\longleftharp{xx}$$, $$\longrightharp{xx}$$,

This work proposes a Dynamic Time Warping (DTW)-based hierarchical alignment mechanism with online adaptation and attention-driven multimodal fusion, enabling reliable long-horizon intention prediction and task tracking. A lightweight explainability module improves interpretability, fostering trust in human-AI collaboration. The approach generalizes to both robotic and virtual interactive systems.

Abstract

Loading...
$$\rightleftharpoonup{xx}$$ $$\longleftharp{xx}$$, $$\longrightharp{xx}$$,

Achieving effective and natural human-AI collaboration remains a major challenge, particularly in dynamic real-world environments. The existing frameworks are often limited by the rigid single-modal inputs or rule-based multimodal fusion, restricting their robustness, personalization, and adaptability. This study introduces an adaptive multimodal interaction framework that integrates vision-based pose estimation and speech-based command understanding within a hierarchical human intention forecasting model. The framework employs Transformer-based sequence models for action prediction and Dynamic Time Warping to align human behavior with the structured task graphs for long-term planning. Unlike prior approaches that rely on static modality switching, our architecture applies an attention-based fusion mechanism to dynamically weight the most informative inputs. Furthermore, an online adaptation module enables real-time adjustment to user behavior, complemented by a lightweight explainability layer to foster user trust. The experimental evaluation in a collaborative assembly task demonstrates significant gains in task performance (up to 24%), intention prediction accuracy (up to 18%), and user satisfaction. These results highlight the effectiveness and versatility of the adaptive multimodal interaction frameworks, supporting their potential to advance collaborative intelligence toward more reliable, transparent, and human-oriented AI systems.

Introduction

Loading...
$$\rightleftharpoonup{xx}$$ $$\longleftharp{xx}$$, $$\longrightharp{xx}$$,

Human-robot collaboration (HRC) has been one of the important areas of research, especially since robots are increasingly being deployed in human-centric environments. While frameworks and technologies for short-term HRC have improved significantly1,2, long-term HRC is not yet as matured3. In industrial applications, exploitation of long-term HRC can lead to a prominent rise in productivity and adaptability in complex manufacturing processes4. Domestic settings, robots as companions or elderly care agents have the potential to enhance quality of life through long-term HRC5. Medical positions community machines as a replacement for household and caregivers, covering care to the affected role with dementia, autism, and additional physical or mental illnesses6. At the same time, the travel and friendliness domains utilize industrial robots to advance user knowledge7. In intelligent engineering, the purpose is to participate AI in a human-centric way to recover competencies and make personalized crops, a stage known as Industry 5.0. Human-AI and human-robot collaboration systems have been explored across diverse application domains8,9,10; however, recent research has increasingly focused on collaborative robotic settings that require reliable intention prediction and long-horizon task understanding.

AI is already deeply embedded in everyday life, and its self-governing or semiautonomous applications are increasingly manifest within many areas of business. This simply states that well-working business functions currently require systems capable of adapting to changes in technology and responsiveness to organizational challenges11. Thereby, the necessity to improve human action by hybridizing, in the process of AI adoption and implementation, helps facilitate overcoming some barriers. Consequently, hybrid systems combine artificial and human intelligence to achieve complex tasks together, deliver improved results, and advance their skills by learning from colleagues. Therefore, extended concepts that include hybrid intelligent systems12, collaboration, and amplified intelligence are informally shaping research and understanding of boosting collaboration by integrating diverse AI modalities.

The term "Trust" was derived from interpersonal relationship studies, defining an emotional bond in human relationships. For example, the author13 identified that the level of trust technically determines aspects of feelings, supposed appeal, and ethics of cooperating parties. Sociologically, interpersonal trust variations have profound impacts on separate well-being, communal interaction, socio-economic growth, and practical behavior14,15,16. HRI, in essence, is an attempt to bridge procedural conversations among social workers and clever automata beyond the conventional automaton linguistic hurdles17. Such interaction combines the computational strengths of autonomous systems with human intuition, enabling effective collaboration across diverse application settings. The use of automata brings with it real advantages: significant decreases in financial expenses, reduced ecological footprint, and considerably less human risk10. But in the midst of achievements, human confidence in robots is a pillar, particularly when assessing the general quality of the interaction18,19.

The level of human trust accorded to robots is a reflection of their importance and central role in many areas. In health, robot trust influences the distribution of medical resources according to HRI as well as the efficiency of transmissible illness suppression20,21,22. In soldierly applications, the level of social belief in automata controls the level of operative competence of soldierly investigation device strategies and the level of accuracy of fighting transportations23,24. Additionally, belief in machines affects submission and use of AI active belief standardization mechanisms in speculative soldierly policymaking25. Even within environmental investigation, e.g., high-level bodily movement research, the rationality of results depends on the degree of belief in the HRI26. A significant warning, though, is that rising decisional difficulties have the potential to degrade this trust, jeopardizing interaction27. Table 1 describes the comparative overview of several studies.

Paper (Year, Source)Main ObjectivesKey Methods & ModalitiesInsights
Xie & Park (2021, arXiv) [28]Customize robot interactions with emotive social robotsPolicy for Reinforcement learning based on both verbal and nonverbal clues (body position, speech, and facial expressions)An RL-trained robot's behavior changes in response to human emotions, increasing simplicity and engagement.
Li et al. (2022, IEEE T‑IE) [29]Facilitate proactive support for manufacturing assemblyIntention prediction, transfer learning, and multimodal learning (visual + skeletal)Improved robot proactivity compared to baselines; quicker, more precise action prediction during assembly
Abdelrahman et al. (2022, IEEE Access) [30]Estimating involvement in real time in group HRIRule-based classification combined with deep learning and vision-based gaze/head posture≥90% F score; manages several users at once in physical configurations
Chiossi et al. (2025, arXiv) [31]Structure for multimodal, dynamic intent transmission in manufacturing environmentsTheoretical multifaceted layout: SAT transparency layers; haptic, aural, and visual modesEstablishes the conceptual framework and explains the medium, planning, and intent of the design environment.
McGrath et al. (2024, arXiv) [32]Model the development of confidence in human-AI associations that are adaptable.Team phases plus trust antecedents; a dynamic trust model that isn't dependent on particular modalitiesProvides trust-aware design principles for all stages of interaction and time.
Fragiadakis et al. (2024, arXiv) [33]Standardize the assessment of cooperation between humans and AI.Metrics selected by a decision tree according to the HAIC mode; mixed-method metricsThe framework aids in selecting pertinent metrics for AI/human-centric, symbiotic interaction modalities.
Younis et al. (2023, MDPI) [34]HRI that adapts to context using demographic inferenceAge and gender are estimated using vision and audio models, and robot behavior is adjusted appropriately.The survey summarizes approaches and emphasizes the significance of customized adaption in HRI.

Table 1: A comparative overview of recent research in Human-AI Collaboration that emphasizes the key goals, methodologies, and findings of each study. Datasets used in the study, their specifications, size, modalities, and participant distribution.

Though the current hierarchical and multimodal human-robot collaboration (HRC) framework shows impressive progress in using visual and audio modalities for long-term collaboration, a number of research gaps still exist, which hinders its applicability to more general Human-AI collaboration settings. For one, the present framework is mostly applicable to physical robotics (e.g., KINOVA GEN3 arm) and non-transferable to non-embodied or virtual AI agents that exist in varied digital environments28. Second, while multimodality is addressed by vision and speech, the system implements a pre-defined modality-switching strategy instead of dynamic, context-dependent fusion according to user state or task difficulty29. Third, the approach emphasizes task efficiency and resilience but does not specifically model user trust, cognitive burden, or affective state, which are essential for long-term collaboration and system explainability. In addition, human interpretability and confidence in the system's adaptive behavior are limited when explainability processes are absent from decision-making. Last but not least, the assessment is restricted to a single domain (toy car assembly) without multi-domain validation, representing variability in the real world30. Such gaps necessitate the development of a generalized, adaptive, and human-centered multimodal interaction framework that combines heterogeneous input channels dynamically and models human intention, trust, and context in real time to augment Human-AI collaboration.

The foremost goal of this study is to maximize human-AI collaboration through the design of an adaptive multimodal interaction system that unifies vision and speech modalities with real-time user modeling. Building on the solid foundation of robust long-term human-robot collaboration, this work will generalize the multimodal architecture to general AI systems outside robotics, such as virtual agents and intelligent interfaces. The major emphasis is on dynamic intention prediction via hierarchical planning mechanisms, which integrate low-level action understanding and high-level task progress tracking. In contrast to rule-based systems, the framework presented herein will utilize adaptive learning methods, including user-specific model updates and context-aware modality fusion, to dynamically adapt interaction behavior according to user preferences, cognitive status, and environmental dynamics. Moreover, the framework also aims at enhanced transparency and trust with explainable interaction reasoning, which allows users to understand AI choices and give feedback. Finally, this research hopes to enhance interaction fluency, task effectiveness, and long-term user satisfaction in different collaborative contexts.

This paper introduces a new Adaptive Multimodal Interaction Framework intended to promote human-AI collaboration by dynamically fusing visual and audio modalities. The 'unimodal' system refers to a model that uses only a single input modality (e.g., visual or speech) for task execution and intent prediction, without integrating information across multiple channels. The 'non-adaptive' system refers to a rule-based or fixed-strategy system that does not adjust its behaviour based on user actions or context, such as the Rule-Based Switching baseline. These baselines provide reference points to evaluate the advantages of multimodal perception and adaptive interaction strategies implemented in our proposed Transformer with attention-guided fusion model. Based on previous hierarchical HRC paradigms31, our method brings several important innovations:

This work presents a Dynamic Time Warping (DTW)-based alignment mechanism that maps expected low-level human actions to nodes in a hierarchical task graph, in contrast to previous multimodal interaction systems that mainly concentrate on short-term action recognition32. In the face of temporal unpredictability and noisy multimodal inputs, this allows for reliable long-horizon intention prediction and task progress tracking.

An online adaptation module that continuously updates intention inference and action prediction models during interaction is incorporated into the suggested system. This overcomes the drawbacks of static or offline-trained multimodal models by enabling the system to dynamically adapt to unique user behaviours, preferences, and contextual changes.

Visual and auditory modalities are dynamically weighted based on their contextual importance at each time step using a unique attention-based fusion technique. This data-driven fusion approach improves robustness in a variety of interaction settings and substitutes rule-based modality switching.

This study incorporates a lightweight explainability module that converts hierarchical planning and fusion decisions into comprehensible justifications in order to close the interpretability gap in adaptive human-AI systems. This enhances the quality of long-term collaboration and encourages proper trust calibration.

The suggested approach is intended to generalize beyond embodied robots to virtual agents and intelligent interfaces, expanding its usefulness across many human-AI cooperation domains, even though it was proven in a real-world collaborative assembly activity utilizing a robotic platform.

Access restricted. Please log in or start a trial to view this content.

Protocol

Loading...
$$\rightleftharpoonup{xx}$$ $$\longleftharp{xx}$$, $$\longrightharp{xx}$$,

The proposed architecture for augmenting human-AI cooperation uses an adaptive multimodal interaction pipeline combining vision and speech modalities in a dynamic, hierarchical reasoning framework. This study used publicly available data from previously published experiments. No new human subjects were involved, and no additional ethical approval was required.

Proposed adaptive multimodal interaction architecture

Fundamentally, the system begins with perception modules, such as visual stream captures user pose and motion using RGB-D sensors, and an audio stream analyzes speech inputs using a light-weight ASR (Automatic Speech Recognition) engine. Such raw inputs are fed into the Transformer-based action prediction module, which encodes temporal dependencies in pose and command sequences to reason about low-level human actions. Subsequently, to enable high-level collaboration planning, a Dynamic Time Warping (DTW) mechanism aligns the detected actions with nodes of a pre-defined task graph to predict user goals and next actions. A new attention-based fusion module decides, at every timestep, which modality (audio or vision) is giving the most contextually important information and adjusts the input weights dynamically. To individualize interaction, the system includes online model adaptation, adapting action predictors in real-time to changing user behavior. A light-weight explainability module also creates user-friendly summaries of predicted intentions and AI decisions to improve transparency and trust. The whole system is tested in a simulated but real-world cooperative assembly environment, allowing comparison across unimodal and adaptive-multimodal environments. Such an approach facilitates more adaptable, more intuitive, and more reliable collaboration with AI agents in all areas.

Figure 1 shows the architecture of the described Adaptive Multimodal Interaction Framework aimed at improving Human-AI collaboration. It starts with the Input Acquisition phase, which is based on two main sensing devices: an RGB-D camera for observing depth-enhanced visual information (e.g., human posture and movement) and a directional microphone for collecting speech commands. Raw signals are subsequently run through the Pre-Processing module, which processes the audio-visual data for subsequent analysis by cleaning, normalizing, and formatting the input streams. Then, the pipeline proceeds to the Feature Extraction phase, where the data is split into two streams: The Visual Stream, which isolates pose and motion features through pose estimation algorithms, and the Audio Stream, which transcribes speech into interpretable command embeddings. A multimodal Transformer encoder, which uses multi-head self-attention and temporal positional encoding to express long-range interdependence across interaction sequences, then processes the fused representation. The user's current goal state is estimated by a long-horizon intention and task progress predictor, and low-level actions are mapped to nodes in a hierarchical task graph for reliable task tracking by a DTW-based alignment module. Fusion weights and prediction parameters are continuously updated via an online model adaptation loop in response to interaction feedback, and an explainability module provides clear explanations to promote openness and confidence. The whole pipeline allows the system to adaptively and efficiently work together with users on dynamic tasks and environments.

Multimodal neural network diagram, RGB-D camera, visual and audio encoders, fusion, adaptive output.
Figure 1: An overview of the suggested multimodal interaction framework based on transformers. An attention-guided technique is employed to encode and dynamically integrate visual and aural data. A Transformer module, which integrates hierarchical task modeling and online adaptation, processes the fused representation for long-horizon intention prediction and task progress estimation.. Please click here to view a larger version of this figure.

The suggested system incorporates a lightweight explainability module that works in tandem with the multimodal fusion and hierarchical planning components to improve transparency and user confidence. The module derives interpretable cues from the DTW-based task graph alignment (e.g., current task progress and anticipated next action) and the attention-based fusion layer (e.g., modality importance weights). These cues are converted into comprehensible explanations for humans, such as whether spoken commands or visual gestures primarily affected a system choice and how the anticipated action fits into the broader work plan. Users receive explanations in the form of succinct written or visual feedback, which helps them comprehend, predict, and, if needed, rectify system behavior. Improved user happiness, interaction smoothness, and trust-related metrics indicated in the evaluation are indirect indicators of user perception of explainability. The findings imply that in long-term human-AI interaction, transparent reasoning for AI decisions enhances user confidence, facilitates collaboration, and increases acceptance of adaptive system behavior.

Input data

Input acquisition in the envisioned framework for Human-AI collaboration is carried out through a structured multimodal sensing mechanism that combines both vision and hearing modalities to effectively realize human actions and intentions. Visual input is acquired with an RGB-D camera, which captures real-time pose information of humans, spatial location, and context of the surroundings. Visual data is processed through pretrained models such as BlazePose to get upper-body keypoints of interest for interaction tasks. At the same time, auditory information is gained from a directional microphone and interpreted through a pretrained DeepSpeech model to identify speech commands that disambiguate or provide supplementary uncertain visual cues. In cooperative settings, the dual-input system offers dynamic adaptation and context-aware perception. The training and evaluation dataset comprises more than 5,000 multimodal human motion trajectories1 for a range of interaction activities, such as assembly and handover behaviors, gathered from four users in controlled and semi-structured settings. To enhance generalization, the system enables online model adaptation in deployment, which allows the model to update its predictions in real time as a function of changing user behavior and environmental dynamics.

Dataset description

The training and evaluation dataset for the suggested framework was obtained from real-world human-robot collaboration experiments under a toy car assembly task using the KINOVA GEN3 robot1. The experimental environment aimed to mimic long-term collaborative activities encompassing multimodal interaction, which included both visual and auditory modalities. In particular, the training dataset is comprised of 3,003 sequences of trajectories from two users, and the test dataset holds 2,088 sequences, including data from two new users for assessing model generalization. Every trajectory records a 5-step sequence of upper-body pose keypoints retrieved with BlazePose, along with corresponding speech commands transcribed by DeepSpeech33. The collected data depicts natural human variation in gesture and command expression on activities including object delivery, tool use, and cooperative actions. The multimodal dataset is the foundation for supervised learning of the DLinear action and trajectory forecasting models, and is the central benchmark in the user studies performed under vision-only, audio-only, and multimodal conditions. Incorporating different types of users and realistic noise conditions guarantees robustness and adaptability testing for permanent HRC systems. Table 2 shows the detailed dataset description.

AttributesDetails
Dataset NameToy Car Assembly Dataset (collected in-house)
Collection EnvironmentReal-world HRC experiment with KINOVA GEN3 arm
Number of Users (Training)2 users
Number of Users (Testing)4 users (2 from training set, 2 unseen)
Total Trajectories (Training)3,003 trajectories
Total Trajectories (Testing)2,088 trajectories
Modalities CapturedVisual (RGB-D for pose), Audio (Speech commands)
Data FormatPose keypoints (from BlazePose), speech-to-text commands
Temporal ResolutionEach trajectory includes 5-time steps
Tasks Covered3 main tasks: Pick-and-place, object rotation, collaborative assembly
UsageSupervised training of action/trajectory prediction models; user study

Table 2: Dataset Description. Information on the simulation environment, including programming framework, hardware settings, and evaluation environment.

Dataset pre-processing

To enable adaptive multimodal interaction in human-AI collaboration, the raw dataset, such as visual (RGB and depth images), auditory (speech commands), and temporal behavior data (pose keypoints), is subjected to structured preprocessing. Visual data are first extracted by a pose estimation model fpose, commonly derived from BlazePose or OpenPose. For frame t, the human pose vector is defined as:

Pose estimation formula: Z_t(v) = f_pose(I_t^RGB, I_t^Depth) ∈ ℝ^k×2, mathematical equation.   (1)

Where k is the number of keypoints and every keypoint contains 2D coordinates. The sequences are normalized by torso length and smoothed over time with a Savitzky-Golay filter:

Savitzky-Golay filter equation, method for data smoothing in signal processing, mathematical formula.   (2)

where w is the size of the filtering window. Second, raw audio streams At are sliced into segments using a Voice Activity Detector (VAD). Credible segments are fed into a speech recognition model fspeech, (e.g., DeepSpeech) to pull out command tokens:

Mathematical speech recognition equation, featuring c_t as output from f_speech(A_t) process.   (3)

where V is the vocabulary of recognized commands. For integration, command tokens for all are time-aligned to visual pose timestamps using an interpolation-based alignment function τ:

Equation illustrating concept of aligned temporal data in computational analysis; formula for alignment.   (4)

Third, to form the temporal intention training sequences Mathematical equation representing indexed sequence Z_T as a set in vector notation., we define trajectory sequences , along with associated speech commands Mathematics equation formula showing aligned constant set notation, used in data analysis methods.. These are windowed into fixed-length segments utilizing sliding windows of size l:

Equation of statistical alignment method, Chi-squared formula, diagram for data comparison.   (5)

Equation of action/intention labels, formula, educational use.   (6)

Lastly, the inputs are normalized to zero mean and unit variance per keypoint dimension for convergence in trajectory-based prediction models:

Standardization process equation, Z-score formula, statistical data analysis method.   (7)

where μj, σj are the mean and standard deviation of keypoint j completed the training set.

Feature extraction

In the outlined adaptive multimodal interaction paradigm, strong and real-time cooperation is enabled by combining visual, audio, and contextual task features. The pipeline of feature extraction starts with concurrent data acquisition from two main modalities: visual signals Thermodynamic equation, OtVt, describing static gas laws, formula representation for calculations. and audio signals Chemical equilibrium; \(O_t^A\) formula; symbolic representation; key concept in chemistry., indexed at time step t. These observations are passed through corresponding perception modules to obtain high-level semantic features pertaining to human behavior, intention, and speech command.

Extracting visual features

The raw visual data Thermodynamic equation, OtVt, describing static gas laws, formula representation for calculations. from RGB-D cameras is fed through a human pose estimator (e.g., BlazePose), providing a spatial keypoint vector Mathematical formula ZtH ∈ ℝk×3; equation analysis; linear algebra application., where k represents the detected number of keypoints and each contains 3D coordinates:

Equation for pose estimation; formula ztH=fpose(OtV); mathematical concept.   (8)

These keypoints are time-embedded into a pose trajectory Time series equation Tt^H={Zt−L^H,…,Zt^H}, statistical analysis formula, from which motion dynamics are extracted via trend-seasonal decomposition:

Time series decomposition equation, diagram showing trend, seasonality, error components.   (9)

where ∈_t represents the residual noise.

Extracting audio features

Audio input Chemical equilibrium; \(O_t^A\) formula; symbolic representation; key concept in chemistry. is processed through a pre-trained speech recognition model (e.g., DeepSpeech) and produces a tokenized command representation Mathematical equation showing vector st^H in real number space R^n, formula representation., where n is the token length:

Speech signal processing equation \(s_t^H = f_{\text{speech}}(O_t^A)\), formula representation.   (10)

Multimodal fusion

To construct a unified interaction context, we combine the features by a confidence-weighted attention based on confidence and mutual information:

Mathematical formula for weighted combination in prediction analysis, includes alpha parameters.   (11)

where ϕV and ϕA are the projection functions for visual and auditory features into a shared latent space, and αt∈[0,1] is a dynamic modality weighting term calculated on the basis of entropy:

Equation of selective attention; ratio of predicted and actual outcomes; mathematical expression.   (12)

Here, Mutual Information Equation \(I(g; O_t^V)\); formula for information theory analysis. and Information theory equation I(g; O<sup>A</sup><sub>t</sub>); mutual information concept. are mutual information between the goal state g and observed modalities.

Planning using Contextual Embedding

The last multimodal feature vector Ft is enriched with task context and progress Pand becomes the state input to the adaptive planning module:

Concatenation formula \(x_t = concat(F_t, P_t)\) in mathematical equation representation.   (13)

This extracted feature xt is used as input to the downstream intention prediction and motion planning models, supporting adaptive and robust human-AI collaboration in dynamic and uncertain settings.

Action and plan prediction via Task Graph Alignment

Following the multimodal feature extraction process, where speech intent (audio), pose trajectory (visual), and contextual (e.g., task logs or emotion) data are integrated in the subsequent processing includes adaptive human action prediction and plan inference using hierarchical task graph alignment.

Let the integrated multimodal feature vector at time step t be represented as:

Multimodal fusion formula, Ft=[z_vision, s_audio, c_context], equation, data processing concept.   (14)

The system initially forecasts the low-level human action via a function fact, which is often represented as a recurrent encoder-decoder or transformer:

Machine learning prediction equation, formula; dynamic process forecasting in data science research.   (15)

where F(t-k:t) denotes the feature fusion sequence in the last k time steps, and Mathematical formula, a with hat symbol, element of set A, used in set theory analysis. is the forecasted action from the action set A. The module is online fine-tuned by adaptive loss:

Loss function equation, L_act=λ_cls·CE(â_t,a_t)+λ_traj·||t̂_(t+1)−t_(t+1)||²_2, in machine learning.   (16)

Here, CE is the action classification cross-entropy loss, then the second term nudges good future trajectory prediction.

For high-level plan prediction, we frame the task as a progression along a Hierarchical Task Graph G = (N, E), wherein every node n∈N signifies a sub-task or intermediate goal. The objective is to match the observed action sequence against a reference task path R*∈G  by reducing the distance:

Optimal rotation equation diagram, R*=argmin(d(Pt,R)); mathematical optimization concept.   (17)

Where Pt denotes the current progress trace and d(⋅) denotes the Dynamic Time Warping (DTW) distance or soft alignment metric.

The ultimate plan-level intent Statistical notation π̂_t equation; used in probability analysis; symbol only. is then derived by projecting the predicted action a ̂_t onto the most probable next step in the aligned path Chromatography system diagram, method R*, component separation, experiment setup, data analysis.:

Equation: estimator symbol π̂ equals function R of control variable â; data analysis concept.   (18)

This hierarchical decoding guarantees that even with unclear or noisy sensory input, the system picks the most contextually relevant high-level intent consistent with the current task sequence. This predictive planning not only enhances collaboration efficiency but also enhances anticipation and adaptability within multi-turn Human-AI interaction.

Multimodal fusion (attention-based) mechanism

In adaptive human-AI collaboration, multimodal integration of several sensory modalities (e.g., visual: human pose, trajectory; auditory: speech commands) is necessary for proper interpretation of user intention. A rule-based or static fusion approach cannot cope with real-world dynamic human interactions. For this, a fusion mechanism based on attention is presented here to dynamically weight each modality according to its contextual importance at every time step.

Let us denote the extracted feature vectors from the visual and audio modalities as Equation depicting vector space element: \( h_v \in \mathbb{R}^d \). and Mathematical symbol, vector space equation, ha in Rd, educational math formula., respectively. These vectors encode semantic information obtained through earlier processing pipelines, e.g., visual features may result from pose estimation and action prediction, while audio features may be derived from speech embeddings using models like DeepSpeech or wav2vec.

The objective of attention-based fusion is to calculate modality-specific attention weights that capture the relative significance of every modality. This is done using a soft attention mechanism.

Attention-weight computation

In a first step, both vectors are transformed into a shared attention space through a learnable transformation. That is, for every modality i∈{V,A}, the intermediate attention score is calculated as:

Attention mechanism equation, e_i=wTtanh(Wh_i+b), formula for neural network calculation.   (19)

where Matrix notation \( W \in \mathbb{R}^{k \cdot d}, w \in \mathbb{R}^{k} \), linear algebra concept. and Equation showing vector space notation: "b ∈ ℝᵏ". are attention network learnable parameters, and tanh is a non-linear activation function. This conversion allows the model to extract high-level correlations and contextual saliency of each modality.

These unnormalized attention scores eV and αA are then normalized with a softmax function such that they add up to 1

Softmax activation function formula, equation, neural network, machine learning, aᵢ = exp(eᵢ)/[exp(eᵥ)+exp(eₐ)].   (20)

Here, αV and αA are the attention weights over the visual and audio modalities, respectively. The weights are adaptive and then they are updated over time depending on the existing task context, signal quality, and user activity.

Equation: Fusion calculation using weighted sum formula.   (21)

The resulting fused representation Mathematical formula h_fused ∈ ℝ^d, representing vectors in d-dimensional real space. is obtained as a weighted combination of the input modality vectors:

This combined vector encodes the common semantics of the visual and auditory information in a manner that actively highlights the most informative stream. It is then fed into downstream modules like the task graph matching, intention prediction, or robot motion planning engine.

Algorithm 1 : Fusion Based on Multimodal Attention

Input: Visual feature vector is hV∈Rd, Vector of audio features is hA∈Rd

Learnable parameters: W∈Rk*d,w∈Rk and b∈Rk

Output: fused representation hfused∈Rd.

Step 1: Determine intermediate representations by computing using Equation 19

Step 2: Use Softmax to normalize attention scores using Equation 20

Step 3: Calculate the fused vector using Equation 21

Step 4: Return hfused

The Attention-Based Multimodal Fusion algorithm provides a means for smart fusion of features of various input modalities, such as visual and audio streams, together in a context-aware fashion. Fusion is essential in systems wherein each modality presents complementary but varying information, such as in human-AI collaborative environments where vision captures gestures or body position and audio captures speech commands or voice signals.

The Algorithm 1 begins by determining intermediate representations for each modality. To calculate both of these models, eV for the visual flow and eA for the audio stream, a shared neural network layer first performs a non-linear transformation on the corresponding feature vectors. Then, the attention weights αV and αA are calculated by performing a softmax function over the intermediate scores. This helps the weights to be positive and add up to 1, essentially making the contributions from each modality to be normalized. The more informative a modality is, as modelled, the greater its attention weight.

Ultimately, the algorithm calculates the fused depiction hfused through a weighted combination of the graphic and audio feature vectors, normalized with their respective attention weights. This fused vector combines both modalities in a data-driven, context-sensitive fashion and serves for downstream tasks such as intention recognition, decision-making, or robot motion planning. In total, this algorithm enables the system to dynamically focus on the most pertinent modality or modality combination at a specific moment, instead of using fixed or rule-based fusion approaches. This renders the framework stronger, more flexible, and responsive in dynamic human-AI interaction situations.

Figure 2 illustrates the workflow of Attention-Based Multimodal Fusion mechanism that operates in the context-sensitive integration of visual and auditory inputs. Visual Feature Extraction, the first step in the workflow, involves extracting spatial or pose-related features from input images or video (such as RGB-D cameras). These are subsequently input into a Visual Attention module that learns to highlight the most informative content of the visual information. In parallel, Audio Attention modules process speech embeddings or audio tokens from spoken commands, attributing contextual significance to different auditory signals. The resulting attention-weighted features are integrated through an attention-based fusion layer and passed to a position-wise feedforward network with residual connections and layer normalization, forming a standard Transformer encoder block. The output is a unified multimodal embedding that supports robust long-horizon intention prediction and adaptive decision-making under varying interaction conditions. This design enables the system to selectively concentrate on the more robust modality at any point, improving robustness and interpretability in real-time human-AI interaction.

Multimodal learning flowchart showing visual/audio feature extraction and attention mechanisms in AI.
Figure 2: The attention-guided multimodal fusion module's detailed construction. In order to generate a context-aware fused representation at each time step, modality-specific encoders collect features from visual and aural streams. These features are then dynamically weighted using a data-driven attention mechanism. Please click here to view a larger version of this figure.

Online model adaptation

To ensure high-quality human-AI collaboration under different user behaviors and contexts, our framework incorporates an online model adaptation mechanism that progressively optimizes user-specific models during interaction. The module tracks real-time behavioral signals (e.g., pose, gesture trajectories, speech tone) and updates the underlying predictive models with incremental learning algorithms. In particular, trajectory prediction and intent recognition models are fine-tuned with recent input via gradient-based fine-tuning or low-rank adaptation (LoRA) so that the system is able to adapt to changing user behavior or task-specific requirements without having to be fully retrained.

The feedback layer acts harmoniously together with adaptation through both implicit feedback (e.g., corrections, hesitations, delays) and explicit commands (e.g., speech or GUI). It has two purposes: Adaptively calibrating the salience of multimodal cues, and Transmitting trust/confidence measures to the decision and control modules.

The feedback cycle enables smoother task performance and user-oriented responses. To nurture trust and transparency, the framework integrates an explainability module that translates low-level model decisions into human-understandable justifications. It uses rationale maps from multimodal fusion transformer, supplemented with logic tracing through the task graph. It can optionally present this rationale via a GUI or auditory aide to enable users to comprehend and perhaps even correct system behavior.

Access restricted. Please log in or start a trial to view this content.

Results

Loading...
$$\rightleftharpoonup{xx}$$ $$\longleftharp{xx}$$, $$\longrightharp{xx}$$,

The adaptive multimodal interaction model proposed here addresses these concerns and greatly enhances human-AI collaboration by seamlessly combining vision and speech modalities with real-time user adaptation processes. In contrast to the baseline hierarchical multimodal system, our approach includes a personalized feedback loop as well as cognitive load-aware adaptation that provides more intuitive, context-aware interaction. Upon experimental testing in simulated collaborative activities, such as object handling, compo...

Access restricted. Please log in or start a trial to view this content.

Discussion

Loading...
$$\rightleftharpoonup{xx}$$ $$\longleftharp{xx}$$, $$\longrightharp{xx}$$,

The findings conclusively demonstrate the efficiency of the suggested adaptive multimodal interaction basis to enhance human-AI collaboration.

Apart from the standard assessment scores like Task Completion Time (TCT), Human Intention Prediction Accuracy (HIPA), Multimodal Fusion Efficiency (MFE), Error Recovery Rate (ERR), User Satisfaction Score (USS), and Interaction Smoothness Index (ISI), additional user-focused metrics can be added for in-depth analysis of adaptive multimodal frameworks. ...

Access restricted. Please log in or start a trial to view this content.

Disclosures

Loading...
$$\rightleftharpoonup{xx}$$ $$\longleftharp{xx}$$, $$\longrightharp{xx}$$,

The authors have no conflicts of interest.

Acknowledgements

Loading...
$$\rightleftharpoonup{xx}$$ $$\longleftharp{xx}$$, $$\longrightharp{xx}$$,

The authors gratefully acknowledge the support provided by the School of Design, Jiangnan University, Wuxi, China.

Access restricted. Please log in or start a trial to view this content.

Materials

List of materials used in this article
NameCompanyCatalog NumberComments
RGB-D Camera (Intel RealSense D435)Intel CorporationD435Depth resolution 1280×720 px @ 30 fps; RGB resolution 1920×1080 px @ 30 fps; used for capturing user pose, body movements, and spatial context
Directional MicrophoneGeneric / Multiple vendorsN/AWideband, noise-canceling microphone (16 kHz–44.1 kHz); used for collecting spoken commands
BlazePose (pretrained model)Googlehttps://developers.google.com/mediapipe/solutions/vision/posePose estimation model with 33 keypoints; real-time human pose extraction
OpenPose (pretrained model)CMU Perceptual Computing Labhttps://github.com/CMU-Perceptual-Computing-Lab/openposeMulti-person pose estimation framework used for extracting motion trajectories
DeepSpeech ASR EngineMozillav0.9.3Pretrained automatic speech recognition model for speech-to-text transcription
wav2vec 2.0 ASR EngineMeta AIhttps://github.com/facebookresearch/fairseqSelf-supervised speech representation model for real-time speech processing
Transformer-based Action Prediction Model (DLinear / Seq2Seq)Custom implementationN/ATemporal encoder-decoder architecture with attention for short-term action forecasting
Dynamic Time Warping (DTW) ModuleCustom / SciPy-basedhttps://docs.scipy.org/doc/scipy/reference/generated/scipy.spatial.distance.cdist.htmlSoft alignment algorithm for matching user actions with predefined task graphs
Attention-based Multimodal Fusion NetworkCustom implementationN/AConfidence-weighted attention mechanism for fusing vision and speech inputs
Online Adaptation Module (LoRA / Gradient-based)Custom implementationhttps://github.com/microsoft/LoRALightweight, parameter-efficient fine-tuning for adapting to user behavior
Explainability Module (Rule-based + Attention Maps)Custom implementationN/AProvides interpretable decision rationale using rules and attention visualizations
KINOVA Gen3 Robotic ArmKINOVA RoboticsGen37-DOF robotic manipulator with integrated torque sensors; used for collaborative toy car assembly
Cooperative Assembly Simulation EnvironmentCustom environmentN/AReal-world toy car assembly setup used for multimodal collaboration benchmarking
Evaluation Metrics (TCT, HIPA, MFE, Latency, ERR, USS, ISI)Custom definitionsN/AQuantitative and user-centric metrics for evaluating collaboration efficiency, accuracy, and trust

References

Loading...
$$\rightleftharpoonup{xx}$$ $$\longleftharp{xx}$$, $$\longrightharp{xx}$$,
  1. Yu, P., Abuduweili, A., Liu, R., Liu, C. Robustifying long-term human-robot collaboration through a hierarchical and multimodal framework. arXiv. , (2024).
  2. Cheng, Y., Sun, L., Liu, C., Tomizuka, M. Towards efficient human-robot collaboration with robust plan recognition and trajectory prediction. IEEE Robot Autom Lett. 5, 2602-2609 (2020).
  3. Abuduweili, A., Li, S., Liu, C. Adaptable human intention and trajectory prediction for human-robot collaboration. arXiv. , (2019).
  4. Amirova, A., Rakhymbayeva, N., Yadollahi, E., Sandygulova, A., Johal, W. Ten years of human-NAO interaction research: A scoping review. Front Robot AI. 8, 744526(2021).
  5. Liu, R., Liu, C. Human motion prediction using adaptable recurrent neural networks and inverse kinematics. IEEE Control Syst Lett. 5, 1651-1656 (2021).
  6. Wang, T., Zhao, Y., Chen, L., Li, Q., Xu, Y., Zhang, H. Multimodal human-robot interaction for human-centric smart manufacturing: A survey. Adv Intell Syst. 6, 2300359(2024).
  7. Ergonomic adaptation of robotic movements in human-robot collaboration. van den Broek, M. K., Moeslund, T. B. Companion Proc. ACM/IEEE Int Conf Hum-Robot Interact, 2020, 499-501 (2020).
  8. Human-robot collaboration using multimodal perception. Yang, Y., et al. Proc IEEE Int Conf Robot Biomimetics (ROBIO), , 358-363 (2021).
  9. Sawadwuthikul, G., et al. Intelligent manufacturing with deep reinforcement learning. IEEE Trans Ind Inform. 18, 1883-1894 (2022).
  10. Multimodal robotic manipulation in collaborative tasks. Huang, W., Gu, J., Duan, P., Hou, S., Zheng, Y. Proc IEEE Int Conf Robot Autom (ICRA), , 14213-14219 (2021).
  11. Dellermann, D., et al. The future of human-AI collaboration: A taxonomy of design knowledge for hybrid intelligence systems. arXiv. , (2021).
  12. Ostheimer, J., Chowdhury, S., Iqbal, S. An alliance of humans and machines for machine learning: Hybrid intelligent systems and their design principles. Technol. Soc. 66, 101647(2021).
  13. Afsar, B., Al-Ghazali, B. M., Cheema, S., Javed, F. Cultural intelligence and innovative work behavior: The role of work engagement and interpersonal trust. Eur J Innov Manag. 24, 1082-1109 (2020).
  14. Spadaro, G., Gangl, K., Van Prooijen, J. W., Van Lange, P. A. M., Mosso, C. O. Enhancing feelings of security: How institutional trust promotes interpersonal trust. PLoS ONE. 15, e0237934(2020).
  15. Pavez, I., Gómez, H., Laulié, L., González, V. A. Project team resilience: The effect of group potency and interpersonal trust. Int J Proj Manag. 39, 697-708 (2021).
  16. Yuan, H., Long, Q., Huang, G., Huang, L., Luo, S. Different roles of interpersonal trust and institutional trust in COVID-19 pandemic control. Soc Sci Med. 293, 114677(2022).
  17. Yun, Y., Ma, D., Yang, M. Human-computer interaction-based decision support systems with applications in data mining. Future Gener Comput Syst. 114, 285-289 (2021).
  18. Nazar, M., Alam, M. M., Yafi, E., Su'ud, M. M. A systematic review of human-computer interaction and explainable artificial intelligence in healthcare. IEEE Access. 9, 153316-153348 (2021).
  19. Ho, Y. H., Tsai, Y. J. Open collaborative platforms for multi-drone search and rescue operations. Drones. 6, 132(2022).
  20. VAHAK: A blockchain-based outdoor delivery scheme using UAVs for healthcare 4.0 services. Gupta, R., et al. Proc IEEE INFOCOM Workshops, , 255-260 (2020).
  21. BloCoV6: A blockchain-based 6G-assisted UAV contact tracing scheme for COVID-19. Zuhair, M., Patel, F., Navapara, D., Bhattacharya, P., Saraswat, D. Proc 2nd Int Conf. Intell Eng Manag (ICIEM), , 271-276 (2021).
  22. Sahoo, S. K., et al. Intelligent trust-based utility and reusability model using UAV-assisted sensor networks. Appl Sci. 12, 1317(2022).
  23. Ko, Y., et al. Drone secure communication protocol for future sensitive military applications. Sensors. 21, 2057(2021).
  24. Gupta, R., Kumari, A., Tanwar, S., Kumar, N. Blockchain-envisioned softwarized multi-swarming UAVs to tackle COVID-19 situations. IEEE Netw. 35, 160-167 (2020).
  25. Tomsett, R., et al. Rapid trust calibration through interpretable and uncertainty-aware. Patterns. 1, 100049(2020).
  26. Huang, R., Zhou, H., Liu, T., Sheng, H. Multi-UAV collaboration to survey Tibetan antelopes in Hoh Xil. Drones. 6, 196(2022).
  27. Seeber, I., et al. Machines as teammates: A research agenda on AI in team collaboration. Inf Manag. 57, 103174(2020).
  28. Ajoudani, A., et al. Progress and prospects of the human-robot collaboration. Auton Robots. 42, 957-975 (2018).
  29. Oviatt, S., et al. The Handbook of Multimodal-Multisensor Interfaces, Volume 1: Foundations, User Modeling, and Common Modality Combinations. , Association for Computing Machinery. New York, NY, USA. (2017).
  30. Villani, V., Pini, F., Leali, F., Secchi, C. Survey on human-robot collaboration in industrial settings. Mechatronics. 55, 248-266 (2018).
  31. Nikolaidis, S., et al. Human-robot collaboration in manufacturing: Quantitative evaluation of predictable, convergent robot behavior. Auton Robots. 41, 123-140 (2017).
  32. Roggen, D., Junker, M., Transter, G., Roques, D. L. Collecting complex activity datasets in smart environments. J Ambient Intell Humaniz Comput. 1, 1-17 (2009).
  33. Bazarevsky, V., et al. BlazePose: On-device real-time body pose tracking. arXiv. , (2020).
  34. Garcia, S., Gomez-Donoso, F., Cazorla, M. Enhancing human-robot interaction: Development of a multimodal robotic assistant for user emotion recognition. Appl Sci. 14, 11914(2024).
  35. Li, S., et al. Toward proactive human-robot collaborative assembly: A multimodal transfer-learning-enabled action prediction approach. IEEE Trans Ind Electron. 69, 8579-8588 (2021).
  36. Abdelrahman, A. A., Almotiri, J., Yaseen, Q., Alanazi, A., Alotaibi, B. Multimodal engagement prediction in multiperson human-robot interaction. IEEE Access. 10, 61980-61991 (2022).
  37. Chiossi, F., et al. Designing intent: A multimodal framework for human-robot cooperation in industrial workspaces. arXiv. , (2025).
  38. McGrath, M. J., Chen, Y., Patel, A., Smith, J., Rao, V. Collaborative human-AI trust (CHAI-T): A process framework for active management of trust in human-AI collaboration. arXiv. , (2024).
  39. Fragiadakis, G., Nikas, C., Karageorgiou, E., Papadopoulos, A. Evaluating human-AI collaboration: A review and methodological framework. arXiv. , (2024).
  40. Younis, H. A., Khan, S., Raza, M., Ahmed, F., Mahmood, T. Multimodal age and gender estimation for adaptive human-robot interaction: A systematic literature review. Processes. 11, 1488(2023).

Access restricted. Please log in or start a trial to view this content.

Reprints and Permissions

Request permission to reuse the text or figures of this JoVE article

Request Permission

Tags

Human AI CollaborationMultimodal InteractionAdaptive FrameworksPose EstimationSpeech Command UnderstandingIntention ForecastingTransformer ModelsAttention FusionDynamic Time WarpingExplainable AI

Related Articles