$$\rightleftharpoonup{xx}$$
$$\longleftharp{xx}$$,
$$\longrightharp{xx}$$,
Achieving effective and natural human-AI collaboration remains a major challenge, particularly in dynamic real-world environments. The existing frameworks are often limited by the rigid single-modal inputs or rule-based multimodal fusion, restricting their robustness, personalization, and adaptability. This study introduces an adaptive multimodal interaction framework that integrates vision-based pose estimation and speech-based command understanding within a hierarchical human intention forecasting model. The framework employs Transformer-based sequence models for action prediction and Dynamic Time Warping to align human behavior with the structured task graphs for long-term planning. Unlike prior approaches that rely on static modality switching, our architecture applies an attention-based fusion mechanism to dynamically weight the most informative inputs. Furthermore, an online adaptation module enables real-time adjustment to user behavior, complemented by a lightweight explainability layer to foster user trust. The experimental evaluation in a collaborative assembly task demonstrates significant gains in task performance (up to 24%), intention prediction accuracy (up to 18%), and user satisfaction. These results highlight the effectiveness and versatility of the adaptive multimodal interaction frameworks, supporting their potential to advance collaborative intelligence toward more reliable, transparent, and human-oriented AI systems.