$$\rightleftharpoonup{xx}$$
$$\longleftharp{xx}$$,
$$\longrightharp{xx}$$,
Recent advances in artificial intelligence (AI) and somatosensory technology are reshaping music education by enabling learners to interact with music through body movements, where gestures are translated into notes, rhythms, or controls for virtual instruments1,2. These interactive features enhance engagement, retention, and creativity compared to traditional classroom instruction, and somatosensory tools allow students to practice rhythm, coordination, and expression through body percussion, conducting gestures, and ensemble simulations3. Combined with AI-driven adaptive pathways, learners receive individualized content, real-time feedback, and progressive skill development that improve motivation and outcomes4,5.
Despite these developments, existing platforms often rely on limited modalities, lack continuity of personalization, or fail to adapt to diverse cultural and physical learning styles6,7. Traditional approaches also fall short in delivering real-time, data-driven adjustments that reflect the learner's evolving capabilities. For example, motion capture and wearable devices can generate rich datasets but are often underutilized in adaptive instruction8,9. Furthermore, while music libraries and learning management systems have expanded accessibility, they rarely provide dynamic personalization across sessions, which is critical in multicultural and heterogeneous learning contexts10.
To address these gaps, this study proposes a novel Trust Region Policy Optimized Deep Residual Long Short-Term Memory (TRPO-ResLSTM) framework for music education platforms11. The system integrates advanced preprocessing methods, including Wiener filtering and Z-score normalization, with Fast Fourier Transform for frequency-domain feature extraction. Res-LSTM provides robust recognition of gestures and temporal sequences, while TRPO reinforcement learning dynamically adjusts task difficulty in response to learner performance. Incremental learning further strengthens personalization by updating models across sessions.
Experiments were conducted on the Kaggle music gesture and rhythm dataset comprising 2,730 samples, split into training, validation, and testing subsets. Results show that the proposed method consistently outperforms baseline multimodal architectures, achieving accuracy, precision, recall, and F1 values in the range of 93%-95%. Ablation analyses confirm the effectiveness of both TRPO and Res-LSTM components. By enhancing rhythm accuracy, user engagement, and policy stability in real-time, the framework provides a practical solution for improving music education efficiency in resource-constrained and remote learning environments. Related work on AI-driven music education has highlighted the potential of somatosensory engagement, adaptive learning personalization, and even applications in music therapy and automated composition12,13. This study builds on these findings by offering a reproducible protocol that combines reinforcement learning with deep temporal modeling to advance the field of intelligent music education.