Research Article

Classroom Behavior Recognition via Transformer and Attention Mechanism: Application in Teaching Quality Evaluation for AI Medical Education

DOI:

10.3791/69052

August 15th, 2025

In This Article

Summary

Loading...
$$\rightleftharpoonup{xx}$$ $$\longleftharp{xx}$$, $$\longrightharp{xx}$$,

This study presents a Transformer-based AI system for recognizing classroom behavior in medical and nursing education. Using multimodal data, the system assesses learner engagement and emotional states. The results demonstrate higher accuracy compared to traditional models, enabling real-time feedback and enhanced support for teaching quality and student mental health.

Abstract

Loading...
$$\rightleftharpoonup{xx}$$ $$\longleftharp{xx}$$, $$\longrightharp{xx}$$,

This study introduces a Transformer-based artificial intelligence system designed for classroom behavior recognition, aimed at enhancing teaching quality assessment and mental health monitoring in medical and nursing education. The model utilizes multimodal data, specifically video, audio, and skeletal movement, to evaluate key indicators of student engagement, such as attention span, interaction frequency, and emotional fluctuations. A novel multimodal fusion strategy, combined with spatiotemporal attention mechanisms, enables the system to capture complex classroom behaviors with improved precision and robustness. Experimental results on the EduSense and SBU Kinect datasets demonstrate that the proposed method significantly outperforms traditional models in both accuracy and false alarm rate. Additionally, ablation studies confirm the contribution of spatial and temporal attention modules to the system's recognition capability. The framework supports real-time feedback and visual behavior mapping, offering practical value in experiential learning environments. This approach provides a scalable and adaptive solution for classroom behavior monitoring and supports early intervention strategies in medical education.

Introduction

Loading...
$$\rightleftharpoonup{xx}$$ $$\longleftharp{xx}$$, $$\longrightharp{xx}$$,

With the increasing mental health challenges among nursing and medical students due to academic and clinical stress, integrating artificial intelligence into classroom environments can provide not only teaching quality monitoring but also real-time psychological support. This study positions behavior recognition within medical education to support experiential learning and promote student well-being1. Intelligent classroom refers to a new and high-quality educational model that combines intelligent and personalized teaching processes through the application of cutting-edge technologies, such as information technology, artificial intelligence, big data, and the Internet of Things, to assist teaching management and classroom interaction2,3,4. Currently, with the development of technologies such as cameras and sensors, the behavioral data collected in the classroom not only includes videos but also multimodal data, such as voice and audio5. However, handling this complex spatiotemporal information and data and accurately extracting valuable behavioral features from it are the main challenges facing the field of intelligent classroom behavior recognition6,7. Existing behavior recognition approaches primarily rely on feature extraction from single-modal data sources, such as classroom video streams or sequential images8. While these methods can detect coarse behavioral events, they often fail to model the temporal evolution and spatial nuances of students' gestures, expressions, or interaction frequency factors critical for understanding learner well-being and focus9. Moreover, current models lack robust mechanisms to fuse heterogeneous data sources like audio cues and skeletal movement, which limits their sensitivity to affective or disengaged behavior10. To address these limitations, this study proposes a transformer-based classroom behavior recognition system enhanced with spatial-temporal attention mechanisms. By leveraging the Transformer's ability to model long-range dependencies and combining it with multimodal feature fusion strategies, the proposed system aims to improve the precision of behavior classification while enabling real-time visualization of engagement and emotional indicators. The ultimate goal is to support dynamic teaching quality assessment and embedded mental health feedback within medical education, offering a scalable, adaptive solution for monitoring and enhancing both instructional efficacy and learner well-being.

The evaluation proposes a novel Transformer-based behavior recognition system that combines spatiotemporal attention mechanisms with multimodal data fusion (video and skeleton inputs) tailored for smart classroom observation in medical and nursing education. In contrast to existing models, the new model facilitates accurate, real-time student attentiveness and emotional state detection, which is valuable for both instructional quality evaluation and psychological well-being.

Related works
The intelligent classroom is a new teaching environment that utilizes modern information technology to enhance teaching interaction, personalized learning, and teaching management. It can use intelligent teaching methods to improve teaching efficiency and quality while providing students with richer and more personalized learning experiences. Therefore, many scholars around the world have researched intelligent technology and classroom teaching. Radosavljevic et al. proposed an intelligent classroom model based on environmental intelligence, which evaluates students' fatigue levels and adjusts learning strategies by analyzing their daily academic activity data to optimize learning outcomes11. Tai T. Y studied the effect of intelligent personal assistants on improving foreign language speaking ability and found that using Google Assistant significantly improved the speaking ability of non-native speakers, with similar effects as native speakers in communication12. Chen et al. found through their research that using chatbots as teaching tools could quickly respond to student needs, enhance interactivity, and demonstrate the potential application of intelligent technology in the classroom13. Research raised a classroom teaching effectiveness evaluation method based on the intelligent teaching mode, which improved the accuracy of evaluation by using the cuckoo search algorithm and extreme learning machine14.

Behavior recognition in intelligent classrooms is a key technology for achieving personalized teaching and optimizing classroom management. It can monitor students' behavior status in real time and provide effective feedback. Through behavior recognition, teachers can accurately understand students' participation and focus, thereby improving teaching effectiveness and assisting decision-making. In this context, scholars worldwide have conducted extensive research on intelligent classroom behavior recognition. Research proposed a real-time visual monitoring system based on artificial intelligence, which used machine learning to train student behavior recognition models to help teachers better monitor student participation and classroom interaction, thereby optimizing teaching quality15. Chonggao P achieved real-time student behavior recognition and improved the accuracy of classroom behavior analysis in remote teaching by combining clustering analysis and random forest algorithm, as well as the human skeleton model16. Wu S developed a student behavior recognition model that combines particle swarm optimization and the K-nearest neighbor algorithm, promoting the accuracy of emotion recognition to meet the practical needs of classroom teaching17. The ACORN system proposed by Ramakrishnan et al. utilized multimodal machine learning and combined convolutional neural networks to analyze audio, facial expressions, and image data. Integrating these features through temporal convolutional networks achieved automatic evaluation of classroom atmosphere and made preliminary progress18.

In summary, existing research has proposed various machine learning and deep learning-based techniques for intelligent classroom behavior recognition, covering areas such as student fatigue detection, sentiment analysis, and classroom interaction monitoring. However, there are still some shortcomings in current research, and many methods lack sufficient real-time and scalability in practical applications, making it difficult to adapt to complex and changing classroom scenarios. Therefore, the study proposes an intelligent classroom behavior recognition method with Transformer and AM. The innovation of the research lies in utilizing the powerful temporal modeling capability of the Transformer combined with spatiotemporal AM to accurately capture key behavioral moments and regions in the classroom, and integrating multiple sources of information, such as video and audio, through multimodal data fusion to enhance the comprehensiveness and accuracy of behavior recognition.

Access restricted. Please log in or start a trial to view this content.

Protocol

Loading...
$$\rightleftharpoonup{xx}$$ $$\longleftharp{xx}$$, $$\longrightharp{xx}$$,

This study uses publicly available datasets that do not contain any personally identifiable information. All data used in the research have been anonymized to ensure the privacy of the participants. Informed consent was obtained from all participants at the time of data collection, and the study adheres to ethical guidelines. The protocol explains the intelligent classroom behavior recognition method, which consists of feature extraction based on multimodal data fusion and spatiotemporal behavior recognition based on Transformer and attention. The performance of the entire method is optimized through joint training.

1. Multimodel data collection

Multimodal Data Collection involved collecting synchronized classroom behavior data from two datasets: the EduSense Dataset and the SBU Kinect Interaction Dataset. EduSense recorded real-life university-classroom recordings, and SBU Kinect recorded interaction-based skeletal sequences. The datasets contained video streams, audio cues, and 3D skeletal joint information for modeling student postures and motion. Data preprocessing guaranteed temporal alignment and noise removal for cleaning. This convergence provided end-to-end behavior monitoring, recording visual gestures and body movement across many learning situations.

2. Feature extraction based on multimodal data fusion

Targeting the issue of student behavior recognition in intelligent classrooms, a method for intelligent classroom behavior recognition with Transformer and AM is proposed. Due to the complex nature of classrooms with numerous students, changes in factors within the scene can interfere with student behavior recognition. Additionally, single feature extraction has limitations such as incomplete information, susceptibility to environmental interference, and difficulty in capturing complex behaviors19,20. Therefore, the study first proposes a student classroom feature extraction method based on multimodal data fusion. This method combines data from video and skeletal modalities, utilizing their complementarity to achieve more comprehensive and accurate feature extraction of student behavior. In particular, every input video sequence is divided frame-wise. Frames are fed into a pretrained Faster R-CNN to localize human regions of interest. The detected regions are fed into a convolutional neural network backbone to obtain temporal visual features. For the skeletal modality, 3D joint positions are posed as sequences and encoded using a pretrained BERT model to obtain temporal and spatial dependencies. These are normalized and embedded before being input into the multimodal fusion network. Among them, the video modality is mainly responsible for capturing visual features such as students' movements and expressions, while the skeletal modality accurately describes students' posture changes through joint point information. While the video modality relies on the global features of visual perception, while the skeletal modality global visual features torepresent human behavior, the skeletal modality focuses on the dynamic changes in joint position, leading to significant semantic differences between the two21. Therefore, it is necessary to develop a specialized multimodal fusion network to effectively model the correlation between the two modalities. The multimodal fusion network is shown in Figure 1.

In Figure 1, the network is mainly broken into three modules. Among them, the input of the video modality processing module is a video sequence, which is extracted through a Faster Region-based Convolutional Neural Network (Faster R-CNN) for feature extraction. The treatment of the video modality starts by translating classroom video shots into frame-by-frame inputs for feature extraction. Frames are treated with a Faster Region-Based Convolutional Neural Network (Faster R-CNN) with a ResNet-50 base trained on ImageNet. The network treats RGB frames resized to 224 × 224 pixels, resulting in high-level visual feature maps with spatial and object-specific features of student activity. These features are then input to a self-attention module, which learns temporal relationships between frames prior to encoding them to spatiotemporal embeddings via a Transformer layer. This preserves faint behavioral signals, such as head tilts or hand raises, in embedding representations for later fusion. The extracted video features are captured by the self-attention module to capture the temporal and spatial relationships within the video modality, and then further modeled by a transformer to generate embedding vectors for the spatiotemporal information of the video sequence. The input of the skeletal modality processing module is a skeletal sequence, which is processed by Bidirectional Encoder Representation from Transformers (BERT) to extract temporal and spatial features of skeletal joints. For skeletal data, the pose of every student is modeled as a sequence of 2D joint coordinates from the classroom videos using an OpenPose-based pose estimator. These sequences are converted into embedding tokens in which every token corresponds to a frame with 18 keypoints. Temporal context and dependencies of these joint sequences are modeled using a Bidirectional Encoder Representations from Transformers (BERT) model. The model employs 12 transformer layers, 8 attention heads, and a hidden size of 768, which is the standard BERT-base architecture. The model gets fine-tuned during training to acquire behavior-specific joint patterns. The output embeddings represent the structural and motion-related characteristics of student behavior that are subsequently combined with video modality features. Subsequently, the skeletal features are further aggregated into spatial and temporal features through 3D CNN and pooling operations, generating embedded vectors of the skeleton as feature representations of the skeletal modality.

3. Cross-attention fusion module

In the cross-attention fusion module, the study uses the multi-head attention mechanism (MHAM) to construct the interaction relationship between video modality and skeletal modality, and further extracts more compact fusion features through 3D CNN and pooling operations from the generated multimodal fusion features. The fusion process is accomplished using a two-stream attention network, wherein each modality's temporal embeddings are passed through 3D CNNs and bidirectional GRUs. The multimodal data streams are aligned using Multi-Head Attention Modules (MHAM), with scaled dot-product attention to attain inter-modality dependencies. The embeddings are then pooled down to tackle dimensional complexity and sent to a fully connected Softmax layer for classification.

In multimodal fusion networks, the self-attention module is used to model the temporal and spatial dependence within a single mode, while the cross-attention fusion module is utilized to establish the association and information interaction between different modes. Therefore, both are very important. The self-attention module captures the internal dependencies of the input sequence by calculating the relationship between the query Q, key K, and value V. The calculation of Q, key K, and value V is shown in equation (1). Where Matrix dimensions in linear algebra equation, Q, K, V matrices with Rn×dk dimensions. are the query, key, and value matrices respectively; dk is the key vector dimensionality, and dk is the dimensionality of the value vectors. The dimension factor dk is adopted in order to avoid the dot products from becoming large, which would influence the stability of gradients during training.

Equations representing query, key, and value in transformer neural network model computations.      (1)

In equation (1), Mathematical formula X=∼n×d_model, equation related to model dimensions in linear algebra. represents the input features, where n is the length of the input sequence, dmodel is the dimension of the features, and WQ, WK, and WV represent three sets of learnable projection matrices. Through these matrix mappings, the input features are transformed into queries, keys, and values. The dot product of query and key K are utilized to calculate attention weights and scale them, as shown in equation (2).

Attention mechanism formula, softmax(QK^T/sqrt(d_k))V equation, neural network concept.      (2)

In equation (2), dk represents the dimension of querying Q and key K, and softmax denotes the softmax function utilized to obtain the attention weight matrix. Scaling can effectively prevent the problem of the gradient being too small due to excessive dot product values caused by dimensionality. The attention weight matrix is applied to the value V to obtain the output Z of the self attention module, as shown in equation (3).

Mathematical equation for linear combination; z=ΣaiVj; used in data analysis; formula image.      (3)

In equation (3), aij means the attention weight of the th query vector on the j th key vector. The specific structure of the cross attention fusion module is denoted in Figure 2.

From Figure 2, this module mainly starts processing input features from skeletal and video modalities. Firstly, the skeleton sequence and video sequence are separately subjected to 3D convolutional layers to extract spatiotemporal features, and then these spatiotemporal features are modeled using Bidirectional Gated Recurrent Units (Bi-GRUs) for temporal modeling. Next, the features of the two modalities are further extracted through MHAMs to extract correlations within the modalities. Each modality achieves feature representation internally through mechanisms of queries, keys, and values. Then, time-averaged pooling is performed on the output features to reduce the complexity of the time dimension. After pooling, the multimodal features are statistically pooled to calculate the mean and standard deviation of each modality, and finally output the student's action recognition and classification through a fully connected layer (FCL) and Softmax classifier. Therefore, the feature extraction based on multimodal data fusion proposed in the study utilizes self-attention and cross AMs to accurately capture and model the spatiotemporal features of student behavior by fusing data from video and skeletal modalities, thereby improving the accuracy and comprehensiveness of student behavior recognition in intelligent classrooms.

4. Spatiotemporal behavior recognition based on Transformer and attention

The feature extraction based on multimodal data fusion achieves a preliminary improvement in the comprehensiveness and accuracy of student behavior by fusing data from video and skeletal modalities. However, in the context of intelligent classrooms, students' behavior often changes over time, and the same behavior may exhibit different characteristics at different time points and spatial regions. Relying solely on static information fused from multimodal features cannot fully capture the complex spatiotemporal dynamic relationships in the behavior sequence22,23. Therefore, further research proposes a spatiotemporal behavior modeling and AM based on Transformer. The study chose Vision Transformer (ViT) as the network architecture, which can model the global spatiotemporal information of input through self AM and has a stronger ability to capture spatiotemporal dependencies. Following multimodal fusion, the fused features are again represented using a Vision Transformer (ViT) to forecast the spatiotemporal behavior. The feature maps in the input are split into disjoint 16 × 16 patches, linearly embedded into one-dimensional vectors, and positionally encoded to preserve spatial order. ViT structure has 12 Transformer encoder blocks, made up of multi-head self-attention (8 heads) and feed-forward layers whose hidden dimension is 768. To enhance behavioral saliency, Spatial Attention (SA) and Temporal Attention (TA) modules are suggested. SA module encourages feature relevance in spatial areas via Tanh activation, while TA module highlights important temporal frames via ReLU activation. These two attention streams are subsequently combined with the ViT backbone via a two-stage training method such that the model learns to capture the dynamic evolution of classroom behavior efficiently. The structure of ViT is shown in Figure 3.

In Figure 3, in ViT, the original input is an image that is segmented into multiple fixed-sized small blocks, each represented as a part of the image, and linearly flattened into a one-dimensional vector. Since the Transformer cannot perceive location information, each small block is embedded with location information before being input into Transformer to ensure that it can capture the relative spatial position relationship between different small blocks. The core module of ViT is the Transformer encoder, which is composed of multiple stacked Transformer encoding blocks. The encoding blocks are composed of an MHAM and a feed-forward neural network. The MHAM captures the dependencies between input sequences in parallel, while the feed-forward neural network performs nonlinear transformations on the results to enhance the model's ability to capture and express global information. After the Transformer encoder processes all the small pieces of information, the output of the last encoding layer is passed to the Multi-Layer Perceptron Head (MLP Head), which outputs the final student action recognition and classification results.

Practically, each frame is split into non-overlapping 16 × 16 patches, flattened, and positionally embedded in order to maintain spatial relations. Patch embeddings are processed sequentially by stacked transformer encoders in the ViT architecture. Spatial Attention (SA) module utilizes tanh activation to vary the attention across areas of interest within each frame, and Temporal Attention (TA) module utilizes ReLU to highlight temporally important frames. This two-attention mechanism enables effective modeling of space and time behavior dynamics.

To further enhance the performance of ViT in complex behavioral sequences, spatial attention (SA) and temporal attention (TA) are introduced to enable ViT to not only autonomously select important spatial joints, but also focus on behaviors at different time points, better capturing the spatiotemporal dynamic changes of behavior. The SA module and TA module are denoted in Figure 4.

In Figure 4A, the SA module processes the spatial dimension information of the input by passing the current frame's features into the ViT layer to extract the hidden states. These are combined with the features of the previous frame and jointly fed into the FCL. The FCL will perform a linear transformation on the features to generate latent feature representations. Subsequently, it is passed to the Tanh activation function to map the output to the range of [-1,1], enhancing nonlinearity and better distinguishing between important and unimportant features. Finally, the features are normalized through a normalization layer to ensure that the output features have a relatively stable scale. In Figure 4B, the TA module focuses more on the key information contained in each frame in the temporal dimension. Unlike Tanh in SA, the TA module uses the ReLU activation function to suppress negative values and only retains the output of positive values, effectively removing negative noise information and enhancing the model's temporal capture ability.

5. Joint training method

To effectively solve the problem of biased and redundant spatiotemporal features caused by the mutual influence between ViT, SA module, and TA module in spatiotemporal behavior recognition based on Transformer and attention, a joint training method is adopted to optimize the three modules. The specific process is denoted in Figure 5.

In Figure 5, the joint training strategy first inputs the model training parameters, sets the network structure, and hyperparameters. Subsequently, separate pre-training is conducted on SA and TA to better learn local features. It is worth noting that during the pre-training of one model, the weights of the other model are kept fixed and are not updated. After completing individual pre-training, the ViT layer is combined for training. Then, two attention models will be fixed, and the main ViT network will be trained separately. Finally, by jointly training the main ViT network, SA, and TA modules, the overall network is optimized to generate the final model for intelligent classroom behavior recognition. The training procedure is carried out stepwise: (1) autonomous pre-training of the SA and TA modules with frozen ViT parameters, (2) stepwise training of ViT with frozen attention module weights, and (3) end-to-end fine-tuning of all the components together with Adam optimizer (learning rate = 0.001, batch size = 64, epochs = 100). The approach stabilizes feature learning and reduces interference among modules during optimization.

Overall, the proposed intelligent classroom behavior recognition method combines video and skeletal modalities to model spatiotemporal relationships through self-attention and cross AMs. By utilizing ViT's global spatiotemporal modeling capability and combining spatial and TA modules, the model is optimized to capture key behaviors and temporal dynamics, achieving more comprehensive and accurate student behavior recognition. Critical fusion operations, wherein multimodal feature representations are combined, are performed in the architecture under consideration. Specifically, the spatial features learned by CNN are connected to the temporal features of the Bi-LSTM. This output is then fused with attention-augmented features of MHAM before the classification task. These operations of fusion are crucial for the fusion of complementary views of behavioral data and are emphasized in Figure 1 and Figure 5.

Algorithm 1 illustrates the combined multimodal data, skeletal, text, and video information, using ViT, BERT, and GCN-based models to extract informative features in a class. The joint representation is then marked to identify student behavior, with higher accuracy using cross-modal learning.

Algorithm 1: Process of multimodal behavior recognition using ViT, BERT, and GCN.

Initialize:

Pre-trained BERT model

Pre-trained Faster R-CNN model

Pre-trained ViT model

GCN model for skeleton data

Classifier (e.g., MLP or Transformer)

Load dataset:

For each sample in multimodal dataset:

Extract video frames

Extract audio transcript

Extract skeleton data (OpenPose or MediaPipe)

Step 1: Video modality processing

For each frame in video:

Apply Faster R-CNN to detect objects/participants

Apply Vision Transformer (ViT) to extract frame-level features

Aggregate features over time for temporal representation

Step 2: Text modality processing

Preprocess transcript text:

Tokenize using BERT tokenizer

Pass tokens through BERT model

Extract contextual embeddings from the [CLS] token or average pooling

Step 3: Skeletal modality processing

For each skeleton sequence:

Normalize coordinates

Pass through GCN or ST-GCN to extract motion features

Step 4: Feature fusion

Concatenate or attention-fuse:

Video features

Text features

Skeletal features

Obtain a unified multimodal representation

Step 5: Classification

Pass the fused representation to the final classifier

Predict classroom behavior label (e.g., attentive, distracted)

Step 6: Evaluation

Compare predictions with ground truth

Compute metrics: Accuracy, Precision, Recall, F1 Score

Access restricted. Please log in or start a trial to view this content.

Results

Loading...
$$\rightleftharpoonup{xx}$$ $$\longleftharp{xx}$$, $$\longrightharp{xx}$$,

To assess the efficacy and superiority of the raised intelligent classroom behavior recognition method with Transformer and AM, experimental performance verification and analysis were conducted. Firstly, the experimental environment and parameters were configured, as denoted in Table 1.

Based on Table 1, the study selected the EduSense Dataset and SBU Kinect Interaction Dataset as experimental datasets, named D1 and D2, respectively. D1 collected the behaviors...

Access restricted. Please log in or start a trial to view this content.

Discussion

Loading...
$$\rightleftharpoonup{xx}$$ $$\longleftharp{xx}$$, $$\longrightharp{xx}$$,

A behavior recognition method for intelligent classrooms based on Transformer and AM was proposed to address the challenges of complex student behavior and multimodal data recognition. A ViT network with global spatiotemporal modeling capability was constructed by integrating data from video and skeletal modalities, and self-attention and cross AMs were utilized to capture spatiotemporal features of behavior. By integrating video and audio data into a new Intelligent Sensing Framework (ISF) framework using a Multimodal F...

Access restricted. Please log in or start a trial to view this content.

Disclosures

Loading...
$$\rightleftharpoonup{xx}$$ $$\longleftharp{xx}$$, $$\longrightharp{xx}$$,

The authors have nothing to disclose.

Acknowledgements

Loading...
$$\rightleftharpoonup{xx}$$ $$\longleftharp{xx}$$, $$\longrightharp{xx}$$,

None.

AUTHOR CONTRIBUTION:
Liu Lei: Responsible for the conception and design of the research; proposed the Transformer-based multimodal data fusion method for intelligent classroom behavior recognition. Led the development of the model and experimental design, managed data collection and processing, and drafted the initial manuscript. Contributed to the analysis and discussion of experimental results.

He Zhihua: Provided overall guidance and supervision of the study; contributed to the research framework and methodological support. Participated in model optimization and experimental analysis, reviewed and revised the manuscript, ensured compliance with ethical standards, and gave final approval for submission. Responsible for correspondence.

Access restricted. Please log in or start a trial to view this content.

Materials

List of materials used in this article
NameCompanyCatalog NumberComments
128 GB DDR4 RAMMultiple vendorshttps://www.crucial.com/memory/ddr4Sufficient memory for multimodal data processing
2 TB SSD StorageMultiple vendorshttps://www.samsung.com/ssdFor storing datasets and models
BERT (Base configuration)Hugging Facehttps://huggingface.co/bert-base-uncasedLanguage model for skeletal sequence embedding
CUDA 11.1 ToolkitNVIDIAhttps://developer.nvidia.com/cuda-toolkitEnables GPU acceleration for PyTorch
cuDNN 8.0 LibraryNVIDIAhttps://developer.nvidia.com/cudnnGPU-accelerated library for deep neural networks
EduSense DatasetCarnegie Mellon Universityhttps://www.cs.cmu.edu/~./edusenseReal-life university classroom dataset
Faster R-CNN (ResNet-50)Open Source (PyTorch)https://github.com/facebookresearch/detectron2Object detector used for video feature extraction
Intel Xeon E5-2680 v4 CPUIntelhttps://www.intel.com/products/xeon-e5-2680-v4High-performance processor for model training
MatplotlibOpen Sourcehttps://matplotlib.orgUsed for plotting performance metrics
NVIDIA Tesla V100 GPUNVIDIAhttps://www.nvidia.com/en-us/data-center/tesla-v100Used for training and inference; supports real-time behavior recognition
OpenPoseCMUhttps://github.com/CMU-Perceptual-Computing-Lab/openposeExtracts 2D skeletal keypoints from video frames
PyTorch 1.8 FrameworkOpen Sourcehttps://www.pytorch.orgDeep learning library used for model development
SBU Kinect Interaction DatasetStony Brook Universityhttps://www3.cs.stonybrook.edu/~kyunghan/sbu_kinectInteraction-based skeletal dataset
TensorBoardOpen Source (Google)https://www.tensorflow.org/tensorboardUsed for training visualization
Ubuntu 20.04 LTS Operating SystemCanonicalhttps://www.ubuntu.com/downloadLinux OS used as the development environment
Vision Transformer (ViT)Google / Hugging Facehttps://huggingface.co/google/vit-base-patch16-224Used for global spatiotemporal behavior modeling

References

Loading...
$$\rightleftharpoonup{xx}$$ $$\longleftharp{xx}$$, $$\longrightharp{xx}$$,
  1. He, F., Yuan, X. An analysis of ways to optimize teacher-student relationship in ideological and political wisdom classroom teaching in senior high school. Int J Educ Humanit. 14 (1), 40-44 (2024).
  2. Tambunan, M. A., Sirait, J., Silitonga, I. Implementation of the Glasser model with the help of animation media based on local wisdom to improve the writing skills of UHNP Indonesian language education program students. IDEAS J Engl Lang Teach Learn Linguist Lit. 12 (2), 1110-1117 (2024).
  3. Munisa, M., Putri, U. N., Sari, W. V., Fitri, N. A. Digital literacy based on local wisdom in inclusive education. Pionir J Pendidik. 13 (1), 120-124 (2024).
  4. Intelligent classroom perception system based on artificial intelligence. Eur Alliance Innov. Zhang, T. EAI Urb-IoT 2021: Proc 6th EAI Int Conf IoT Urban Space, 20-21 Dec 2021, Shenzhen, China, 51, (2022).
  5. Badawi, H. Exploring classroom discipline strategies and cultural dynamics: Lessons from the Japanese education system. Tafkir Interdiscip J Islam Educ. 5 (1), 1-12 (2024).
  6. Yang, F. A spatio-temporal attention-based method for detecting student classroom behaviors. arXiv. , (2023).
  7. Wang, Z., Wang, M., Zeng, C., Li, L. Multi-scale deformable transformers for student learning behavior detection in smart classroom. Xiv. , (2024).
  8. Lin, L., Yang, H., Xu, Q., Xue, Y., Li, D. Research on student classroom behavior detection based on the real-time detection transformer algorithm. Appl Sci. 14 (14), 6153(2024).
  9. Wang, Z., Yao, J., Zeng, C., Li, L., Tan, C. Students' classroom behavior detection system incorporating deformable DETR with Swin transformer and lightweight feature pyramid network. Systems. 11 (7), 372(2023).
  10. Lyons, J., Tarc, P. How might IB classroom pedagogy 'make a better world?' (Toward) illuminating a promising IBDP teacher praxis. Globalisation Soc Educ. 22 (4), 605-624 (2024).
  11. Radosavljevic, V., Radosavljevic, S., Jelic, G. Ambient intelligence-based smart classroom model. Interact Learn Environ. 30 (2), 307-321 (2022).
  12. Tai, T. Y. Effects of intelligent personal assistants on EFL learners' oral proficiency outside the classroom. Comput Assist Lang Learn. 37 (5-6), 1281-1310 (2024).
  13. Chen, Y., Jensen, S., Albert, L. J., Gupta, S., Lee, T. Artificial intelligence (AI) student assistants in the classroom: Designing chatbots to support student success. Inf Syst Front. 25 (1), 161-182 (2023).
  14. Chen, J., Lu, H. Evaluation method of classroom teaching effect under intelligent teaching mode. Mob Netw Appl. 27 (3), 1262-1270 (2022).
  15. Trabelsi, Z., Alnajjar, F., Parambil, M. M. A., Gochoo, M., Ali, L. Real-time attention monitoring system for classroom: A deep learning approach for student's behavior recognition. Big Data Cogn Comput. 7 (1), 48-56 (2023).
  16. Chonggao, P. Simulation of student classroom behavior recognition based on cluster analysis and random forest algorithm. J Intell Fuzzy Syst. 40 (2), 2421-2431 (2021).
  17. Wu, S. Simulation of classroom student behavior recognition based on PSO-kNN algorithm and emotional image processing. J Intell Fuzzy Syst. 40 (4), 7273-7283 (2021).
  18. Ramakrishnan, A., Zylich, B., Ottmar, E., LoCasale-Crouch, J., Whitehill, J. Toward automated classroom observation: Multimodal machine learning to estimate class positive climate and negative climate. IEEE Trans Affect Comput. 14 (1), 664-679 (2021).
  19. Lee, J., Kim, H. Discrete cosine transformed images are easy to recognize in vision transformers. IEIE Trans Smart Process Comput. 12 (1), 48-54 (2023).
  20. Rachman, R. K., Setiadi, D. R. I. M., Susanto, A., Nugroho, K., Islam, H. M. M. Enhanced vision transformer and transfer learning approach to improve rice disease recognition. J Comput Theor Appl. 1 (4), 446-460 (2024).
  21. Wu, H., Triebe, M. J., Sutherland, J. W. A transformer-based approach for novel fault detection and fault classification/diagnosis in manufacturing: A rotary system application. J Manuf Syst. 67 (1), 439-452 (2023).
  22. Zhong, X., Gu, Y., Luo, Y., Zeng, X., Liu, G. Bi-hemisphere asymmetric attention network: Recognizing emotion from EEG signals based on the Transformer. Appl Intell. 53 (12), 15278-15294 (2023).
  23. Jamali, A., Roy, S. K., Bhattacharya, A., Ghamisi, P. Local window attention transformer for polarimetric SAR image classification. IEEE Geosci Remote Sens Lett. 20 (1), 1-5 (2023).
  24. Zhao, X. M., Yusop, F. D. B., Liu, H. C., Prilanita, Y. N., Chang, Y. X. Classroom student behavior recognition using an intelligent sensing framework. IEEE Access. 13, 49767-49776 (2025).
  25. Huang, Y., et al. A method for classroom behavior state recognition and teaching quality monitoring. Int J Intell Comput Cybern. 18 (2), 382-396 (2025).

Access restricted. Please log in or start a trial to view this content.

Reprints and Permissions

Request permission to reuse the text or figures of this JoVE article

Request Permission

Tags

Classroom Behavior RecognitionTransformer ModelAttention MechanismTeaching Quality EvaluationAI Medical EducationMultimodal FusionSpatiotemporal AttentionStudent EngagementReal Time FeedbackBehavior Mapping
Video Coming Soon

Related Articles