A subscription to JoVE is required to view this content. Sign in or start your free trial.

Method Article

Dual-Branch Cross-Attention Network for Electroencephalography and Facial Multimodal Emotion Recognition During Painting Viewing

98 views

DOI:

10.3791/70600

May 12th, 2026

In This Article

Summary

This protocol describes synchronized acquisition of electroencephalography and facial data in painting-viewing settings and a dual-branch cross-attention network for multimodal emotion recognition, including preprocessing, feature fusion, and classification of four emotional states.

Abstract

Emotion recognition during painting viewing remains difficult because aesthetic responses are dynamic, context-dependent, and only partly observable through outward behavior. Unimodal approaches based only on facial expressions or only on physiological signals, therefore, often provide incomplete representations of the viewer's experience. This article describes a reproducible method for constructing a dual-branch cross-attention network that integrates facial video with electroencephalography (EEG) data for multimodal emotion recognition in a painting exhibition context. The protocol begins with the establishment of a synchronized acquisition environment using a 64-channel EEG system and a high-resolution camera for concurrent recording during painting viewing. The subsequent procedure details signal preprocessing, including EEG filtering, segmentation, and differential entropy feature extraction, together with facial detection, alignment, and frame preparation for video analysis. The neural network contains two modality-specific branches. The EEG branch uses a convolutional neural network and a bidirectional long short-term memory network to model spatial-temporal neural features, whereas the facial branch uses a coordinate attention-enhanced lightweight convolutional network to extract visual expression features. A multi-head cross-attention module is then used to fuse the two streams by learning intermodal relevance. Under the experimental conditions reported here, the multimodal framework achieved higher classification accuracy than the unimodal comparison models in four-category emotion recognition. This method is most suitable for offline analysis in controlled acquisition settings and provides a practical framework for affective computing, empirical aesthetics, and exhibition-oriented human-computer interaction research.

Introduction

Emotion is a multidimensional psychological and physiological process shaped by external stimuli, internal appraisal, and ongoing cognitive regulation, and it therefore occupies a central position in human-computer interaction and cognitive science1. In painting exhibition settings, emotion also mediates the relationship between visual artworks and viewer interpretation. For this reason, the identification of viewer emotional states has practical value for exhibition evaluation, spatial design, and empirical studies of aesthetic experience2. At the same time, emotion measurement in semi-natural exhibition environments remains technically difficult because responses are subtle, transient, and influenced by both the artwork and the viewing context.

Emotion recognition research has generally developed along two major lines: behavioral analysis and physiological analysis. Behavioral approaches, especially facial expression analysis based on computer vision, are widely used because they are noninvasive and relatively easy to deploy3. Deep learning models such as residual networks and lightweight convolutional networks have shown strong performance in basic facial emotion classification tasks4. However, facial behavior during painting viewing may be weak, socially moderated, or deliberately restrained, which can reduce the correspondence between visible expression and subjective feeling5. Physiological approaches, particularly those based on electroencephalography (EEG), offer more direct access to neural activity associated with affective processing6. EEG has been widely used in emotion studies because temporal fluctuations in neural signals are informative for changes in affective state, including dimensions related to valence and arousal7. Even so, EEG recordings are sensitive to motion artifacts, electromyographic contamination, and environmental noise, and EEG alone often provides limited contextual information about what the viewer is visually responding to8.

These complementary strengths and weaknesses have increased interest in multimodal emotion recognition9. The central premise of multimodal fusion is that heterogeneous signals can provide mutually informative evidence and reduce the ambiguity that arises when a single modality is interpreted in isolation10. In painting-viewing tasks, for example, restrained facial behavior may still coincide with measurable neural changes associated with attention or affective engagement. However, existing multimodal systems do not always model this relationship effectively. Early fusion strategies based on direct feature concatenation can increase dimensional burden and may blur modality-specific temporal structure11. Late fusion strategies based on separate decisions are easier to implement, but they may overlook fine-grained interactions between visual and physiological signals during emotional processing12. In addition, many multimodal models use fixed or weakly adaptive fusion schemes, which are not well suited to fluctuating viewer responses in exhibition-like settings13.

To address these limitations, the present protocol describes a dual-branch cross-attention network for multimodal emotion recognition during painting viewing. In this framework, EEG features are processed by a convolutional neural network-bidirectional long short-term memory model (CNN-BiLSTM), which is used to represent spatial-temporal neural dynamics. Facial video features are processed by a coordinate attention-enhanced MobileNetV2 branch, which is used to capture low-intensity expression cues while maintaining computational efficiency14. The two branches are then integrated through a multi-head cross-attention mechanism that learns relevance between modalities rather than simply concatenating their outputs15,16. This design is intended to preserve modality-specific structure while improving cross-modal interaction modeling.

The method is designed primarily for offline analysis under controlled data-acquisition conditions rather than for direct real-time deployment in open public exhibition spaces. Reliable implementation requires synchronized EEG and video capture, stable lighting, controlled camera placement, acceptable electrode impedance, and sufficient computing resources for model training. Performance may also depend on sample composition, stimulus type, and the degree to which participants modify behavior because they know they are being recorded. These operational constraints should be considered when adapting the protocol to other exhibition settings or populations.

The protocol presented here covers the full experimental pipeline, including synchronized acquisition from 327 participants, preprocessing of EEG and facial video data, feature extraction, cross-modal fusion, and classification. The aim is to provide a reproducible workflow for multimodal emotion recognition in painting exhibition research. Beyond affective computing17, the protocol may also support quantitative investigation in empirical aesthetics and related psychological studies18.

Access restricted. Please log in or start a trial to view this content.

Protocol

The study protocol was reviewed and approved by the Ethics Committee on Human Research Protection of Xingyi Minzu Normal University (Approval No. 2600AL0621). Written informed consent was obtained from all participants before their inclusion in the study and prior to data collection.

1. Participant recruitment and ethical compliance

  1. Obtaining ethical approval
    1. Submit the complete experimental protocol, consent form, participant information sheet, and data-protection plan to the institutional review board (IRB) or ethics committee before starting the study.
    2. Obtain written approval and record the approval number in the study documentation.
  2. Recruitment of participants
    1. Recruit healthy adult participants for multimodal emotion acquisition during painting viewing. In this demonstration protocol, recruit 327 participants.
    2. Apply the following inclusion criteria: age 18–35 years, normal or corrected-to-normal vision, no self-reported history of neurological disorders, no current psychoactive medication use, and willingness to undergo electroencephalography (EEG) recording and facial video recording.
    3. Apply the following exclusion criteria: excessive scalp injury or dermatological conditions preventing electrode placement, inability to complete the full recording session, severe motion during acquisition, or incomplete consent documentation.
    4. Record participant characteristics, including age, sex, handedness, prior art-viewing experience, and education level. In this demonstration protocol, record the following sample SHORT LONG ABSTRACT: mean age 24.8 ± 3.7 years, 165 females and 162 males, all participants right-handed.
    5. Define demographic balance operationally before recruitment. In this demonstration protocol, use this term to indicate an approximately balanced sex distribution and an age range concentrated in early adulthood to reduce age-related variance in EEG and facial expression responses.
  3. Obtaining informed consent
    1. Provide each participant with a written informed consent form describing EEG recording, facial video recording, synchronization procedures, anonymization, data storage, and the right to withdraw at any time without penalty.
    2. Explain that facial images will be used for feature extraction and that eye regions will be masked in all manuscript figures to protect participant identity.
    3. Obtain written informed consent before any experimental procedure begins.
  4. Assigning identifiers and anonymizing data
    1. Assign each participant a unique numerical study ID. Store EEG files, video files, and metadata using only the study ID.
    2. Store identifiable consent forms separately from physiological and image data. Use de-identified facial frames or facial landmark-derived outputs for internal analysis storage when required by the local ethics policy.

2. Experimental environment and stimulus setup

  1. Preparation of the viewing environment
    1. Prepare a quiet indoor room to simulate a controlled painting exhibition environment. Maintain an ambient temperature of 22–24 °C and a relative humidity of 45%–60%. Control background noise below 40 dB where possible.
    2. Seat the participant in an upright chair with back support and position the chair so that the viewing distance between the participant and the display is 2.0 m.
    3. Instruct the participant to remain seated alone in the recording area during stimulus presentation. Do not allow other participants or observers to be in the immediate field of view during data acquisition.
  2. Configuration of lighting
    1. Install diffuse ring or circular lighting in front of the participant to reduce facial shadowing. Set illumination to a stable level of approximately 500 lux at face level.
    2. Avoid rapid lighting fluctuation and direct glare on the eyes, forehead, or cheeks. Visual checkpoint: Confirm that the full face, including the forehead, eyebrows, eyes, nose, cheeks, and mouth region, is visible in the camera preview without overexposure or deep shadow.
  3. Configuration of visual stimulus presentation
    1. Present the visual stimuli on a monitor or projection screen. In this demonstration protocol, use a 27-inch liquid crystal display monitor with a resolution of 1,920 x 1,080 pixels and a refresh rate of 60 Hz.
    2. Display painting-viewing stimuli in video or slideshow format according to the study design. Use the same screen brightness, contrast, and presentation order rules for all participants.
    3. Present each painting clip for 90–120 s. In this demonstration protocol, use 6 clips, each of 100 s duration, with a 20 s interstimulus rest interval.
      ​CAUTION: Ground all electrical equipment properly and place the EEG amplifier and power cables away from high-interference sources to reduce 50/60 Hz line noise.

3. Multimodal data acquisition

  1. Acquiring facial video
    1. Position a high-definition industrial camera directly in front of the participant at approximately eye level. Set the camera-to-face distance to 1.2 m.
    2. Use a camera with a minimum acquisition resolution of 1,920 x 1,080 pixels and a frame rate of 30 frames/s.
    3. Set exposure manually to avoid frame-to-frame fluctuation. In this demonstration protocol, use ISO 400, shutter speed 1/60 s, and fixed white balance.
    4. Save the video in MP4 or AVI format using a constant frame rate. In this demonstration protocol, save video as MP4, H.264 encoding, 30 frames/s.
    5. Verify that the complete face remains within the field of view throughout the recording session. Visual checkpoint: Confirm that the eyebrows and forehead are fully visible. Reposition the camera immediately if the forehead, brow region, or chin is cropped.
      NOTE: Use 30 frames/s because this frame rate provides adequate temporal resolution for low-intensity facial-expression dynamics while maintaining a manageable storage and processing load in synchronized multimodal recording.
  2. Acquiring EEG signals
    1. Measure head circumference and select the correct cap size for the participant. Place the 64-channel EEG cap according to the international 10–20 system or a manufacturer-specified 64-channel extension layout.
    2. Align the cap using anatomical landmarks. Position Fpz on the midline above the nasion and verify symmetry across the left and right hemispheres.
    3. Part the hair at each electrode site and apply conductive gel to improve electrode-scalp contact.
    4. Prepare the scalp gently at high-impedance sites using a cotton swab or blunt preparation tool. Do not abrade the skin excessively.
    5. Connect the cap to the EEG amplifier and perform impedance measurement. Reduce electrode impedance to below 10 kΩ for all channels. In this demonstration protocol, aim for < 5 kΩ for frontal and central electrodes when possible.
    6. Set the sampling frequency to 500 Hz. Set the online reference according to the amplifier configuration. In this demonstration protocol, use a linked mastoid reference during acquisition and perform common average re-referencing during preprocessing.
    7. Place the ground electrode according to the amplifier specification. In this demonstration protocol, place the ground at AFz. Record at least 60 s of resting baseline before stimulus presentation.
    8. Visual checkpoint: Inspect the live EEG traces before starting the experiment. Confirm stable baseline activity, low drift, and absence of persistent saturated channels or major motion artifacts.
  3. Synchronizing EEG and video streams
    1. Connect the stimulus presentation computer to the EEG amplifier trigger input using a transistor-transistor logic (TTL) trigger cable or trigger box.
    2. Use the presentation software to send a TTL trigger pulse at the onset of each painting clip and a second TTL trigger pulse at the offset of each clip.
    3. Assign unique trigger codes to each event type. In this demonstration protocol, use 11–16 for clip onset and 21–26 for clip offset.
    4. Record the camera stream and EEG stream simultaneously from at least 5 s before the first trigger until at least 5 s after the final trigger. Log the trigger timestamp table automatically in the stimulus software.
    5. Validate synchronization after acquisition by matching trigger timestamps in the EEG recording with video frame indices from the presentation log. Visual checkpoint: Confirm that the mean alignment error between EEG trigger timestamps and video frame timestamps is within ±33 ms at 30 frames/s.
      ​NOTE: If no hardware frame-accurate synchronization interface is available, export timestamp logs from both systems and perform offline alignment correction using trigger onset matching.
  4. Record experimental data
    1. Instruct the participant to view the paintings naturally while minimizing large head movements, excessive blinking, speaking, and facial touching.
    2. Present the predefined painting clips in the selected order. Record synchronized EEG and facial video continuously throughout the viewing session.
    3. Pause the experiment and recheck electrodes if more than 5 channels exceed the impedance threshold during recording. Document interruptions, participant discomfort, and technical problems in the session log.

4. EEG preprocessing and feature extraction

  1. Importing raw EEG data
    1. Import the raw EEG recordings into the analysis environment. In this demonstration protocol, use either MATLAB R2023b with EEGLAB 2024.0 or Python 3.10 with MNE 1.6.
    2. Import the event markers and verify that all stimulus onset and offset triggers are present.
  2. Downsampling the EEG signals
    1. Downsample the signals from 500 Hz to 200 Hz and preprocess them according to the workflow shown in Figure 1.
    2. Apply anti-aliasing during resampling according to the software default or defined filter settings.
  3. Filtering the EEG signals
    1. Apply a band-pass filter from 0.5–45 Hz. Apply a notch filter at the local power frequency if visible line noise remains. In this demonstration protocol, apply a 50 Hz notch filter.
  4. Removing artifacts
    1. Inspect the continuous EEG manually and mark segments with gross movement artifacts.
    2. Run independent component analysis on the filtered data. Identify components associated with eye blinks, horizontal eye movements, jaw tension, or muscle activity by using scalp maps, time courses, and power spectra.
    3. Remove the identified artifact components and reconstruct the cleaned EEG signals. Reject segments that still contain excessive amplitude fluctuation after independent component analysis correction. In this demonstration protocol, reject segments exceeding ±100 µV.
  5. Re-referencing the data
    1. Re-reference the cleaned EEG signals to the common average reference.
  6. Segmenting the EEG data
    1. Epoch the continuous EEG according to stimulus onset triggers. Extract stimulus-aligned segments covering the full clip or fixed analysis windows. In this demonstration protocol, segment the EEG into 2.0 s windows with 50% overlap. Exclude windows with residual artifacts after segmentation.
  7. Computing differential entropy features
    1. Decompose each EEG segment into the following frequency bands: delta (1–4 Hz), theta (4–8 Hz), alpha (8–13 Hz), beta (13–30 Hz), and gamma (30–45 Hz).
    2. Calculate the differential entropy of each frequency band for each channel. Use the following equation for a Gaussian variable:
      DE = 1/2 ln(2πeσ2), where σ2 is the variance of the band-limited EEG signal.
    3. Arrange the channelwise differential entropy values into a two-dimensional topographic matrix or channel-feature matrix for model input.
    4. In this demonstration protocol, generate EEG feature maps with an input size of 128 × 128 x 5 per analysis window. Visual checkpoint: Plot representative differential entropy maps for each band and confirm that missing channels, constant-value maps, or abnormal bandwise collapse are not present.

5. Facial video preprocessing

  1. Extracting image frames
    1. Decode the recorded video into image frames using a scripted pipeline. Extract frames at 25 frames/s for model input while preserving synchronization metadata.
  2. Detecting and aligning the face
    1. Detect the face in each frame using a validated face detector. Locate facial landmarks and align the face based on the eye centers or standard landmark anchors. Discard frames in which the face is not detected reliably.
  3. Cropping and resizing the face region
    1. Crop the face to the detected bounding box with a small margin to preserve contour information. Resize the cropped face image to 224 x 224 pixels.
    2. Inspect representative cropped facial samples from the four target emotion categories, as shown in Figure 2, and confirm that the forehead, eyebrows, eyes, nose, cheeks, and mouth remain visible after cropping.
  4. Augmenting the training images
    1. Apply random horizontal flipping with probability 0.5 to the training set only. Apply random rotation within ±15° to the training set only.
    2. Do not apply augmentation to validation or test images.
  5. Normalizing the facial input
    1. Normalize pixel values to the range [-1, 1] or use the backbone-specific normalization rule consistently across all splits.

6. Construct the multimodal neural network

  1. Configuring the software and hardware environment
    1. Implement the model in Python 3.10 using PyTorch 2.1. Run training on Ubuntu 22.04 or another specified operating system.
    2. Use a workstation with at least 32 GB RAM, one graphics processing unit with at least 12 GB video memory, and sufficient local storage for synchronized EEG-video data. In this demonstration protocol, use an NVIDIA RTX 3080 graphics processing unit.
    3. Store the training script, model definition file, preprocessing script, and configuration file in a version-controlled project directory.
    4. Report the code repository or archived script package in the manuscript if sharing is permitted.
  2. Building the facial branch
    1. Construct the overall dual-branch multimodal framework as shown in Figure 3. Initialize a MobileNetV2 backbone for the facial branch.
    2. Modify the first convolution layer from 32 channels to 24 channels. Build the facial feature extraction branch according to Figure 4 and insert coordinate attention modules after the selected inverted residual blocks, as shown in Figure 5.
    3. Replace designated standard convolutions with parallel depth wise convolutions of 3 x 3 and 5 x 5 kernel size to improve multi-scale feature extraction.
    4. Export a fixed-dimensional facial feature vector for fusion. In this demonstration protocol, use a feature dimension of 256.
  3. Building the EEG branch
    1. Input the differential entropy feature maps into a convolutional neural network encoder. Use the convolutional layers to extract spatial patterns from the EEG feature maps.
    2. Feed the convolutional output sequence into a bidirectional long short-term memory module to capture temporal dependency across windows.
    3. Export a fixed-dimensional EEG feature vector. In this demonstration protocol, use a feature dimension of 256.
  4. Building the cross-modal fusion module
    1. Implement the multimodal feature interaction module shown in Figure 6. Implement a multi-head cross-attention block with 8 heads.
    2. Use the EEG feature sequence as the query and the facial feature sequence as the key and value in one cross-attention stream.
    3. Use the facial feature sequence as the query and the EEG feature sequence as the key and value in the reciprocal cross-attention stream.
    4. Concatenate the outputs from the two cross-attention streams. Pass the concatenated feature through a fusion projection layer.
  5. Adding the self-residual dilated convolution module
    1. Refine the fused representation with the self-residual dilated convolution connection layer shown in Figure 7.
    2. Construct a dilated convolution concatenation layer using dilation rates of 1, 3, 6, and 12. Concatenate the outputs from the four dilation branches.
    3. Add residual skip connections to stabilize gradient flow and preserve lower-level information.
  6. Adding the classifier
    1. Flatten or pool the fused representation. Add one fully connected layer and a final softmax output layer. Set the output classes to fear, happiness, calmness, and sadness.
  7. Defining the training settings
    1. Use the network and training settings summarized in Table 1. The four emotion classes (fear, happiness, calmness, and sadness) were operationally defined using a combined labeling procedure based on stimulus design and participant self-report. Each painting clip was preselected to elicit a target emotion category, and participants reported their dominant emotional experience after viewing. The final labels were assigned by aligning the intended stimulus category with self-reported responses, and inconsistent samples were excluded to ensure labeling reliability.

Access restricted. Please log in or start a trial to view this content.

Results

The performance of the proposed dual-branch cross-attention network was evaluated using the multimodal dataset collected from 327 participants during painting viewing. The primary outcome was four-class emotion classification performance for fear, happiness, calmness, and sadness. Additional analyses examined unimodal model behavior, multimodal convergence, sensitivity to the number of attention heads, comparative recognition performance, and the behavior of the facial and electroencephalography sub-networks. No separate...

Access restricted. Please log in or start a trial to view this content.

Discussion

This protocol addresses a practical problem in affective computing, namely, how to recognize emotional states in painting-viewing settings where visible expression and physiological response may not fully coincide at every moment. Under the experimental conditions reported here, the dual-branch framework achieved higher classification performance than the unimodal comparison models, which supports the value of combining electroencephalography and facial-expression information in exhibition-like environments. The findings...

Access restricted. Please log in or start a trial to view this content.

Disclosures

The authors declare no conflicts of interest.

Acknowledgements

We extend our sincere gratitude to the volunteers from Minzu Normal University of Xingyi and Southwest University for their participation in the experiments. We also thank the technical staff at the University of Girona for their assistance with EEG data acquisition and the anonymous reviewers for their insightful comments.

Access restricted. Please log in or start a trial to view this content.

Materials

List of materials used in this article
NameCompanyCatalog NumberComments
Deep Learning FrameworkPyTorchv2.0 / https://pytorch.orgUsed for model construction, training, and evaluation
EEG Acquisition System (64-channel)Brain Products GmbHactiCHamp Plus / v1.6Used for high-resolution EEG signal acquisition during experiments
EEG Electrode Cap (10–20 system)Brain Products GmbHActiCap / 64-channelStandard international 10–20 electrode placement
High-Definition CameraSony CorporationIMX327 SensorUsed for facial expression video recording (≥1080p, 30 fps)

References

  1. Yang, Y., et al. A dual-stream regional feature learning and adaptive fusion method for electroencephalogram-based emotion recognition. Eng Appl Artificial Intell. 164, 113250(2026).
  2. Li, C., Pun, S. H., Li, J. W., Chen, F. Eeg-based emotion recognition using graph attention network with dual-branch attention module. 2024 46th Ann Int Conf IEEE Eng Me Biol Soc (EMBC), , 1-4 (2024).
  3. Ubillús, J. A. T., et al. Algorithms used for facial emotion recognition: A systematic review of the literature. EAI Endorsed Transact Pervasive Health Technol. 9, 1-17 (2023).
  4. Xu, Y., Cheng, L. Face emotion recognition based on gabor wavelet and particle swarm optimization algorithm. Multimed Syst. 31 (6), 435(2025).
  5. Dube, D. Y., Sannasi, M. V., Kyritsis, M., Gulliver, S. R. Facial emotion recognition from feature loss media: Human versus machine learning algorithms. Comp Human Behav. 174, 108806(2026).
  6. Manuel, M. T. J., Medina-DeVilliers, S., Clarkson, T., Lerner, M. D., Riccardi, G. Evaluation of interpretability for deep learning algorithms in EEG emotion recognition: A case study in autism. Artif Intell Med. 143, 102545(2023).
  7. Watanabe, A., Yamazaki, T. Representation of the brain network by electroencephalograms during facial expressions. J Neurosci Meth. 357, 109158(2021).
  8. Prakash, A., Poulose, A. Electroencephalogram-based emotion recognition: A comparative analysis of supervised machine learning algorithms. Data Sci Manag. 8 (3), 342-360 (2025).
  9. Sun, Y., Ayaz, H., Akansu, A. N. Multimodal affective state assessment using fnirs + eeg and spontaneous facial expression. Brain Sci. 10 (2), 85(2020).
  10. Huang, Y., Yang, J., Liu, S., Pan, J. Combining facial expressions and electroencephalography to enhance emotion recognition. Future Internet. 11 (5), 105(2019).
  11. Han, Q., et al. A multi-branch electroencephalogram emotion recognition framework based on cross-subject or cross-session multi-perspective representation fusion. Neurocomputing. 658, 131767(2025).
  12. Wang, L., Chang, Y., Wang, K. Dual-branch multimodal fusion network for driver facial emotion recognition. Appl Sci. 14 (20), 9430(2024).
  13. Mutawa, A. M., Hassouneh, A. Multimodal real-time patient emotion recognition system using facial expressions and brain EEG signals based on machine learning and log-sync methods. Biomed Signal Proc Contr. 91, 105942(2024).
  14. Qi, Y., Ibrayim, M., Hamdulla, A. Attention-based dual-branch network for micro-expression recognition with global-local feature fusion. 2024 IEEE Int Joint Conf Biometr (IJCB), , 1-9 (2024).
  15. Li, E. Intervention of art education on college students’ aesthetic mood based on emotion recognition algorithm. Front Art Res. 6 (9), 81-90 (2024).
  16. Liu, Z., Han, M., Wu, B., Rehman, A. Speech emotion recognition based on convolutional neural network with attention-based bidirectional long short-term memory network and multi-task learning. Appl Acoustics. 202, 109178(2023).
  17. Xue, P., Wang, S., Bai, J., Qiang, Y. Research on bimodal emotion recognition algorithm based on multi-branch bidirectional multi-scale time perception. J Biome Eng. 42 (3), 528-536 (2025).
  18. Shioiri, S., Nagata, H., Sato, Y., Hatori, Y. Prediction of preference judgments of face images using facial expressions and eeg signals. J Vis. (9), 5063(2023).
  19. Lu, Y., Chen, J. Cross-subject EEG emotion recognition using the SSA-EMS algorithm for feature extraction. Entropy. 27 (9), 986(2025).
  20. Thapa, D., Rai, R. Freq-eer: A novel frequency-driven ensemble framework for emotion recognition and classification of eeg signals. Appl Sci. 15 (19), 10671(2025).
  21. Yang, Y., et al. Investigating of deaf emotion cognition pattern by eeg and facial expression combination. IEEE J Biomed Health Inform. 26 (2), 589-599 (2022).
  22. Tong, L., Yang, L., Wang, X., Liu, L. Self-aware face emotion accelerated recognition algorithm: A novel neural network acceleration algorithm of emotion recognition for international students. PeerJ Comput Sci. 9, e1611(2023).

Access restricted. Please log in or start a trial to view this content.

Reprints and Permissions

Tags

EEG Emotion RecognitionFacial Expression AnalysisDifferential EntropyConvolutional Neural NetworkBidirectional LSTMAffective ComputingHuman Computer Interaction