Research Article

Construction of a Music Genre Preference Recognition Model Based on Deep Learning

DOI:

10.3791/70514

May 26th, 2026

In This Article

Summary

Loading...
$$\rightleftharpoonup{xx}$$ $$\longleftharp{xx}$$, $$\longrightharp{xx}$$,

This study built a deep learning model using EEG signals and music familiarity to predict music genre preference. Data were collected with a low‑cost Muse S device and analyzed using a CNN+RNN and an EEGNet model. Familiarity significantly improved accuracy, reaching up to 99%, showing its substantial value in preference prediction.

Abstract

Loading...
$$\rightleftharpoonup{xx}$$ $$\longleftharp{xx}$$, $$\longrightharp{xx}$$,

With the increasing prevalence of mental health issues, music therapy has gained attention as a non-pharmacological intervention, and deep learning techniques have shown promise in music emotion recognition and preference prediction. This study constructed a deep neural network model (CNN+RNN/EEGNet) to efficiently identify music type preferences and examine the influence of user familiarity on prediction accuracy. EEG signals were collected using a four-channel Muse S wearable device, and user familiarity scores were used as input features. The study followed a four-stage workflow: preparation, experimental design, model construction, and result analysis. In the experimental design, music was categorized into rock, ballad, and folk, and EEG data and familiarity ratings were collected for each category. Data was trained and tested using CNN+RNN or EEGNet models, and model performance was evaluated via subject-level 10-fold cross-validation. Results indicated that predicting all music types with EEG data alone achieved an accuracy of 82.28 ± 3.42%. For individual music types, accuracies were 91.13 ± 3.60% (rock), 91.83 ± 2.07% (ballad), and 87.87 ± 4.76% (folk). When incorporating user familiarity as a feature and using a multi-level rating output, overall prediction accuracy increased to 94.94 ± 1.61%, while individual music type accuracies reached 99.15 ± 1.56% (rock), 98.51 ± 2.30% (ballad), and 98.21 ± 2.60% (folk). These results demonstrate that combining familiarity features with a multi-level scoring system significantly improves the prediction of music preferences. By using an affordable, wearable Muse S EEG device and leveraging user familiarity, this study successfully developed a highly effective deep neural network model (CNN+RNN/EEGNet) for recognizing music type preferences. The findings indicate that both overall and individual music-type predictions benefit from the inclusion of familiarity information, highlighting the potential of this approach for personalized music recommendations and music therapy applications.

Introduction

Loading...
$$\rightleftharpoonup{xx}$$ $$\longleftharp{xx}$$, $$\longrightharp{xx}$$,

Music plays a significant role in human life, and the widespread use of online platforms has made listening to music a routine daily activity. Beyond entertainment, music functions as a therapeutic tool that promotes relaxation, reduces pain, and supports individuals with developmental conditions. Musical preference directly influences both therapeutic outcomes and emotional responses. Previous studies show that preferred music enhances physical performance, such as endurance and sprint capacity1, and reduces stress and pain2. It is also an important intervention for older adults, in which individual preferences influence treatment effectiveness3,4. Preferred music has also been shown to lower heart and respiratory rates, promoting relaxation. In addition, musical preference improves emotion prediction accuracy5, strengthens emotional responses6, and enhances machine learning model performance, increasing prediction accuracy from 53.18% to 71.09%. Repeated exposure further increases preference regardless of initial bias7.

Despite these findings, existing studies on predicting music preferences remain limited by small sample sizes, narrow age ranges, and insufficient consideration of user familiarity. These limitations reduce model generalizability and practical applicability. Incorporating user-specific factors such as familiarity remains an underexplored approach for improving prediction performance.

Music emotion recognition (EMR) has become an active area of research, and its accuracy is influenced by music preference and familiarity8,9,10,11. Participants provide more consistent emotional evaluations when exposed to familiar and preferred music, and increasing familiarity improves recognition accuracy in both experimental and educational settings12,13,14,15,16,17. These findings highlight familiarity as a critical factor in both emotion recognition and preference prediction.

Deep learning, based on artificial neural networks, enables the automatic extraction of high-dimensional features and demonstrates stronger pattern recognition than shallow models18,19,20. Deep learning techniques have also been widely applied to music-related tasks, including genre classification, signal recognition, and personalized recommendation systems21,22,23,24,25,26. Electroencephalography (EEG) is widely used to measure brain electrical activity across frequency bands (Table 1) for cognitive and emotional analysis27,28. The integration of EEG with deep learning provides an effective framework for capturing neural responses to music preferences.

Traditional EEG systems often require gel-based electrodes and high-cost equipment, limiting their applicability in real-world settings. Wearable EEG devices with dry electrodes offer a practical alternative, providing reduced cost, improved usability, and greater flexibility. The Muse S device records signals from key brain regions (TP9, TP10, AF7, AF8) at 256 Hz, enabling reliable, scalable data acquisition for both controlled experiments and real-world applications.

This study addresses the identified limitations by combining EEG signals with user-familiarity features to predict music preferences using deep learning models. CNN+RNN and EEGNet architectures are used to extract both time–frequency features and temporal dependencies, enabling improved representation of EEG signals compared to traditional approaches. The study tests two hypotheses: (1) user music familiarity improves EEG-based preference prediction accuracy; (2) genre-specific models (rock, lyrical, and folk) outperform models trained on mixed categories.

A cohort of 30 healthy adults with balanced gender distribution was included. Although the limited sample size and use of four-channel EEG restrict generalizability and signal resolution, integrating familiarity features with deep learning models enables robust prediction performance. The objectives are to classify music genres and perform individual predictions for each category, and to incorporate user familiarity as a feature to enhance prediction accuracy.

Access restricted. Please log in or start a trial to view this content.

Protocol

Loading...
$$\rightleftharpoonup{xx}$$ $$\longleftharp{xx}$$, $$\longrightharp{xx}$$,

This study was conducted in accordance with the guidelines of the Ethics Committee of Quzhou University. The research protocol was approved by the Quzhou University Institutional Review Board (IRB) under approval number QZU-IRB-2026-037. All participants provided informed consent prior to their inclusion in the study.

Experimental preparation and participant recruitment
The experiment consisted of four stages (Figure 1): preparation, experimental design, model construction, and result analysis. In the preparation stage, research objectives were defined based on a review of literature on music therapy, EEG data acquisition, and deep learning models. In the experimental design stage, 30 healthy participants were recruited to perform music listening tasks. EEG signals and music familiarity ratings were recorded. EEG data were standardized and segmented into windows of 2048 samples with 1024 sample overlap. Familiarity features were combined with EEG data as model inputs. Two models, CNN+RNN and EEGNet, were constructed. CNN layers extracted time–frequency features, and bidirectional LSTM captured temporal dependencies. A fully connected layer with a softmax activation function was used for classification. Models were trained with a batch size of 50 for 30 epochs using the Adam optimizer. Dropout and early stopping were applied. Performance was primarily evaluated using Accuracy as the reported metric, with Precision, Recall, and F1-score metrics also calculated for additional evaluation.

Experimental environment and equipment set-up
The experiment was conducted in a quiet, enclosed room free from external interference, with the room temperature maintained at 27 oC. The following equipment was used: an EEG device (Muse S, InteraXon Inc.), a four-channel EEG system (TP9, AF7, AF8, TP10) with a sampling rate of 256 Hz; Final E3000 in-ear headphones; and a data acquisition computer running the Muse SDK and Lab Streaming Layer (LSL).

Before each experimental session, the ear-contact surfaces of the headphones were cleaned with alcohol to ensure hygiene and reduce the risk of cross-contamination. The Muse S headset was fitted to each participant, ensuring stable contact between all four dry electrodes (TP9, AF7, AF8, TP10) and the skin. If the signal quality was poor, the headband position was adjusted until stable signals were obtained. Participants were seated in a comfortable chair with back support and instructed to keep their eyes closed, avoid speaking, and minimize large body movements during the experiment.

Music stimuli preparation
Three music genres were prepared: rock, lyrical (ballad), and folk music. Each genre included three Chinese songs, for a total of nine songs. To control differences in music familiarity, songs in each genre were selected according to the following criteria: one old song; one current popular song (from the Spotify Top Songs chart); and one new song (from the Spotify New Releases chart).

Each song was edited into a standardized audio segment, starting from the introduction, followed by the verse, and ending at the first chorus, with a total duration of 90–130 s. A unique identifier was assigned to each song (e.g., R1–R3, B1–B3, F1–F3) for subsequent data labeling and file management.

Experimental procedure and questionnaire recording
Before the formal experiment, a volume test was conducted by playing a test audio clip. Participants adjusted the volume to a level that was clear, comfortable, and not excessively loud, and kept it constant throughout the experiment. The playback order of the nine songs was randomized (e.g., using a random number table or software randomization) to reduce order effects. All participants completed the music-listening tasks independently. Each music segment was presented from the introduction through the first chorus, lasting 90–130 s. A 30-second rest period was provided between songs. During this interval, participants completed a familiarity questionnaire using an online platform (Questionnaire Star). Participants rated their familiarity with each song using a five-point Likert scale (1 = completely unfamiliar; 5 = very familiar). The participant’s preference label for each song was recorded (0 = dislike; 1 = like) according to the predefined binary classification scheme. The complete experimental procedure is illustrated in Figure 2.

EEG acquisition and data logging
EEG data were acquired using the official Muse SDK and streamed via the Lab Streaming Layer (LSL). The EEG data corresponding to each song were saved as Comma-Separated Values (CSV) files, containing at minimum the following fields: timestamp (milliseconds); raw EEG signals from four channels (TP9, AF7, AF8, TP10); manually recorded familiarity score (1–5); and preference label (0/1). A consistent file-naming convention was used (e.g., S01_R1.csv, indicating participant 01 and rock song 1) to facilitate automated processing and traceability.

Data preprocessing, normalization, and windowing
EEG signals were recorded using a Muse S four-channel EEG device with a sampling rate of 256 Hz, and custom programs were developed using the official development kit. Raw EEG data were acquired through the Lab Streaming Layer (LSL) and saved in CSV format. The CSV files included the following columns: timestamp (milliseconds), four-channel EEG values (TP9, AF7, AF8, TP10), user music familiarity ratings (1–5), and user music preference ratings.

The total number of data entries corresponds to multiple windows generated for each participant and each music piece, resulting in a large-scale dataset depending on the windowing scheme. In the original experiment, music preference was represented as a binary label (0 = dislike, 1 = like), which may lead to artificially high accuracy and limited model generalization.

To improve the practical significance of the prediction task, the preference labels were converted into a multi-level rating scale: 0 = Strongly Disagree; 1 = Disagree; 2 = Neutral; 3 = Agree; 4 = Strongly Agree.

Prior to machine learning, the data were standardized and segmented into windows of 2048 samples each, with 1024-samples overlapping between adjacent windows to preserve temporal continuity. Additionally, data augmentation was applied to increase the number of training samples and reduce the risk of overfitting.

Model architecture
The research model consisted of three convolutional layers, three pooling layers, and a fully connected layer. Max pooling was used for feature selection, dropout was applied to reduce overfitting, and the final output was generated using a softmax classification layer. Each input window included four-channel EEG signals (2048 samples per window) and a familiarity feature. The output represented a five-class preference rating. The architecture is shown in Figure 3, and model parameters are outlined in Table 2.

Model training and cross-validation
The dataset was randomly split into 80% for training and 20% for testing. Ten-fold cross-validation was applied, with one subset used for validation and the remaining nine for training in each iteration. This process was repeated ten times, and the results were aggregated. Training hyperparameters were set as follows: Epochs = 30; Batch size = 50; Optimizer = Adam; Loss function = categorical crossentropy; Evaluation metric = Accuracy. Model performance was recorded under EEG-only and EEG + familiarity conditions, both for all genres combined and for individual genres.

Statistical analysis methods
Data analysis was conducted using Python 3.10 with TensorFlow 2.12, Keras 2.12, and MNE-Python 1.3. Data processing and analysis were performed using Pandas 2.1 and NumPy 1.26. Performance metrics from the 10-fold cross-validation were aggregated, and mean and standard deviation values were calculated. Statistical comparisons between conditions were performed using paired t-tests or Wilcoxon signed-rank tests, with a significance level of P < 0.05.

Access restricted. Please log in or start a trial to view this content.

Results

Loading...
$$\rightleftharpoonup{xx}$$ $$\longleftharp{xx}$$, $$\longrightharp{xx}$$,

Model performance was evaluated under two conditions: (1) EEG features only and (2) EEG features combined with music familiarity. Classification was conducted for both combined music categories and individual music categories (rock, lyrical, and folk).

Table 3 presents results obtained using EEG features only, while Table 4 presents results obtained using EEG features combined with familiarity. Both CNN+RNN and EEGNet models were evaluated under these conditions, and EEGNet wa...

Access restricted. Please log in or start a trial to view this content.

Discussion

Loading...
$$\rightleftharpoonup{xx}$$ $$\longleftharp{xx}$$, $$\longrightharp{xx}$$,

Previous studies on EEG-based music preference prediction have largely relied on multi-channel systems, limiting practical application21,22,23,24,25,26. In this study, a four-channel wearable EEG device (Muse S) achieved comparable performance when combined with deep learning models and familiarity features, demonstrating a m...

Access restricted. Please log in or start a trial to view this content.

Disclosures

Loading...
$$\rightleftharpoonup{xx}$$ $$\longleftharp{xx}$$, $$\longrightharp{xx}$$,

The author has no conflicts of interest to disclose.

Acknowledgements

Loading...
$$\rightleftharpoonup{xx}$$ $$\longleftharp{xx}$$, $$\longrightharp{xx}$$,

The author gratefully acknowledges the support and facilities provided by the research institution during the course of this study. Special thanks are extended to the laboratory staff for their assistance with technical procedures.

Access restricted. Please log in or start a trial to view this content.

Materials

List of materials used in this article
NameCompanyCatalog NumberComments
Alcohol wipesAny standard supplierN/AUsed for cleaning headphones between participants
Data acquisition computerAny standard manufacturerN/ARuns Muse SDK and Lab Streaming Layer (LSL)
EEG headset (Muse S)InteraXon Inc.https://choosemuse.com/muse-s/Four-channel wearable EEG device (TP9, AF7, AF8, TP10)
Final E3000 in-ear headphonesFinal Audio Designhttps://snext-final.com/en/products/detail/E3000Used for audio playback during experiment
Lab Streaming Layer (LSL) softwareOpen-sourcehttps://github.com/sccn/labstreaminglayerUsed for EEG data streaming
Muse SDKInteraXon Inc.https://developer.choosemuse.com/Used for EEG data acquisition
Questionnaire Star (online platform)Changsha Ranxing Information Technology Co., Ltd.https://www.wjx.cn/Used for familiarity data collection
Python 3.10Python Software Foundationhttps://www.python.org/downloads/Used for data processing and analysis
TensorFlow 2.12Googlehttps://www.tensorflow.org/Deep learning framework
Keras 2.12Googlehttps://keras.io/Neural network API
MNE-Python 1.3Open-sourcehttps://mne.tools/stable/index.htmlEEG data analysis
NumPy 1.26Open-sourcehttps://numpy.org/Numerical computation
Pandas 2.1Open-sourcehttps://pandas.pydata.org/Data manipulation

References

Loading...
$$\rightleftharpoonup{xx}$$ $$\longleftharp{xx}$$, $$\longrightharp{xx}$$,
  1. Ballmann, C. G., McCullum, M. J., Rogers, R. R., Marshall, M. R., Williams, T. D. Effects of preferred vs. nonpreferred music on resistance exercise performance. J. Strength Cond. Res. 35 (6), 1650-1655 (2021).
  2. Banks, D. Neurotechnology: microelectronics. Handbook of Neuroprosthetic Methods. Finn, W. E., LoPresti, P. G. , CRC Press. Boca Raton, FL. 361-400 (2002).
  3. Aldridge, D. An overview of music therapy research. Complement. Ther. Med. 2, 204-216 (1994).
  4. Hillecke, T., Nickel, A., Bolay, H. V. Scientific perspectives on music therapy. Ann. N. Y. Acad. Sci. 1060, 271-282 (2005).
  5. Lai, H. L. Music preference and relaxation in Taiwanese elderly people. Geriatr. Nurs. 25 (5), 286-291 (2004).
  6. Kreutz, G., Lotze, M. Neuroscience of music and emotion. Neurosciences in Music Pedagogy. Rauscher, F., Gruhn, W. , Nova Science Publishers. New York. 143-167 (2007).
  7. Madison, G., Schiölde, G. Repeated listening increases the liking for music regardless of its complexity: implications for the appreciation and aesthetics of music. Front. Neurosci. 11, 147(2017).
  8. Pereira, C. S., et al. Music and emotions in the brain: familiarity matters. PLoS One. 6 (11), e27241(2011).
  9. Hamlen, K. R., Shuell, T. J. The effects of familiarity and audiovisual stimuli on preference for classical music. Bull. Counc. Res. Music Educ. 168, 21-34 (2006).
  10. Liu, W. J. Unit learning in junior high school music class based on deep understanding. Prim. Second. Sch. Music Educ. 3, 5(2022).
  11. A vocal range balancing method, device, and system based on deep learning. China Patent. , CN109147807B (2024).
  12. Yin, D. H. Constructing a music classroom teaching paradigm from the perspective of deep learning (Part 2). Prim. Second. Sch. Music Educ. 6, 4(2022).
  13. Zhang, C., Shi, T., Ai, J., Tian, W. Construction of GUI elements recognition model for AI testing based on deep learning. Proceedings of the 2021 8th International Conference on Dependable Systems and Their Applications (DSA), , 508-515 (2021).
  14. Shi, Y., Ko, Y. C. Construction of English pronunciation judgment and detection model based on deep learning neural networks data stream fusion. Int. J. Pattern Recognit. Artif. Intell. 36 (6), 2252011(2022).
  15. Li, X. H., et al. Survey of remote sensing image registration based on deep learning. Natl. Remote Sens. Bull. 27 (2), 267-284 (2023).
  16. Bishop, C. M., Nasrabadi, N. M. Pattern recognition and machine learning. , Springer. New York. (2006).
  17. Goodfellow, I., Bengio, Y., Courville, A. Deep learning. , MIT Press. Cambridge, MA. (2016).
  18. LeCun, Y., Bengio, Y., Hinton, G. Deep learning. Nature. 521, 436-444 (2015).
  19. Janiesch, C., Zschech, P., Heinrich, K. Machine learning and deep learning. Electron. Mark. 31 (3), 685-695 (2021).
  20. A deep learning-based system and method for Gongche notation character recognition. China Patent. , CN202210650096.0 (2024).
  21. Sun, L. Y. Research on music genre classification and generation technology based on deep learning. , Nanjing Audit University. Nanjing, China. Master's thesis (2022).
  22. Gan, Q. Deep learning-based electronic music signal recognition method. Inf. Comput. 35, 54-56 (2023).
  23. Zhang, H. Building a music teaching and research community using deep learning. Music World. 1, 16-19 (2022).
  24. Liu, L., Kong, M., Cao, C., Shu, Z., Liu, K., et al. Personalized music recommendation algorithm based on machine learning. Multimed. Syst. 31, 166(2025).
  25. Naser, D. S., Saha, G. Influence of music liking on EEG-based emotion recognition. Biomed. Signal Process. Control. 64, 102251(2021).
  26. Cannard, C., Brandmeyer, T., Wahbeh, H., Delorme, A. Self-health monitoring and wearable neurotechnologies. Handbook of Clinical Neurology. Ramsey, N. F., Millán, J. delR. 168, 207-232 (2020).
  27. Kumari, P., Vaish, A. Brainwave-based user identification system: a pilot study in robotics environment. Robot. Auton. Syst. 65, 15-23 (2015).
  28. Asif, A., Majid, M., Anwar, S. M. Human stress classification using EEG signals in response to music tracks. Comput. Biol. Med. 107, 182-196 (2019).

Access restricted. Please log in or start a trial to view this content.

Reprints and Permissions

Request permission to reuse the text or figures of this JoVE article

Request Permission

Tags

Music Genre RecognitionDeep Learning ModelEEG Signal AnalysisMusic Preference PredictionCNN RNN ModelEEGNet ArchitectureUser Familiarity FeatureMuse S DeviceMusic Therapy ApplicationPersonalized Music Recommendation

Related Articles