Here, we present a protocol for establishing a non-invasive XGBoost-based AI model for diagnosing vocal cord polyps using the Saarbrücken database.
Method Article
Here, we present a protocol for establishing a non-invasive XGBoost-based AI model for diagnosing vocal cord polyps using the Saarbrücken database.
With the continuous growth of human social communication, the number of people suffering from voice disorders is also increasing. Due to the objective and non-invasive advantages of acoustic detection methods for pathological voice, the use of speech signal analysis for pathological voice recognition has become a research hotspot. This article first selected 101 continuous vowels /a/ from the German SVD database as the research object. Secondly, using wavelet packet technology for time-frequency analysis, four nonlinear dynamic parameters, namely approximate entropy, sample entropy, fuzzy entropy, and permutation entropy, are extracted from the sub signals as the feature parameter set for the pathological voice classifier. Finally, the machine learning algorithm XGBoost is selected as the pattern recognition method to establish a pathological voice classifier, and the classification performance is verified using five fold cross validation and ROC curve. Experimental results have shown that the accuracy of XGBoost's pathological voice classifier is 0.857, the F1 score is 0.875, and the AUC value is 0.944, all of which are higher than the classifier constructed by SVM, indicating that XGBoost has better performance in pathological voice recognition.
The number of people suffering from voice disorders is constantly increasing. According to statistics, one-third of the population experiences voice problems at some point in their lives1. For professional voice users such as singers and teachers, due to their frequent misuse and abuse of their voice, the incidence of throat diseases is higher. Two studies conducted surveys on teachers in Tianjin and Urumqi, and the results showed that the probability of suffering from voice disorders is 33.81% and 28.23%, respectively2,3. As a result, there is increasing emphasis on the detection and diagnosis of pathological voice in clinical practice. At present, subjective evaluation methods, laryngoscopy examination, and acoustic detection methods are mainly used for the diagnosis of voice diseases in clinical practice4,5. Acoustic detection method converts voice signals into digital signals for research, ensuring objectivity and avoiding patient pain during the examination6. Therefore, acoustic detection has become an effective auxiliary tool for doctors in clinical diagnosis, and research on it is increasing.
The most important techniques for diagnosing pathological voice using acoustic analysis methods are feature extraction and pattern recognition. Since the 1960s, researchers have been extracting some traditional acoustic parameters for the study of pathological voice, and have found many effective and robust acoustic feature parameters, such as fundamental frequency perturbation, amplitude perturbation, and Mel-frequency cepstral coefficients (MFCC)7. In recent years, researchers have not only limited themselves to studying traditional acoustic characteristic parameters, but nonlinear dynamic parameters have also begun to be used in the field of pathological voice recognition8. It also has a very good effect. With the rapid development of machine learning, more pattern recognition methods with fast running speed and good classification performance have emerged. There are also more and more pattern recognition methods used in pathological voice recognition, such as support vector machines (SVM), random forest, neural networks, deep learning9,10. XGBoost is a pattern recognition method that has gained attention in recent years11. It does not require feature normalization and can select features on its own. It can adapt to various loss functions and uses regularization terms to prevent overfitting. It has the characteristics of fast speed, portability, and fault tolerance. Therefore, this article attempts to apply the XGBoost method to pathological voice recognition in order to achieve better recognition results.
Access restricted. Please log in or start a trial to view this content.
The technical workflow of this study is illustrated in Figure 1.
1. Acquisition of voice samples
2. Wavelet packet decomposition of voice data13
3. Feature parameter extraction
(1)
(2)
(4)
(5)
(7)
(8)
(9)
(10)
(11)4. Model establishment
5. Evaluation of the model performance
Access restricted. Please log in or start a trial to view this content.
In this study, wavelet packet analysis was employed to perform time-frequency decomposition of voice signals. By integrating nonlinear dynamic parameters—including approximate entropy, sample entropy, fuzzy entropy, and permutation entropy—the complexity and irregularity characteristics of pathological voices associated with vocal fold polyps were systematically characterized. Two pathological voice classifiers were subsequently constructed based on the support vector machine (SVM) algorithm and the XGBoost algorithm, re...
Access restricted. Please log in or start a trial to view this content.
Pathological voice diagnosis utilizing speech signal analysis represents a non-invasive and objective approach19. Consequently, pathological voice classification has emerged as a significant research focus in speech signal processing and recognition, demonstrating substantial clinical application value and developmental prospects20. This study employed speech samples collected from SVD as experimental corpora. Through speech signal processing and machine learning techniques...
Access restricted. Please log in or start a trial to view this content.
The authors declare that they have no competing interests.
The work was supported by the Jiangsu Province Hospital Capability Improvement Project (JSPH-MC-2023-12). The authors would like to thank Associate Professor Shan Li of the School of Economics and Management, and the staff of the Key Laboratory of Brain-Machine Intelligence Technology (Ministry of Education) at Nanjing University of Aeronautics and Astronautics, for their guidance on methodology, data analysis, and figure generation. These contributions have ensured the smooth progress of this research.
Access restricted. Please log in or start a trial to view this content.
| Name | Company | Catalog Number | Comments |
|---|---|---|---|
| MATLAB | The MathWorks | R2023a | Primary software platform for numerical computation and algorithm prototyping. |
| Python software | Python Software Foundation | Python 3.10 | Core programming language used throughout the study. |
| Saarbruecken Voice Database,SVD | Institute of Phonetics, Saarland University | Saarbrücken Voice Database (SVD) | Publicly available database of pathological and normal voice recordings. Contains sustained vowels and continuous speech from subjects with conditions including vocal cord polyps, as well as healthy controls. Used for academic research under its specified terms of access. |
Access restricted. Please log in or start a trial to view this content.
Request permission to reuse the text or figures of this JoVE article
Request Permission