Thyroid dysfunction (TD) is one of the most common chronic endocrine illnesses, with varying prevalence across populations1. The two most common thyroid gland disorders are hypothyroidism and hyperthyroidism. Hypothyroidism is a prevalent condition that results from insufficient thyroid hormone production. While it is generally manageable with treatment, in severe instances, it can be fatal if not addressed. The most common symptoms of hypothyroidism in adults include fatigue, constipation, lethargy, weight gain, dry skin, and voice changes. However, clinical presentations may vary by age, gender, and other factors3. It is reported that the disease affects 5% of the general population, and an additional 5% remain undiagnosed, while 99% of these patients suffer from primary hypothyroidism. It is widespread in the Arabian Gulf States, although it often remains underdiagnosed4. Congenital hypothyroidism affects approximately one in every 3,450 babies in Saudi Arabia. Other cross-sectional studies that focused solely on women reported higher prevalence rates ranging from 13% to 35%5. Further, proving that females are more susceptible than males6. Hypothyroidism, if left untreated, can lead to dyslipidemia, cognitive impairment, hypertension, infertility, and neuromuscular dysfunction7. Consequently, this study aims to employ machine learning methods to proactively diagnose hypothyroidism using a dataset from King Fahad Specialist Hospital (KFUH) in Dammam, Saudi Arabia.
Hypothyroidism can be difficult to diagnose because the symptoms are often mistaken for those of other illnesses8. Moreover, a late diagnosis of hypothyroidism can negatively affect the intellectual abilities of a patient9. The thyroid stimulating hormone (TSH) test is the most used diagnostic method for detecting hypothyroidism9. However, the arrival of artificial intelligence and machine learning approaches, strategies, and tools has helped resolve diagnostic and prognostic problems across various medical areas. ML is used to assess the significance of clinical factors and their combinations for prognosis, such as predicting illness progression, extracting medical knowledge for outcome research, therapeutic planning and support, and overall patient care10. With large volumes of medical data available, medical practitioners now have a valuable resource for discovering new and potentially valuable knowledge through machine learning11,12.
Using a simple clinical dataset from the Saudi dataset, this study aims to assist health workers in the country in diagnosing hypothyroidism early and to implement this predictive system in hospitals with limited computer resources and equipment. Additionally, prior studies in this field did not utilize the Saudi dataset. This presents a chance to incorporate the Saudi dataset alongside earlier research. The studies employed a Saudi dataset to predict kidney disease, diabetes, thyroid disease, and Alzheimer's disease, respectively, achieving high accuracy13,14,15. This experience motivates applying these methods to other diseases in Saudi Arabia. The studies are among the latest works in machine learning and artificial intelligence in the medical field16,17,18,19.
The proposed study discusses deploying machine learning techniques on a simple clinical Saudi dataset for hypothyroidism to predict the presence of the disease before it becomes fully symptomatic. The aim is to develop a machine learning based system that provides early warnings of a potential occurrence of hypothyroidism. It will also make it easier for local hospitals with limited equipment to use this diagnostic system. This study is the first to examine this disease using the Saudi dataset, focusing on the pre-symptomatic stage. The dataset covers a broad patient population for local Saudi hospital patients, with a wide age range, a mix of positive and negative cases of hyperthyroidism.
This study utilized four machine learning algorithms: support vector machine (SVM), K-nearest neighbors (KNN), eXtreme gradient boosting (XGBoost), and a soft voting ensemble classifier composed of SVM and KNN. Each of the four algorithms has proved its capabilities in the medical field. The KNN algorithm was selected for its simplicity and for achieving a test accuracy of 77.5% in the preemptive diagnosis of diabetes mellitus (DM)20. At the same time, SVM was chosen for its robustness21, and it has recently been used to predict mortality risk in COVID-19 patients22. Moreover, XGBoost was selected for its strong performance23, achieving an accurate predictive score of 94% for hypertension outcomes, and stacking was employed to enhance the results from the top-performing algorithms. In addition, a voting ensemble classifier has been utilized to identify thyroid disorders with 100% accuracy24.
The primary objective of the current study is early diagnosis of hypothyroidism in a Saudi population using an original routine clinical data and state-of-the-art machine learning algorithms. The four techniques used in this study were implemented on the Jupyter platform using the Python programming language, using the hypothyroidism dataset. 10 fold cross validation with hyperparameter tuning was used to improve performance. Sequential feature selection with a forward selection technique was then implemented, resulting in the voting ensemble classifier achieving the highest accuracy of 94.7%. SVM, KNN, and XGBoost, respectively, followed it. Regarding recall, the soft voting ensemble classifier achieved the highest score at 95.5%. Overall, the optimal model for preemptive hypothyroidism diagnosis was the soft voting ensemble classifier, which achieved the highest score across all performance measures evaluated in this study.
Several studies have been conducted on diagnosing general thyroid disorders and on implementing technology, especially machine learning, as a computer-aided solution. LSTM-CNN, AlexNet, the adjusted GoogLeNet, IndRNN–CNN, and the adjusted ResNet deep learning models were deployed using Raman spectroscopy to distinguish serum samples from 32 hypothyroidism patients, 38 hyperthyroidism patients, and 29 control subjects. Their model obtained accuracy of 84%, 91%, 75%, 82%, and 71%, respectively25.
Machine learning algorithms like SVM, KNN, logistic regression, and random forest have been utilized for predicting the incidence of radiation-induced hypothyroidism in nasopharyngeal carcinoma patients26. Their model achieved an AUC value of 0.7. Moreover, the authors employed various machine learning algorithms to predict the treatment trend of patients with hypothyroidism undergoing LT4 treatment. Among 10 classifiers, the extra-tree classifier achieved the highest accuracy of 84%.
In another study, the authors created and tested a machine-learning strategy for identifying people with hyperthyroidism and hypothyroidism who would benefit from immediate medical treatment27. They deployed gradient boosting decision tree (GBDT), SVM, ANN, and logistic regression algorithms. Their models obtained an AUROC of 93.8% and 90.9% for hyperthyroidism and hypothyroidism, respectively. In another study, the authors used machine learning to detect thyroid disorders, which comprise hypothyroidism and hyperthyroidism. Their models included SVM, ANN, KNN, logistic regression, and decision tree, with the highest accuracy of 96.92% achieved by Logistic Regression28. Recently, a group of researchers applied different methods, including SVM, decision tree, Naïve Bayes, and ensemble, to predict and detect hypothyroidism. The decision tree achieved the highest accuracy, with 97.6%29. Another study applied the hybrid differential evolution kernel-based Naïve Bayes algorithm to reduce the attributes of a thyroid dataset of hyperthyroidism and hypothyroidism diseases from 21 to 10. They later applied the kernel-based Naïve Bayes classifier to the feature subset, achieving 97.97% accuracy30. Likewise, in another study, the authors used several machine-learning algorithms to predict the risk of thyroid disease on the sick-euthyroid dataset from California University. Their experiment demonstrated that ANN surpassed all other methods, achieving an accuracy of 95%31. Furthermore, Jha et al. used proposed feature augmentation techniques based on deep neural networks to predict thyroid diseases, achieving an outstanding accuracy of 99.95%32.
According to prior literature, a homogeneous voting ensemble with multiple attribute selection was used to identify thyroid diseases, employing the following algorithms: random forest, gradient boosting, decision tree, and logistic regression, obtaining 100% accuracy24. Furthermore, the authors applied the extreme learning machine model to the hypothyroidism open-access dataset on Kaggle, achieving 92% accuracy. The dataset contained 3772 samples: 3481 hypothyroid and the rest negative33. Recently, in a study, the authors utilized bagged CART to classify hypothyroidism. Their proposed model attained an impressive accuracy of 99.9%. Additionally, T4U, TSH, and TBG were identified as the most significant factors associated with hypothyroidism34. In another study, the authors proposed and compared different machine learning and deep learning models, in which decision tree and random forest achieved an accuracy of 99.5% and 99.3%, respectively. They used a Weka dataset containing 30 attributes and 3772 samples35.
This brief literature review demonstrates that, while most methods employed achieved very good to excellent results, none have utilized a straightforward clinical dataset from Saudi Arabia. It was also noted that most publications under consideration used imaging datasets, which raised the work required for data collection and increased the difficulty of applying very complicated models created by non-technical people. Furthermore, most of these works were done on general thyroid disorders, with very few publications emphasizing hypothyroidism. This situation provides an opportunity to investigate basic machine learning models on a simple clinical dataset from the Saudi dataset to assist local hospitals with limited computerized equipment in pre-emptively diagnosing the disease.
The brief introduction to the algorithms investigated in the current study is given in Supplementary File 1. The four algorithms are K-Nearest Neighbor (KNN)36,37,38,39,40, Support Vector Machine (SVM)41,42,43,44,45, Extreme Gradient Boosting (XGBoost)46,47,48, and voting ensemble49,50.