Artificial intelligence (AI) has become an important component of modern healthcare by supporting disease diagnosis, prognosis, risk stratification, medical image analysis, and clinical decision-making1. Advances in machine learning and deep learning have enabled healthcare systems to process large volumes of structured and unstructured data, often achieving predictive performance comparable to or exceeding conventional statistical approaches2. Despite these advances, many AI models operate as complex black-box systems, making it difficult for clinicians and healthcare stakeholders to understand how predictions are generated3. This lack of transparency can limit trust, adoption, and accountability in high-stakes healthcare environments4,5.
Explainable artificial intelligence (XAI) has emerged as a promising approach for improving the transparency and interpretability of machine learning models6. Widely used explanation techniques, including SHAP (Shapley additive explanations), LIME (Local interpretable model-agnostic explanations), and Gradient-weighted class activation mapping (Grad-CAM), provide insights into model behavior through feature attribution and visual explanation mechanisms7,8,9. These methods have been increasingly applied to healthcare applications to support clinical interpretation, model auditing, and decision support10,11. However, explainability alone does not guarantee reliability, reproducibility, or clinical utility. Recent studies have highlighted challenges related to explanation instability, sensitivity to perturbations, limited clinician validation, and inconsistencies between explanation outputs and true model reasoning12,13,14,15.
In addition to explainability challenges, reproducibility remains a significant concern in healthcare machine learning research16,17,18. Many published studies provide limited information regarding preprocessing procedures, software environments, hyperparameter configurations, validation strategies, and implementation workflows, making independent replication difficult19,20,21. Furthermore, healthcare AI studies frequently focus on predictive performance while providing limited guidance regarding explanation quality assessment, clinician-centered evaluation, and practical implementation considerations22. These limitations may hinder the translation of explainable AI systems from research environments to real-world healthcare settings23,24,25,26,27.
Current healthcare XAI workflows often evaluate individual explanation techniques or predictive models in isolation and rarely provide standardized protocols that integrate data preparation, predictive modeling, explainability assessment, human-centered evaluation, and reproducibility resources28,29,30,31. Consequently, researchers and practitioners may encounter difficulties when reproducing workflows, comparing explanation methods, or evaluating the practical utility of explainable AI systems across different healthcare datasets and application domains. A comprehensive and reproducible protocol is therefore needed to facilitate consistent implementation, evaluation, and reporting of healthcare explainable AI workflows32,33,34,35.
Here, we present a reproducible protocol for developing, evaluating, and interpreting explainable artificial intelligence systems in healthcare analytics. The protocol integrates data preprocessing, predictive modeling, explainability assessment using SHAP, LIME, and Grad-CAM, clinician-centered evaluation, validation procedures, and reproducibility resources within a unified workflow. The objective is to provide a standardized framework that supports transparent model development, explanation quality assessment, and reproducible implementation across diverse healthcare datasets36,37,38,39. The methodological contribution of this work lies in the integration of predictive analytics, explainability evaluation, clinician validation, and reproducibility practices into a single protocol suitable for healthcare AI research, education, and deployment-oriented studies40,41,42,43.
The primary contribution of this work is not the development of a novel explainability algorithm.
Rather, it provides a standardized, reproducible, and experimentally validated protocol that integrates predictive modeling, explainability assessment, visualization, evaluation, and human-centered interpretation into a unified workflow suitable for healthcare analytics and JoVE-based protocol dissemination. The primary objective of this work is not to introduce a novel predictive algorithm but to provide a standardized, reproducible, and experimentally validated protocol for implementing explainable artificial intelligence workflows in healthcare analytics. The protocol integrates data preprocessing, predictive modeling, explainability assessment, clinician-centered evaluation, and reproducibility resources into a unified framework suitable for adoption, replication, and educational dissemination.