$$\rightleftharpoonup{xx}$$
$$\longleftharp{xx}$$,
$$\longrightharp{xx}$$,
Breast cancer remains the most commonly diagnosed cancer and the second leading cause of cancer-related death among women globally1. In the United States alone, it accounts for nearly 30% of all new female malignancies, with over 280,000 new cases diagnosed annually2. Despite therapeutic advances, particularly in HER2-positive and hormone receptor-positive subtypes, resistance to treatment and recurrence remain critical challenges-especially for aggressive subtypes like triple-negative breast cancer (TNBC), which lacks targeted therapies3,4. This underscores the urgent need for precision-driven drug discovery to identify effective therapeutic agents and combinations tailored to individual molecular profiles. Drug discovery, traditionally guided by experimental and trial-and-error methods, has seen remarkable acceleration through the integration of machine learning (ML) techniques5,6. ML enables the modeling of complex, nonlinear relationships across high-dimensional biomedical data and can assist in target identification, biomarker discovery, drug sensitivity prediction, and combination therapy design7,8. However, the practical deployment of ML models in oncology faces several hurdles, including model interpretability, reproducibility, overfitting on sparse datasets, and generalization across cancer subtypes9,10,11.
To overcome these limitations, recent research has focused on combining deep learning for feature extraction with ensemble learning for robust prediction. In studies evaluating multiple algorithms, models such as Artificial Neural Networks (ANN) achieved accuracy levels up to 93.2%, outperforming conventional classifiers like Naïve Bayes and Decision Trees12. Additionally, integrated feature mining techniques have uncovered key driver genes and molecular targets through databases like GEO (Gene Expression Omnibus) and GSE45827, identifying up to 1,700 differentially expressed genes, some of which exhibit known drug interactions13. Further, novel drug repurposing studies have revealed the potential of non-oncology compounds like calcitriol to reduce breast cancer cell viability more effectively than standard treatments like neratinib, especially in HER2+ cell lines14. Investigations into the Akt-signaling pathway have also shown promise in overcoming trastuzumab resistance, suggesting molecular pathway targeting as an alternative to receptor-focused therapy15,16. Yet, despite these advancements, a systematic and explainable framework capable of predicting continuous drug response values, ranking effective drug combinations, and visualizing pharmacological similarities remains underexplored in the current literature. Many models are either classification-based or lack translational clarity, especially when applied to real-world pharmacogenomic datasets.
Drug discovery and decision-making can be improved by machine learning (ML), which offers instruments for high-quality data. All phases of drug discovery, including target validation, biomarker identification, and clinical trial analysis, can benefit from the use of machine learning. Interpretability and reproducibility of ML-generated outcomes are obstacles, too17. Reducing failure rates and expediting the process can be achieved by addressing these problems and raising knowledge of validation variables. Using machine learning algorithms, the researchers assessed biopsy samples at various stages of cancer. Test accuracies were high, with ANN 93.2%, Naïve Bayes (NB) 90.4%, Decision Tree (DT) 87.8%, and RF 85.9%, according to the findings. A total of 350 predicted genes and 164 differentially expressed genes were found by combining the GEO database by Rakhshaninejad et al.18. In the combined dataset, the Binary Grey Wolf Optimization with Simulated Annealing Ensemble (BGWO_SA_Ens) algorithm found 1404 genes, while in the GSE45827 dataset, it found 1710. Around 35 superior genes, along with their roles in important pathways and the relationships between superior genes and anticancer medications, were found. To find target genes from the Epidermal Growth Factor Receptor (EGFR (EGFR) overexpression signaling pathway and their related family members, molecular networking investigations were carried out by Nagaraj et al.19 A medication called calcitriol, which is authorized to treat conditions unrelated to cancer, had strong binding affinities with each of the four receptors. According to in vitro cytotoxicity studies, calcitriol reduced SK-BR-3 cell viability in a dose-dependent manner, indicating superior cytotoxicity and reduced proliferation of breast cancer cells in comparison to neratinib. An active and druggable Akt-signaling pathway was suggested by Jernström et al.20 that two cell lines that were trastuzumab-insensitive were responsive to an Akt1/2 kinase inhibitor. Instead of focusing on HER2 amplification or expression, the study recommends targeting the Akt-signaling pathway and taking molecular aspects into account when making treatment decisions. Thirty percent of new female malignancies in the US are breast cancers, making it the most frequent malignant disease among women. The goal of Witt and Tollefsbol21 was to develop a fundamental tool that would help researchers select a breast cancer cell line for use in xenograft experiments, cancer prevention, and epigenetic discoveries, among other fields. Also covered are debates about the provenance of specific breast cancer cell lines and the advantages of employing patient-derived xenograft (PDX) as opposed to cell-derived xenograft (CDX). The use of drug prediction techniques to provide new drug discovery hypotheses was examined in Gruener et al.22, with a focus on triple-negative breast cancer (TNBC). On the basis of cell line transcriptome data, machine learning models of drug response were constructed and then applied to patient tumor data. The findings demonstrated that the Wee1 inhibitor AZD-1775 had preferential action in TNBC and that TP53 mutations were strongly linked to its effectiveness. In order to forecast unknown drug-target interactions in breast cancer research, Song et al23 present a feature-based approach dubbed Pseudo Position-Specific Physicochemical Property-Derived Composition for Drug-Target Interaction Prediction (PsePDC-DTIs), which makes use of protein sequences, the Deep Canonical Correlation Analysis (DCCA) coefficient, and a molecular fingerprint descriptor. The technique predicts DTIs on four gold standard datasets using a random forest classifier and handles unbalanced data using SMOTE. Additionally, the model uses risk genes from genome-wide genetic research to investigate novel targets for the therapy of breast cancer. The model's superiority and validity are demonstrated by the ten possible DTIs it offers for therapy. Ten to twenty percent of instances of breast cancer are triple-negative breast cancer (TNBC). There are currently no targeted therapies for TNBC, despite advancements in HER2+ and hormonal receptor+ treatments24. Although the EGFR is expressed by the majority of patients, early studies did not find any discernible activity. Future experimental treatments for TNBC are suggested by recent findings and clinical advancements25.
Despite the growing integration of machine learning in drug discovery, current models often lack interpretability and reproducibility, limiting their translational application. While prior studies have explored classification accuracy and gene mining, few have systematically predicted continuous drug sensitivity (like LN_IC50) using hybrid interpretable models. Moreover, the combination of dimensionality reduction techniques with robust regressors remains underexplored in the context of breast cancer treatment. This study addresses that gap by introducing and evaluating a dual-pipeline strategy -- XGBoost and Autoencoder-XGBoost -- for high-fidelity prediction of drug response, coupled with explainability and synergy mapping tools for real-world clinical applicability.