Research Article

EfficientNetB7-Based Deep Learning Framework for Enhanced Classification of Lung and Colon Cancer Histopathological Images

DOI:

10.3791/68812

February 6th, 2026

In This Article

Summary

Loading...
$$\rightleftharpoonup{xx}$$ $$\longleftharp{xx}$$, $$\longrightharp{xx}$$,

Here, we introduce a deep learning system with the EfficientNetB7 model for the precise classification of lung and colon cancer histopathological images. The model gained 96% accuracy with the application of preprocessing, data augmentation, and transfer learning. The method has a high prospect for aiding clinical cancer diagnosis.

Abstract

Loading...
$$\rightleftharpoonup{xx}$$ $$\longleftharp{xx}$$, $$\longrightharp{xx}$$,

Early diagnosis of lung cancer plays a pivotal role in ensuring improved treatment and survival of patients. This remains a major focus in clinical research. Artificial intelligence (AI) has transformed pathology by significantly improving diagnostic accuracy and efficiency. This study presents a robust deep learning model in the shape of the pretrained EfficientNetB7 model to classify colon and lung tissue histopathological images with an extremely high accuracy of 96%. The model's performance was optimized using advanced preprocessing methods, fine-tuning, and domain-specific data augmentation techniques. These strategies help reduce problems such as class imbalance and subtle histological variations. To address the issue of overfitting, multiple data augmentation techniques were combined, and an early stopping criterion was incorporated. This approach enabled efficient and cost-effective training. Robust validation of the model demonstrates high utility for clinical applications and enables pathologists to deliver timely and accurate diagnoses. Integrating advanced deep learning models into medical imaging workflows holds great promise for early and accurate cancer diagnosis, ultimately improving patient outcomes.

Introduction

Loading...
$$\rightleftharpoonup{xx}$$ $$\longleftharp{xx}$$, $$\longrightharp{xx}$$,

Lung and colon cancer are among the most prevalent cancers in the world in terms of mortality. Lung cancer is the leading fatal cancer with over 1.8 million deaths annually, followed by colon cancer as the third most occurring malignancy and the second most common cause of cancer mortality, based on global health statistics. Accurate and early diagnosis is crucial for effective treatment and improved survival of these cancers. Histopathological examination, or microscopic evaluation of tissue samples by pathologists, remains one of the most frequent methods of detecting cancer1. Figure 1 shows the sample histopathological images of several types of lungs and colon tissues2.

Histology microscopy images of lung, colon tissues; cellular structure, cancer research analysis.
Figure 1: Sample images from the dataset. This figure shows representative examples from each class in the LC25000 dataset, highlighting the visual diversity among benign and malignant lung and colon tissue images. Please click here to view a larger version of this figure.

Digital pathology has transformed the industry by enabling it to digitize histology slides, which are now diagnosed using sophisticated deep learning algorithms3. In combination with deep learning models, algorithms are great at recognizing subtle patterns within large data, significantly boosting the accuracy and efficiency of diagnosis. Table 1 shows the description of each tissue class.

Class NameDescription
Lung Benign TissueNon-cancerous tissue in the lungs, not involved in cancerous growth.
Lung AdenocarcinomaA type of lung cancer arising from glandular epithelial tissue, known for its malignant growth patterns.
Lung Squamous Cell CarcinomaA kind of lung cancer distinguished by a certain cellular morphology that arises from the squamous cells lining the airways.
Colon AdenocarcinomaMost occurrences of colon cancer are caused by a common kind of colon cancer that begins in the glandular cells of the colon lining.
Colon Benign TissueHealthy tissue in the colon that does not display any signs of cancerous or pre-cancerous conditions.

Table 1: An overview of the five tissue and cancer classes in the dataset, including class names, characteristics, and the number of images in each class.

EfficientNetB7, a cutting-edge convolutional neural network architecture, has been identified as particularly effective for image classification tasks due to its balanced scaling in depth, width, and resolution. Compared to widely used architectures such as ResNet and DenseNet, EfficientNetB7 achieves a superior balance between classification accuracy and computational efficiency through compound scaling, which proportionally scales network depth, width, and resolution. While ResNet and DenseNet have demonstrated robust performance in medical image classification tasks, they typically require a higher parameter count and greater computational resources to reach comparable accuracy, making EfficientNetB7 more suitable for large-scale histopathological analysis in resource-constrained clinical settings4. The application of advanced computational models, including AI tools such as EfficientNetB7, offers new possibilities for overcoming traditional diagnostic challenges, providing tools that support pathologists and potentially improve diagnostic outcomes. Integrating AI into pathology practice could democratize access to quality diagnostic services, making it feasible for regions with limited medical expertise to perform advanced diagnostics.

Recent benchmark studies and comprehensive reviews in medical image analysis have consistently emphasized the growing impact of deep learning architectures in histopathology5. Large-scale analyses have shown that convolutional neural networks and their modern variants provide state-of-the-art performance in identifying subtle morphological features across multiple tissue types, often surpassing traditional diagnostic methods. Comparative evaluations highlight that advanced architectures, including EfficientNet and other scalable models, not only achieve superior classification accuracy but also demonstrate improved computational efficiency6. In addition, systematic reviews of digital pathology underscore its transformative role in routine diagnostics by enabling reproducibility, reducing inter-observer variability, and allowing integration with automated pipelines. These collective findings affirm the relevance of adopting advanced deep learning models for histopathological cancer classification and provide strong evidence that the application of scalable and efficient architectures can address clinical challenges such as class imbalance, histological variability, and limited availability of expert pathologists.

The objectives of this research are (i) creating and assessing a deep learning model for the categorization of lung and colon cancer histopathology images based on the EfficientNetB7 architecture; (ii) Demonstrating the model's ability to achieve high classification accuracy, potentially surpassing traditional diagnostic methods; (iii) Exploring the impact of data augmentation, preprocessing techniques, and transfer learning on model performance in a domain-specific context.

This study not only demonstrates the high diagnostic accuracy achievable with EfficientNetB7 on histopathological images but also rigorously addresses key challenges such as class imbalance and overfitting through targeted data augmentation, robust validation strategies, and early stopping techniques. The methodology is specifically designed to reflect real-world clinical settings by utilizing large-scale, augmented datasets and benchmarking the model's performance on clinically relevant metrics. The workflow's automation and efficiency ensure that the system is practical for integration into digital pathology labs, offering the potential to enhance pathologist productivity, improve diagnostic consistency, and enable broader access to expert-level analysis in both high-resource and resource-limited clinical environments.

Even with remarkable progress in deep learning for medical imaging, several challenges remain when classifying lung and colon cancer histopathological images7. Current models tend to fail in dealing with class imbalance, wherein some types of cancer can be underrepresented within datasets and create biased predictions. Further, subtle histological differences between subtypes of cancer, like adenocarcinoma and squamous cell carcinoma of the lung, present a formidable challenge to proper classification. Most previous work has been carried out on single-cancer classification or with limited small datasets, precluding wide applicability and utility in a clinical setting. Traditional approaches are also less computationally efficient, hence less desirable for large data or real-time clinical applications8.

The present study bridges these gaps by leveraging the EfficientNetB7 architecture, which is a state-of-the-art deep learning model with proven scalability and effectiveness. Random zooms, flips, rotations, and shifts are utilized as advanced data augmentation methods to maximize the model's robustness and generalization capability across different histopathological patterns. Transfer learning is also used to fine-tune the pre-trained EfficientNetB7 model for it to be capable of learning and adapting in a way specific to lung and colon cancer datasets and minimizing training time and computational cost9. The use of global average pooling and early stopping also improves the performance of the model in having good training and avoiding overfitting.

Application of deep learning in medical imaging, especially for the diagnosis of cancer, has made huge progress over the last few years. The conventional diagnosis of lung and colon cancer depends mainly on histopathological examination, wherein pathologists inspect tissue samples under a microscope10. While effective, it is labor-intensive, open to inter-observer variability, and usually restricted by the limited supply of expert pathologists, especially in resource-challenged settings. These limitations have led to the development of computerized AI- and deep learning-based diagnostic systems.

Deep learning models, particularly convolutional neural networks (CNNs), have been impressively successful at image analysis of histopathological specimens11. They are especially good at identifying intricate patterns and subtle tissue morphology variations, tasks that are suitable for cancer classification. Despite the progress, there remain deficiencies such as class imbalance, limited generalization over multi-dataset scenarios, and high computational costs in the current approaches. For instance, most of the work addresses single-cancer classification, which limits its use to multi-cancer diagnosis scenarios. Also, poor preprocessing and data augmentation techniques render it susceptible to overfitting, particularly under small or imbalanced dataset scenarios.

Recent developments in transfer learning have partially overcome these limitations by capitalizing on pretrained models on big-scale datasets such as ImageNet. EfficientNet, ResNet, and DenseNet are some models that have attracted attention in medical image analysis as they can learn high-level features from limited amounts of training data12. EfficientNet has also seen widespread adoption for its scalability and efficiency since it achieves state-of-the-art accuracy on several image classification benchmarks. EfficientNet's application in classifying combined lung and colon cancer is not extensively investigated, especially concerning handling class imbalance and subtle histological variations13.

Data augmentation has been a crucial strategy for model robustness and generalization improvement14. Through incorporating variations such as rotation, shifting, zooming, and flipping, data augmentation techniques simulate variability that exists in actual histopathological images, enabling models to learn invariant features. Early stopping and regularization techniques further improve model performance by avoiding overfitting and maximizing training efficiency. Despite all these advancements, there is still a need for end-to-end frameworks that embrace advanced preprocessing, augmentation, and transfer learning to achieve high accuracy and computational efficiency for multi-cancer classification. Table 215,16,17,18,19,20,21,22,23,24 shows the literature of the study from the existing literature.

StudyObjective
Talukder, M. A. et al.15 (2022) Present a hybrid ensemble feature extraction model combining deep learning and machine learning to identify colon and lung cancer.
Attallah, O. et al.16 (2022)Provide a low-weight deep learning framework that combines a variety of models and transformation techniques to help detect lung and colon cancers early on.
Hage Chehade, A. et al.17 (2022)Create a machine learning-based computer-aided diagnostic system that can identify different kinds of lung and colon tissues.
Wahid, R. R. et al.18 (2023)Use a computer-aided diagnosing system with CNNs to detect lung and colon cancers.
Kumar, N. et al.19 (2022)Compare and contrast the feature extraction techniques used to categorise colon and lung cancer.
Mehmood, S. et al.20 (2022)Make a pretrained neural network with modified layers diagnosis model for lung and colon cancers that is both accurate and efficient.
Zhou, L. et al.21 (2024)Predict new targets for Wogonin (WOG) in treating lung, bladder, and colon cancer using bioinformatics methods.
Reddy, K. R. et al.22 (2022)Provide a hybrid ensemble feature extraction method based on machine learning for the colon cancer (LCC) detection.
Shandilya, S., Nayak, S. R.23 (2022)CNNs and vision transformer design help to classify lung and colon cancer histopathology images.
Hasan, M. et al.24 (2023)Using many CNN models, classify images of lung and colon cancer to enhance diagnosis procedures.

Table 2: A comparison of relevant previous studies on histopathological image classification, detailing datasets used, model types, and reported outcomes.

This paper builds on these contributions by incorporating a deep learning model utilizing the EfficientNetB7 architecture, complemented by state-of-the-art data augmentation and transfer learning. The current solution outperforms the flaws of previous solutions through enhanced classification accuracy, class equilibrium, and lower computational cost, and hence it is an optimal solution for usage in the clinical setting.

Access restricted. Please log in or start a trial to view this content.

Protocol

Loading...
$$\rightleftharpoonup{xx}$$ $$\longleftharp{xx}$$, $$\longrightharp{xx}$$,

This study did not involve any direct experimentation on human participants or animals. All work was conducted using the publicly available, anonymized LC25000 dataset of histopathological images, which contained no identifiable patient information or direct handling of human tissue. Institutional Review Board (IRB) or Institutional Animal Care and Use Committee (IACUC) approval was not required. All procedures complied with ethical standards and adhered to the dataset's terms of use for academic research. Figure 2 shows the steps of the workflow diagram.

Deep learning workflow diagram for image classification; includes normalization, augmentation, classifier.
Figure 2: Workflow of the proposed method. The workflow includes data preprocessing, augmentation, model training, and evaluation. Please click here to view a larger version of this figure.

Dataset description
The LC25000 dataset was utilized for this study. It consisted of 25,000 histopathological images, all uniformly formatted as JPEG files with a resolution of 768 × 768 pixels. The dataset originated from an initial set of 1,250 original images, including 250 benign lung tissue images, 250 lung adenocarcinomas, 250 lung squamous cell carcinomas, 250 benign colon tissue images, and 250 colon adenocarcinomas. The remaining images were generated through data augmentation to expand the dataset, ensuring a diverse and extensive collection for robust model training and validation.

The high-resolution images allowed for detailed examination of cellular structures, which was essential for accurate cancer classification. Data augmentation was necessary due to the limited number of original images; techniques such as rotations, flips, zooms, and shifts were employed to synthetically increase the dataset size and diversity. This process generated an additional 23,750 synthetic images in a class-balanced manner, resulting in a final dataset of 25,000 images, with each class containing 5,000 images post-augmentation.

Data preprocessing
The raw histopathological images within the dataset had a resolution of 768 × 768 pixels. To ensure compatibility with the EfficientNetB7 architecture and to maximize computational efficiency, each image was resized from 768 × 768 to 224 × 224 pixels. This resizing process represented a trade-off between preserving essential histological features and reducing computational overhead. The resizing operation was mathematically represented as shown in

Equation 1: Equation describing image resizing process, formula X_resized=resize(X,(h,w)), used in image processing.     (1)

where X denoted the original image, Xresized was the resized image, and h and w were the desired height and width, respectively. All pixel values were normalized to a range of 0 to 1, which stabilized and accelerated training by ensuring a consistent input distribution across images. This normalization was mathematically calculated as shown in Equation 2. This step was essential to improve model convergence during training.

Data normalization formula, X_norm=(X-μ)/σ, statistical method application, educational diagram.     (2)

Data augmentation was employed as a critical technique to enhance the robustness and applicability of deep learning models, especially in the context of medical imaging with limited and imbalanced datasets. In this study, a variety of augmentation techniques were applied using the ImageDataGenerator module from TensorFlow to simulate in vivo variation in histopathological images. These techniques included random rotations of up to 20 degrees to simulate tissue orientation variability, horizontal and vertical shifts up to 20% of the image dimensions to mimic position variability during slide preparation, and random zooms up to 20% to reflect variability in magnification levels.

Additionally, horizontal flipping was performed to introduce mirror-image variations, thereby further enriching the training dataset. To manage newly synthesized pixels generated during these transformations, a nearest-neighbor filling strategy was adopted to ensure smooth transitions and preserve the integrity of key histological features. The 2D rotation matrix used for augmentation is represented in Equation 3, and Equation 4 defines a function with an incremental change:

Rotation matrix formula R(θ) diagram with trigonometric functions cosθ and sinθ for geometry analysis.     (3)

F(x)=x+Δx, mathematical function equation for calculus study, formula representation.     (4)

These augmentation techniques compensated for challenges such as class imbalance, slight histological variations, and dataset limitations by enabling the model to learn invariant features and reducing the risk of overfitting. By synthetically expanding the dataset and incorporating controlled variability, data augmentation significantly enhanced the model's ability to generalize to new, unseen data, thereby producing a stronger and more robust system for clinical deployment in lung and colon cancer diagnosis.

Model architecture
The model was developed based on the EfficientNetB7 architecture, a state-of-the-art deep learning network renowned for its performance, scalability, and effectiveness in image classification tasks. EfficientNetB7 was constructed using a series of Mobile Inverted Bottleneck Convolution (MBConv) blocks, which incorporated depthwise separable convolutions. This structural design effectively compressed the model, resulting in significant reductions in the number of parameters with minimal loss in performance. Such innovation enabled the model to achieve high accuracy with lower computational costs, making it particularly suitable for computationally intensive applications such as histopathological image analysis. Figure 3 shows the model architecture of the pretrained EfficientNetB7.

Neural network architecture, MBConv layers diagram, showcasing convolutional layers and connections.
Figure 3: Model architecture of EfficientNetB7. This figure shows the detailed structure of the EfficientNetB7 model, illustrating its main layers and functional blocks used for image classification. Please click here to view a larger version of this figure.

The model accepted preprocessed histopathological images that had been resized from their original dimensions of 768 × 768 pixels down to 224 × 224 pixels to match the input requirements of EfficientNetB7. Each image underwent several processes within the network, including batch normalization, Swish activation functions, and convolution operations. Batch normalization was calculated as shown in Equation 5. The Swish activation function introduced non-linearity and enabled faster, more efficient training compared to traditional functions like ReLU. These layers were designed to extract and enhance hierarchical features, allowing the model to learn progressively from low-level to high-level patterns as the images propagated through the network.

Equation for batch normalization process, showing mean, variance adjustment in machine learning.     (5)

The convolutional backbone of EfficientNetB7, pretrained on the ImageNet dataset, provided a highly generalizable set of feature detectors effective across a broad range of image domains. The pretrained base played a key role in capturing subtle yet informative patterns in histopathological images, which are often characterized by high intra-class variability and fine-grained features. The output from the convolutional layers was passed to a global average pooling layer, which reduced the spatial dimensions to a single feature vector for each image. This vector, encapsulating the most relevant features learned by the convolutional layers, was then flattened into a one-dimensional array for compatibility with fully connected layers. The algorithm for enhanced detection of lung and colon cancer is provided in Supplementary File 1.

The subsequent network architecture included two dense (fully connected) layers with 128 and 64 units, respectively. Both layers utilized the Rectified Linear Unit (ReLU) activation function to introduce additional non-linearity and to fine-tune the extracted features, enabling the model to recognize more precise patterns and relationships within the data. The final layer of the model was a SoftMax output layer with five neurons, each representing one of the five classes in the dataset. The SoftMax function produced a probability distribution across the classes, which reflected the model's prediction for each input image, as described in Equation 6:

Softmax function equation, σ(z)_i, for probability distribution analysis.     (6)

The combination of state-of-the-art convolutional techniques, robust regularization, and effective optimization strategies resulted in a deep learning model well-suited for high-stakes medical image classification. The model achieved precision, scalability, and operational efficiency, demonstrating strong potential for clinical application in lung and colon cancer diagnosis. Within this framework, regularization was primarily accomplished through advanced data augmentation and the use of early stopping; no explicit dropout or weight decay was applied to the dense layers.

Training and validation
The validation and pretraining pipeline of the convolutional neural network (CNN) for LC25000 histopathological image classification was designed to optimize model performance and ensure good generalization. The pipeline began with comprehensive preprocessing, which included reducing the original 768 × 768-pixel images to 224 × 224 pixels to match the input requirements of the EfficientNetB7 architecture. Pixel intensities were normalized to the range [0, 1] to stabilize and accelerate training, and data augmentation techniques, including random rotations, shifts, zooms, and horizontal flips -- were applied to introduce variability and mitigate overfitting. These preprocessing operations ensured that the model learned invariant features and generalized well to new data.

The EfficientNetB7 model, pretrained on the ImageNet dataset, was employed as the base, providing a robust feature detector set capable of identifying subtle histological patterns. The model was subsequently fine-tuned with two dense layers (128 and 64 units) activated by the Rectified Linear Unit (ReLU) function, followed by a SoftMax output layer for five-class classification. This architecture was optimized for both precision and computational efficiency, making it well-suited for medical image analysis. The Adam optimizer update rule for θ is shown in Equation 7.

Adam optimization equation, symbol, learning rate adjustment in neural networks, educational use.     (7)

Model training was performed in batches of 128 images using the Adam optimizer, which adaptively scaled the learning rate to promote effective convergence. An early stopping criterion was implemented to monitor validation loss and halt training if no improvement was observed, thus preventing overfitting and optimizing computational resources. Throughout training, key performance metrics -- including accuracy, loss, precision, and recall -- were tracked using a hold-out validation set comprising 20% of the data. The hold-out set provided critical insights into the model's generalization capabilities and mitigated the risk of overfitting to unseen data.

The integration of advanced preprocessing, data augmentation, and optimization strategies resulted in excellent model accuracy and robustness, demonstrating strong potential for clinical application in lung and colon cancer detection. All experiments were conducted on an NVIDIA Tesla P100 GPU with 16 GB memory and 64 GB system RAM. The full training pipeline -- including preprocessing, augmentation, and model optimization -- required approximately three hours to complete ten epochs with a batch size of 128. This computational infrastructure was selected to efficiently manage the large, augmented dataset while maintaining reasonable training times appropriate for clinical research environments.

For reproducibility, all critical hyperparameters were explicitly specified. The model was trained using the Adam optimizer with a constant learning rate of 0.001. Training was carried out for a maximum of 10 epochs, with early stopping based on validation loss and a patience of three epochs to prevent overfitting. A batch size of 128 was utilized for both training and validation phases. To ensure consistent results, the random seed was set to 42 at all stages of data loading and model training. No learning rate decay schedule was applied during training.

The protocol combined systematic preprocessing, well-calibrated hyperparameters, rigorous validation schemes, and an open computational environment to promote both accuracy and reproducibility. The workflow, by integrating robust deep learning architecture, effective optimization, and reproducible experimental design, established a solid foundation for histopathological image classification. In addition to demonstrating strong performance on lung and colon cancer detection, the protocol defined a scalable framework that could be extended to similar biomedical imaging challenges in future research.

During model development, several practical adjustments were implemented to achieve reliable and reproducible results. Early stopping and advanced data augmentation were used to address class imbalance and overfitting. High memory usage during training was mitigated by resizing images to 224 × 224 pixels and using a batch size of 128. When slow convergence was encountered, the learning rate was empirically set to 0.001. Manual verification of the image directory structure was performed to prevent data loading errors, and random seeds were fixed to ensure reproducibility. These strategies can serve as practical guidance for researchers facing analogous challenges in training deep learning models with large histopathological image datasets.

Statistical analysis
Statistical modeling of the model's performance was conducted using a comprehensive set of metrics to assess accuracy, stability, and generalizability. Key evaluation measures included precision, recall, F1-score, and accuracy to quantify classification performance. The confusion matrix was utilized to report class-specific error metrics, calculated according to Equations 8-10:

Precision formula: TP/(TP+FP), equation, relevant for statistical data analysis and classification models.     (8)

Recall formula: TP/(TP+FN), key performance metric, equation diagram, data accuracy analysis.     (9)

F1 score formula: 2 × (Precision × Recall) / (Precision + Recall), statistical method.     (10)

Additionally, measures of error such as Mean Squared Error (MSE), Root Mean Squared Error (RMSE), and Mean Absolute Error (MAE) were evaluated to assess prediction accuracy, calculated using Equations 11-13:

Mean squared error formula, MSE=1/n Σ(Yi-Ŷi)², statistical analysis, equation.     (11)

RMSE formula, equation showing root mean square error calculation, educational math concept.     (12)

Mean Absolute Error formula: MAE=1/n∑|Yi−Ŷi|; used in statistical analysis and modeling.     (13)

The Receiver Operating Characteristic (ROC) curve and Area Under the Curve (AUC) were employed to evaluate the model's ability to discriminate between classes at various thresholds. The AUC was calculated as shown in Equation 14:

AUC integral formula; analysis of true/false positive rates; equation in ROC curve studies.     (14)

These statistical analyses provided a thorough assessment of the model's classification performance, robustness, and generalization to unseen data.

Ablation study
To further assess the contribution of each component within the classification framework, an ablation study was conducted by selectively omitting or modifying specific architectural elements of the final model. Specifically, two alternative experimental configurations were evaluated: one configuration included both the Global Average Pooling and Flatten layers, followed by two dense layers (128 and 64 units, respectively); the other configuration omitted both the Flatten layer and the 128-unit dense layer, leaving only the Global Average Pooling layer and a single 64-unit dense layer.

The complete architecture, as reported in the main results, achieved a validation accuracy of approximately 96%. When the Flatten layer was excluded, and only the 128- and 64-unit dense layers were used after Global Average Pooling, the validation accuracy dropped to 81.6%. Conversely, when the 128-unit dense layer was excluded, and the Flatten layer was retained, the validation accuracy decreased to 88%. These results indicated the necessity of including both the extra dense layer and the appropriate architectural components to achieve optimal classification performance. The findings from the ablation experiments consistently demonstrated that the combination of Global Average Pooling, Flattening, and two fully connected layers enabled a more expressive feature representation, resulting in more robust and accurate histopathological image classification.

Access restricted. Please log in or start a trial to view this content.

Results

Loading...
$$\rightleftharpoonup{xx}$$ $$\longleftharp{xx}$$, $$\longrightharp{xx}$$,

Figure 4 presents the training and validation Accuracy. Figure 5 presents the training and validation Loss.

Training vs validation accuracy graph; epochs, accuracy trends, neural network model evaluation.
Figure 4: Training and...

Access restricted. Please log in or start a trial to view this content.

Discussion

Loading...
$$\rightleftharpoonup{xx}$$ $$\longleftharp{xx}$$, $$\longrightharp{xx}$$,

In the critical review of mislabeled instances under the EfficientNetB7 deep learning architecture, a critical examination is carried out on instances where model predictions do not match real labels within the validation dataset. Critical analysis is of extreme importance in analyzing certain errors of classification, particularly when the model misclassifies various histopathological features of lung and colon tissues11. The procedure is to make class predictions on all the images in the validat...

Access restricted. Please log in or start a trial to view this content.

Disclosures

Loading...
$$\rightleftharpoonup{xx}$$ $$\longleftharp{xx}$$, $$\longrightharp{xx}$$,

The authors declare that there is no conflict of interest regarding the publication of this manuscript. No financial or personal affiliations have influenced the research, results, or conclusions presented in this work.

Acknowledgements

Loading...
$$\rightleftharpoonup{xx}$$ $$\longleftharp{xx}$$, $$\longrightharp{xx}$$,

This research is supported by Princess Nourah bint Abdulrahman University Researchers Supporting Project number (PNURSP2026R195), Princess Nourah bint Abdulrahman University, Riyadh, Saudi Arabia. The authors extend their appreciation to the Deanship of Research and Graduate Studies at King Khalid University for funding this work through Large group research under grant number RGP2/749/46.

Access restricted. Please log in or start a trial to view this content.

Materials

List of materials used in this article
NameCompanyCatalog NumberComments
A100 GPU (CUDA)NVIDIACUDA Version 11.0GPU acceleration for model training and evaluation.
Kaggle PlatformGoogleN/ACloud based Notebook for Machine Learning Model Development
KerasTensorFlow (Google)Version 2.6.0Deep learning API running on top of TensorFlow.
LC25000Borkowski AA, Bui MM, Thomas LB, Wilson CP, DeLand LA, Mastorides SM. Lung and Colon Cancer Histopathological Image Dataset (LC25000)N/AThis dataset contains 25,000 histopathological images with 5 classes. All images are 768 x 768 pixels in size and are in jpeg file format.
MatplotlibPython Software FoundationVersion 3.5.0Visualization library for plotting results.
NumPyPython Software FoundationVersion 1.19.5Numerical computing library.
OpenCVOpen SourceVersion 4.5.4Image processing and computer vision library.
PandasPython Software FoundationVersion 1.3.4Data analysis and manipulation tool.
Python (Anaconda Distribution)Anaconda IncVersion 3.7.12Includes pre-installed packages and environment management tools.
Scikit-learnPython Software FoundationVersion 0.23.2Machine learning tools for performance evaluation.
TensorFlowGoogleVersion 2.6.2Deep learning framework for diffusion models.

References

Loading...
$$\rightleftharpoonup{xx}$$ $$\longleftharp{xx}$$, $$\longrightharp{xx}$$,
  1. Al-Jabbar, M., Alshahrani, M., Senan, E. M., Ahmed, I. A. Histopathological analysis for detecting lung and colon cancer malignancies using hybrid systems with fused features. Bioengineering. 10 (3), 383(2023).
  2. Borkowski, A. A., Bui, M. M., Thomas, L. B., Wilson, C. P., DeLand, L. A., Mastorides, S. M., et al. Lung and colon cancer histopathological image dataset (LC25000). arXiv. , (2019).
  3. Sakamoto, T., Furukawa, T., Lami, K., Pham, H. H. N., Uegami, W., Kuroda, K., et al. A narrative review of digital pathology and artificial intelligence:focusing on lung cancer. Transl Lung Cancer Res. 9 (5), 2255(2020).
  4. WHO Consolidated Guidelines on Tuberculosis. Module 3:Diagnosis: Rapid Diagnostics for Tuberculosis Detection. , World Health Organization. https://www.who.int/publications/i/item/9789240089488 (2024).
  5. Al-Zoghby, A. M., Ismail Ebada, A., Saleh, A. S., Abdelhay, M., Awad, W. A. A comprehensive review of multimodal deep learning for enhanced medical diagnostics. Comput Mater Continua. 84 (3), 2025(2025).
  6. Attallah, O. Lung and colon cancer classification using multiscale deep features integration of compact convolutional neural networks and feature selection. Technologies. 13 (2), 54(2025).
  7. Patharia, P., Sethy, P. K., Nanthaamornphong, A. Advancements and challenges in the image-based diagnosis of lung and colon cancer:a comprehensive review. Cancer Inform. 23, 11769351241290608(2024).
  8. Liao, Q. M., Hussain, W., Liao, Z. X., Hussain, S., Jiang, Z. L., Zhu, Y. H., et al. Computer-aided application in medicine and biomedicine. Int J Comput Intell Syst. 18 (1), 221(2025).
  9. Gautam, N., Ghosh, S., Sarkar, R. CNN models aided with a metaclassifier for lung carcinoma classification using histopathological images. Multimed Tools Appl. 84, 28493-28517 (2025).
  10. Mohammed, M. S., Mohammed, A. M. S. Safety diagnostic tool for non-small cell lung cancer (NSCLC) lyophilized serum. Clin Cancer Investig J. 13 (4), 15-19 (2024).
  11. Iqbal, S., Qureshi, A. N., Alhussein, M., Aurangzeb, K., Kadry, S. A novel heteromorphous convolutional neural network for automated assessment of tumors in colon and lung histopathology images. Biomimetics. 8 (4), 370(2023).
  12. Haziq, U., Uddin, J., Rahman, S., Yaseen, M., Khan, I., Khan, J., et al. Improving lung cancer detection with enhanced convolutional sequential networks. Sci Rep. 15 (1), 32099(2025).
  13. Abd El-Aziz, A. A., Mahmood, M. A., Abd El-Ghany, S. Advanced deep learning fusion model for early multi-classification of lung and colon cancer using histopathological images. Diagnostics. 14 (20), 2274(2024).
  14. Ravindran, U., Gunavathi, C. Cancer disease prediction using integrated smart data augmentation and capsule neural network. IEEE Access. 12, 81813-81826 (2024).
  15. Talukder, M. A., Islam, M. M., Uddin, M. A., Akhter, A., Hasan, K. F., Moni, M. A. Machine learning-based lung and colon cancer detection using deep feature extraction and ensemble learning. Expert Syst Appl. 205, 117695(2022).
  16. Attallah, O., Aslan, M. F., Sabanci, K. A framework for lung and colon cancer diagnosis via lightweight deep learning models and transformation methods. Diagnostics. 12 (12), 2926(2022).
  17. Hage Chehade, A., Abdallah, N., Marion, J. -M., Oueidat, M., Chauvet, P. Lung and colon cancer classification using medical imaging:a feature engineering approach. Phys Eng Sci Med. 45 (3), 729-746 (2022).
  18. Wahid, R. R., Nisa', C., Amaliyah, R. P., Puspaningrum, E. Y. Lung and colon cancer detection with convolutional neural networks on histopathological images. AIP Conf Proc. 2654, 020020(2023).
  19. Kumar, N., Sharma, M., Singh, V. P., Madan, C., Mehandia, S. An empirical study of handcrafted and dense feature extraction techniques for lung and colon cancer classification from histopathological images. Biomed Signal Process Control. 75, 103596(2022).
  20. Mehmood, S., et al. Malignancy detection in lung and colon histopathology images using transfer learning with class selective image processing. IEEE Access. 10, 25657-25668 (2022).
  21. Zhou, L., et al. Identification of novel targets and mechanisms of wogonin on lung cancer, bladder cancer, and colon cancer. J Future Foods. 4 (3), 267-279 (2024).
  22. Reddy, K. R., DeLaune, R. D., Inglett, P. W. Biogeochemistry of Wetlands: Science and Applications. , CRC Press. Boca Raton. (2022).
  23. Shandilya, S., Nayak, S. R. Analysis of lung cancer by using deep neural network. Innovation in Electrical Power Engineering, Communication, and Computing Technology. Lecture Notes in Electrical Engineering. , Springer. Singapore. (2022).
  24. Vision transformer-based classification for lung and colon cancer using histopathology images. Hasan, M., et al. 2023 International Conference on Machine Learning and Applications (ICMLA), Jacksonville, FL, USA, , (2023).
  25. Cen, J., Hu, N., Shen, J., Gao, Y., Lu, H. Pathological functions of lysosomal ion channels in the central nervous system. Int J Mol Sci. 25 (12), 6565(2024).
  26. Li, H., Wang, Z., Guan, Z., Miao, J., Li, W., Yu, P., et al. UCFNNet: ulcerative colitis evaluation based on fine-grained lesion learner and noise suppression gating. Comput Methods Programs Biomed. 247, 108080(2024).
  27. Song, W., Wang, X., Guo, Y., Li, S., Xia, B., Hao, A. CenterFormer:a novel cluster center enhanced transformer for unconstrained dental plaque segmentation. IEEE Trans Multimedia. 26, 10965-10978 (2024).
  28. Gayap, H. T., Akhloufi, M. A. Deep machine learning for medical diagnosis, application to lung cancer detection:a review. BioMedInformatics. 4 (1), 236-284 (2024).
  29. Luan, S., et al. Deep learning for fast super-resolution ultrasound microvessel imaging. Phys Med Biol. 68 (24), 245023(2023).
  30. Sun, D., Hu, Y., Li, X., Peng, J., Dai, Z., Wang, S. Unlocking the full potential of memory T cells in adoptive T cell therapy for hematologic malignancies. Int Immunopharmacol. 144, 113392(2025).
  31. Wang, G., Ma, Q., Li, Y., Mao, K., Xu, L., Zhao, Y. A skin lesion segmentation network with edge and body fusion. Appl Soft Comput. 170, 112683(2025).
  32. Singh, O., Kashyap, K. L., Singh, K. K. Lung and colon cancer classification of histopathology images using convolutional neural network. SN Comput Sci. 5 (2), (2024).
  33. Diagnose colon and lung cancer histopathological images using pre-trained machine learning model. Maheshwari, U., Kiranmayee, B. V., Suresh, C. Proc Int Conf Contemp Comput Inform, , 1078-1082 (2022).
  34. Masud, M., Sikder, N., Nahid, A. -A., Bairagi, A. K., AlZain, M. A. A machine learning approach to diagnosing lung and colon cancer using a deep learning-based classification framework. Sensors. 21 (3), 748(2021).
  35. Lung cancer detection from histopathological images using deep learning. Mohalder, R. D., et al. International Conference on Machine Intelligence and Emerging Technologies, , Noakhali, Bangladesh. (2022).
  36. Detection of colon cancer using Inception V3 and ensembled CNN model. Swarna, I. J., Hashi, E. K. Proc Int Conf Electr Comput Commun Eng, 2023, 1-6 (2023).
  37. Reis, H. C., Turk, V. Transfer learning approach and nucleus segmentation with MedCLNet colon cancer database. J Digit Imaging. 36 (1), 306-325 (2022).
  38. Deep convolutional neural networks for early-stage detection and prognostication of lung and colon cancer. Laxmikant, K., Arthi, A., Vinodhini, V., Natarajan, B., Bhuvaneswari, R., Selvam, P. 2024 International Conference on Integrated Circuits and Communication Systems (ICICACS), Raichur, India, , (1109).
  39. Musthafa, M. M., Manimozhi, I., Mahesh, T. R., Guluwadi, S. Optimizing double-layered convolutional neural networks for efficient lung cancer classification through hyperparameter optimization and advanced image pre-processing techniques. BMC Med Inform Decis Mak. 24 (1), 142(2024).
  40. Histopathological analysis advancements in deep learning for the diagnosis of lung and colon cancer with explanatory power via visual saliency. Kashyap, S., Shukla, A. K., Naim, I. Proc Int Conf Innov Emerg Trends Comput Inf Technol, 2125, Springer. Cham. 148-158 (2024).
  41. Alhassan, A. M. Enhanced pre-processing technique for histopathological image stain normalization and cancer detection. Multimed Tools Appl. 84, 29733-29761 (2024).
  42. Hijazi, A., Bifulco, C., Baldin, P., Galon, J. Digital pathology for better clinical practice. Cancers. 16 (9), 1686(2024).
  43. Ameyaw, S. A., Afari, D. A., Boateng, J. Advancements in image-based analyses for morphology and staging of colon cancer: a comprehensive review. Biomed Res Int. 2025 (1), 9214337(2025).
  44. Leveraging artificial intelligence for enhanced histopathological image analysis in lung and colon cancer diagnosis:development of a mobile application. Cherrat, L., Khrouch, S., Chraibi, M., Ezziyyani, M. Proc Int Conf Adv Intell Syst Sustain Dev, , Springer. Cham. 748-766 (2024).

Access restricted. Please log in or start a trial to view this content.

Reprints and Permissions

Request permission to reuse the text or figures of this JoVE article

Request Permission

Tags

EfficientNetB7 ModelDeep LearningLung Cancer DiagnosisColon Cancer DiagnosisHistopathological ImagesMedical ImagingData AugmentationEarly Cancer DetectionModel Fine TuningArtificial Intelligence Pathology

Related Articles