$$\rightleftharpoonup{xx}$$
$$\longleftharp{xx}$$,
$$\longrightharp{xx}$$,
Automated and fast detection of mpox in medical cases is clinically significant since the disease can be similar to a variety of other vesicular or pustular eruptions at the initial visual evaluation. Clinicians in practice do not pay much attention to the morphology of lesions alone, but they combine pattern, distribution, history, exposure risk, epidemiological context, and laboratory confirmation1. Nonetheless, computational tools use images and can also perform a helpful triage procedure in environments where skilled dermatological review is constrained or delayed. It is especially applicable when it comes to outbreak tracking, teledermatology screening, and the settings that are resource-limited and have a significant number of possible cases that should be prioritized to undergo confirmatory testing.
The deep-learning-based approaches have also emerged as appealing to this task since they can discover discriminative images of lesions without the use of descriptors that need to be created manually2. The most popular review on the literature still belongs to convolutional neural networks due to their effectiveness in local texture, color range, and lesion margin recognition3. Recent work has further developed this paradigm with the help of transfer learning, attention mechanisms, and transformer modules, which try to model a larger spatial dependence of the image. Nonetheless, the literature on mpox has high reported performance that should be taken with caution. A large variety of studies have made use of sparse public datasets, have different sets of heterogeneous classes, or compare results with a variety of preprocessing and validation regimes4. Due to this, there is frequently a dilemma of whether performance increases are the result of a truly superior model, a desirable split, vigorous augmentation or disparities in the formulation of the underlying task.
Mpox remains one of the significant infectious diseases with adverse implications on public health, as its lesions involving skin infections can seemingly be confused with other vesicular and pustular lesions on the interface during eye screening in the course of the primary examination5. Practically, clinicians do not make their diagnosis of mpox based on morphology: pattern, distribution, history, epidemiology, and laboratory confirmation are components of the diagnosis decision6. Nevertheless, they can play a beneficial role in the triage process with the help of computational tools that rely on the images, particularly when the experience of the dermatologists is either inadequate, slow, or not available. They might be applicable in teledermatology, in cases where an outbreak is to be addressed with screening, and in places with resource limitations where suspected cases on the priorities list must be quickly addressed7.
Deep-learning-based methods, because they allow learning discriminative lesion representations, would be attractive to this task because they learn such representations without handcrafted descriptors, depending on the characteristics of the input images8. It has been a frenzy within the field of dermatological images to utilize convolutional neural networks (CNNs) in analyzing images, as they possess the ability to encode local texture, lesions' periphery, and exquisite morphology meaning. Later studies have provided this paradigm through transfer learning, attention, and transformer-based submodules to learn more general spatial relationships and contextual trends9. Nonetheless, the literature on the mpox imaging is still heterogeneous. It is not uncommon that performance is reported to be based not only on the architecture selection, but also based upon dataset structure, augmentation plan, classes definition, and validation structure10. Then, good headline performance does not necessarily indicate increased generalization.
In the current work, this issue is solved by using a more perceptible and internally consistent binary mpox image-classification benchmark11. It does not discuss the contribution in the format of proposing an entirely new hybrid architecture, as CNN-transformer methods have already been sprinkled throughout the literature12. Rather, the experiment introduces a standard comparison with explicit primary binary decomposition, where the extra multiclass pieces of evidence, preprocessing/training, and model performance are considered in a protocol-friendly way, and model performance is considered in a conservative way compared to experiment environment13. In this paradigm, a CNN-transformer model can still be of functional importance since CNNs are optimally structured to capture local lesion information, whereas transformer-based representations have the potential to represent long-range spatial structure of interest in dermatology as well.
Recent image-classification research of mpox can be divided into three widely streams of research that are founded on a methodology. The latter is based on the conventional CNNs and transfer-learning because of its maturity, computation speed, and recreation of lesion-specific textures14. The latter explores attention-based or transformer models, better suited to learn the global lesion context. The third integrates convolutional and attention in the hybrid pipelines to preserve the local sensitivity but use bigger body contextual modeling. Several of the studies report high performance which cannot be easily directly compared because the studies differ in the size of their dataset, curation strategy, structure of their classes, preprocessing and validation protocol15. They are binary classification oriented, multiclass dermatological discrimination oriented not all of which have the same measures of evaluation. Mpox image datasets made publicly accessible might also include near-duplicate or very similar images, which will require apparent performance to be inflated in case split control is not adequate16.
The given analysis is founded on a controlled binary set of 2,280 photographs of skin lesions, 1,020 photographs of mpox and 1,260 photographs of Other were planned. The Other category overlaps with the non-mpox lesions that appear to the naked eye, and this puts the task a clinically interesting early triage formulation, as compared to a complete dermatological localization system17. The images were all resized, normalized, standardized, and augmented and evaluated in a typical inner training-validation pipeline. The main objective was the comparison of the behavior of conventional and hybrid deep-learning models in a transparent benchmark and a conclusion on whether the hybrid architecture suggested offers a superior balance between discriminative and interpretability, relative to popular baselines.
The definite gap that the current revised version addresses is thus not the mere presentation of the hybrid model, since the strategies of hybrid CNN-transformer have already been covered in the literature. Instead, the disconnect is that what was needed was a more distinct, internally consistent metric that disregards the primary binary task and exploratory multiclass outputs, characterizes the preprocessing and training pipeline in a reproducible way, and presents the model performance in a manner that does not overstate novelty or clinical preparedness18. In that regard, a hybrid architecture is nevertheless worth considering, since CNNs excel at local feature extraction, and transformer-like attention can, in theory, model global lesion configurations and longer-range spatial connections to aid in visual diagnosis19. The practical query is not on whether a hybrid design will have a favorable balance of discrimination, stability, and interpretability as compared to typical baselines on the same curated data aspect, but on whether it provides a good balance of taxation, stability, and interpretability.
The latest works on mpox image analysis can be divided into three general streams of the study of methods. The former one uses traditional CNNs and transfer-learning architectures due to their maturity, computational efficiency, and ability to learn lesion-specific textures and local morphological features20. The second stream offers the vision-transformer or attention-enhanced models to be able to capture the wider spatial associations over the lesion field, particularly when the texture patches are composed of contextual lesion distribution, and, thus, they may be as important as individual texture patches. The third stream adds the convolutional and the attention to support hybrid pipelines in an attempt at maintaining the local sensitivity of CNNs but utilizing the contextual modeling ability of attention blocks.
Even though most of these works claim excellent accuracy, cross-study comparison is not an easy task due to the differences in size, curation strategy, class composition, and validation design across datasets21. There are publications which assess binary discrimination, and those which assess multiclass dermatological categorization some are also balanced in reporting precision and recall, and some focus on accuracy22. Also, reported mpox datasets can be based on online image collections, curated collections, or augmented subsets, which can boost reported performance when train-validation leakage or near-duplicate images are not well managed. Table 1 is thus not kept as a mere compendium of previous work, but rather as a schematic map of model families, dataset constraints, and the opportunities offered by the methods available in this research.
The methodological value of the given work is reflected in a more limited and more defensible manner. This paper will discuss the necessity of the transparent binary benchmark with the help of the provided curated dataset, report the preprocessing and training procedure and directly compare the proposed hybrid model with popular baselines23. These changes also separate the main binary experiment and the supplementation multiclass evidence in a way that allows performance claims to be associated with the appropriate experimental context24. The difference makes a decisive difference in the discursive analysis of model conduct, reported measures, and practicality.
The main experiment of this script is a binary classification experiment on a smaller set of 2280 images of skin lesions, which was curated by hand. In this dataset, the number of images that were identified as mpox was 1,020, and those that were categorized as other were 1,260. The other category entails the presentation of non-mpox related dermatological appearances that are of a visual character, mostly lesions where the manuscript refers to them as chickenpox- and measles-like lesions25. This dual nomenclature makes this binary grouping clinically significant as a primitive screening formulation because the model inquires whether the model is able to disaggregate suspected mpox and visually confusing alternatives, but is also less complex than a full multiclass dermatology benchmark.
The training of the models had to be standardized, with all images preprocessed prior to being trained. Resizing, intensity normalization, augmentation, and conversion of categorical to binary labels are also mentioned in the manuscript. However, in the updated meaning these steps are regarded as components of a reproducible pipeline and not as notes written at the periphery. Its array was further split into training and validation subsets consisting of directories, and the mentioned performance metrics should then be seen as internal validation performance impacts instead of as the indicators of a completely independent external testing group. To be clearer, this 2,280-image collection will be taken as the official data on which the main analysis will be conducted. The binary framing was chosen because it reflects the early triage question common in practice: is the image of a suspicious lesion more likely to be mpox than other visually similar conditions? This framing is limited compared to an extensive dermatology diagnosis, but it is a valid and technically interpretable initial starting point for comparative benchmarking.
The updated description of the dataset also considers the fact that the paper includes some additional outputs obtained with the help of different class structures or some enriched subsets, which are shown in Figures 1A–I. Instead of deleting those materials, the manuscript now contextualizes them directly to allow the readers to view the results that are part of the binary benchmark and the ones that are part of secondary experiments. This method will maintain the entire documentation of this study, but will not allow the inclusion of any incompatible experiments to be integrated into a single assertion.