$$\rightleftharpoonup{xx}$$
$$\longleftharp{xx}$$,
$$\longrightharp{xx}$$,
Autism spectrum disorder (ASD) is a heterogeneous neurodevelopmental condition characterized by impairments in social communication and restricted or repetitive patterns of behavior1. Recent advances in machine learning (ML) and deep learning (DL) have enabled the analysis of diverse ASD-related data modalities, including neuroimaging, behavioral assessments, physiological signals, and biological markers. The following sections synthesize key developments in ML/DL-based ASD research according to data modality, methodological approach, and clinical application1.
Neuroimaging-based approaches
Functional magnetic resonance imaging (fMRI) and electroencephalography (EEG) are among the most widely used modalities for ASD classification using machine learning (ML) and deep learning (DL) approaches. Gupta et al.1 developed a quantized convolutional neural network (Q-CNN) model to classify fMRI data from the ABIDE-1 dataset, achieving an accuracy of 98%. The study employed tangent space embedding for feature extraction and integrated federated learning (FL) to preserve data privacy, demonstrating the potential of lightweight and privacy-preserving diagnostic frameworks for real-world applications. Similarly, Zhao et al.4introduced a cluster-based high-order functional connectivity network (Ho-FCN) framework that captured dynamic interactions among multiple brain regions and achieved an accuracy of 86.2%.
Explainability has emerged as an important consideration for the clinical adoption of neuroimaging-based diagnostic models. Jung et al.10 presented an explainability-guided region-of-interest (ROI) selection framework that identified high-order functional associations between brain regions and outperformed conventional low-order functional connectivity models. In addition, Han et al.9investigated a multimodal framework that combined EEG and eye-tracking (ET) data using stacked denoising autoencoders. By integrating complementary neurophysiological and behavioral information, the proposed approach achieved improved diagnostic performance compared with single-modality methods.
Behavioral and gaze analysis
Behavioral data, particularly patterns of eye movements and facial expressions, have been widely investigated as objective indicators of ASD-related traits. Artiran et al.3 developed a virtual reality (VR)-based assessment system to evaluate gaze behavior during simulated job interviews. Using advanced signal-processing and clustering techniques, the study provided insights into the social modulation of gaze among neurotypical and neurodivergent individuals. Similarly, Zhang et al.5 employed facial dynamics and few-shot learning techniques to classify ASD traits from Autism Diagnostic Observation Schedule (ADOS) interview videos, achieving an accuracy of 91.72%.
Early screening approaches have also demonstrated considerable potential for ASD identification. Liu et al.7 validated a vision-based Response-to-Instructions (RTI) protocol for toddlers and reported 95% agreement with clinical diagnoses. Collectively, these studies highlight the potential of machine learning-based behavioral analysis for improving the scalability, objectivity, and automation of ASD screening and assessment.
Multimodal and hybrid methods
The integration of heterogeneous data modalities enhances the robustness and generalizability of ASD classification models. Han et al.9 developed a multimodal framework that combined EEG and eye-tracking (ET) data, demonstrating improved diagnostic performance by leveraging complementary neurophysiological and behavioral information. Similarly, Xiao et al.14 utilized structural magnetic resonance imaging (MRI) to extract ASD-related multi-level flux features. Their margin-maximized norm-mixed representation learning framework addressed challenges associated with data heterogeneity and high-dimensional feature spaces, achieving an area under the curve (AUC) of 0.907.
Hybrid approaches that integrate biological and environmental information have also emerged as promising strategies for ASD diagnosis. Peralta-Marzal et al.24 analyzed gut microbiome profiles using recursive ensemble feature selection (REFS), identifying bacterial taxa associated with ASD and achieving an AUC of 0.816. Collectively, these studies highlight the potential of multimodal and hybrid frameworks to improve diagnostic accuracy by integrating complementary sources of information beyond conventional neuroimaging data alone.
Computational innovations
Several studies have focused on improving the computational efficiency, scalability, and robustness of ASD diagnostic models. Sravani and Kuppusamy12 developed an optimized deep convolutional neural network (DCNN) integrated with the Dipper-Throated Particle Swarm Optimization (DTPSO) algorithm, achieving an accuracy of 95.9%. The proposed framework enhanced feature optimization and classification performance, demonstrating the potential of hybrid optimization techniques for ASD detection.
Efforts have also been directed toward the development of privacy-preserving and resource-efficient diagnostic systems. Farooq et al.20 applied federated learning (FL) for ASD detection, reporting accuracies of 98% for children and 81% for adults. By enabling decentralized model training without direct data sharing, the approach demonstrated the feasibility of deploying AI-based ASD diagnostic systems in resource-constrained and privacy-sensitive environments.
Emerging technologies
Emerging technologies, including wearable devices, assistive systems, and immersive virtual environments, have expanded the scope of ASD assessment and intervention. Jayanthi et al.23proposed a machine learning-based monitoring system that utilized physiological sensor data to assess stress levels in individuals with ASD, highlighting the potential of wearable technologies to support mental health monitoring and personalized care. Similarly, Robles et al.6 developed a virtual reality (VR)-based platform for ASD screening and classification, demonstrating the value of immersive environments for behavioral assessment and diagnostic support.
Despite these advances, several challenges continue to limit the widespread clinical adoption of AI-based ASD diagnostic systems. As summarized in Table 2, issues related to data heterogeneity, model generalizability, and interpretability remain significant concerns. To address these limitations, Kunda et al.11 emphasized the importance of large multi-site datasets, while Fu et al.17proposed age-specific subtyping approaches to better capture the developmental heterogeneity of ASD. Furthermore, the successful translation of ML-based diagnostic tools into clinical practice will require standardized validation protocols, transparent reporting frameworks, and careful consideration of ethical, privacy, and regulatory requirements.
Behavioral and physiological signal analysis
Behavioral and physiological signal analyses have emerged as promising approaches for ASD diagnosis and screening. Jin et al.27 conducted a meta-analysis of home video-based machine learning techniques and reported pooled sensitivity and specificity values of 0.90 and 0.87, respectively, highlighting the potential of scalable and accessible approaches for early ASD detection. In another study, Manjur et al.29 utilized electroretinogram signals to differentiate among ASD, attention-deficit/hyperactivity disorder (ADHD), and typically developing individuals, demonstrating the utility of physiological biomarkers for neurodevelopmental disorder classification.
EEG-based diagnostic approaches also showed encouraging results. Singh and Kakkar38 proposed the Chronological Sewing Training Optimization–Deep Residual Network (CSTO-DRN) model for ASD detection using EEG signals, achieving an accuracy of 88.6% while addressing challenges related to noise reduction and feature optimization. Similarly, Radhakrishnan et al.41 developed a graph attention network framework that incorporated graph neural networks and attention mechanisms for EEG-based ASD classification, reporting an accuracy of 96.2%. Collectively, these studies demonstrate the growing potential of behavioral and physiological signal analysis for developing objective, non-invasive, and scalable ASD diagnostic tools.
Multimodal and integrative methods
Multimodal fusion techniques have significantly advanced ASD analysis by integrating heterogeneous sources of information. Wang et al.40developed the Federated Hypergraph Neural Network (FedHNN) framework for ASD diagnosis using multimodal data from distributed medical datasets. The proposed framework preserved data privacy while outperforming local models in diagnostic accuracy, demonstrating the feasibility of privacy-preserving and collaborative diagnostic systems.
Gut microbiome analysis has also emerged as a promising modality for ASD research. Olaguez-Gonzalez et al.49 applied machine learning techniques to microbiome composition data and achieved an accuracy of 94.7% using a limited number of bacterial strains. Extending this line of research, Su et al.42 incorporated multi-kingdom microbiota data and reported an area under the curve (AUC) of 0.91, identifying functional biomarkers associated with ASD, including ubiquinol-7 biosynthesis pathways.
In addition to microbiome-based approaches, biomarkers derived from speech and voice acoustics have been investigated as non-invasive diagnostic tools. Briend et al.48 analyzed acoustic features extracted from speech samples of children and achieved 91% classification accuracy in distinguishing individuals with ASD from typically developing controls. These findings highlight the potential of voice-based biomarkers to complement traditional diagnostic assessments and support more accessible ASD screening strategies.
Advancements in computational methods
As summarized in Table 3, recent advances in computational methodologies have substantially improved ASD diagnosis and classification. Vidyadhari et al.46 developed a deep quantized neural network (DQNN) optimized using the Fractional Social Driving Training-Based Optimization (FSDTBO) algorithm and reported an accuracy of 90%. The proposed framework demonstrated the potential of optimization-enhanced deep learning models for improving diagnostic performance.
Similarly, Zhang et al.32 employed a topological data analysis (TDA) approach to extract features from regional homogeneity measures and achieved an accuracy of 96.5%, outperforming several conventional methodologies. The study highlighted the value of advanced feature representation techniques for capturing complex neurobiological patterns associated with ASD.
Hybrid machine learning frameworks have also been explored to improve model selection and classification performance. Alqaysi et al.39 evaluated 72 hybrid machine learning models within a fuzzy multicriteria decision-making framework that integrated feature selection and classification strategies. Their findings indicated that decision tree- and gradient boosting-based approaches were among the most stable classifiers, while effective feature selection played a critical role in optimizing overall model performance.
Clinical significance of multimodal ASD frameworks
Multimodal learning has emerged as an important strategy for improving both the diagnostic performance and clinical relevance of ASD classification systems. Because ASD manifests heterogeneously across individuals through differences in social communication, sensory processing, cognitive functioning, behavioral regulation, and biological characteristics, models trained on a single data modality may capture only a limited portion of the disorder's complexity. By integrating information from neuroimaging, electroencephalography (EEG), eye-tracking, behavioral video, microbiome profiles, voice acoustics, and demographic data, multimodal frameworks provide a more comprehensive representation of ASD-related characteristics.
Beyond improving classification accuracy, multimodal systems support a more phenotype-informed approach to diagnosis by identifying convergent evidence across multiple data sources while preserving clinically meaningful differences among individuals. Findings reported by Han et al.9, Wang et al.40, Su et al.42, Briend et al.48, and Olaguez-Gonzalez et al.49 collectively demonstrate the value of integrating complementary modalities to enhance diagnostic robustness, interpretability, and clinical applicability. Consequently, multimodal ASD frameworks should be viewed not only as tools for improving predictive performance but also as approaches that facilitate more personalized and clinically relevant diagnostic decision-making.
Emerging technologies and applications
Emerging technologies have expanded the scope of ASD diagnosis, intervention, and patient monitoring. Zhao et al.45 reviewed the application of extended reality (XR) technologies for ASD interventions and highlighted their potential to support cognitive, behavioral, and social skill development through immersive and interactive environments. Similarly, Nogay and Adeli47 investigated the influence of age and gender on ASD classification and demonstrated that incorporating demographic information could improve the performance of deep learning models.
Despite substantial progress, several challenges continue to limit the widespread clinical adoption of AI-based ASD diagnostic systems. Data heterogeneity, limited sample sizes, insufficient external validation, and concerns regarding model generalizability remain common limitations across many studies. To address these challenges, Quillet et al.34 explored metabolomic biomarkers associated with ASD, while Thapa et al.50 utilized retrospective clinical data to improve diagnostic classification and model robustness. These studies illustrate the importance of expanding data diversity and incorporating novel biomarkers to enhance model reliability and clinical relevance.
Overall, the reviewed literature demonstrates the transformative potential of machine learning and deep learning for ASD diagnosis and analysis across multiple data modalities. Advances in neuroimaging, behavioral assessment, multimodal learning, physiological signal analysis, microbiome research, and voice-based biomarkers have substantially improved diagnostic accuracy, scalability, and interpretability. Future research should focus on large-scale, multi-cohort validation, standardized evaluation frameworks, enhanced interpretability, and the integration of diverse data modalities to support the development of robust, clinically translatable, and personalized ASD diagnostic systems.
Revised comparative result analysis
The comparative analysis was revised to emphasize evidence quality, validation rigor, and clinical applicability rather than diagnostic accuracy alone. Because the reviewed studies differed substantially in datasets, age groups, target populations, validation strategies, and methodological objectives, direct numerical comparison of reported performance metrics was interpreted cautiously. Instead, the revised synthesis integrates predictive performance with validation maturity, external testing, interpretability, deployment feasibility, and clinical readiness.
Figure 1 summarizes the explicitly reported diagnostic performance metrics extracted from the reviewed studies, including accuracy, sensitivity, and specificity. Although many studies reported high diagnostic performance, substantial variability was observed across modalities and study designs. Several investigations achieved accuracies exceeding 95%, but sensitivity and specificity values were often less consistently reported, highlighting the importance of evaluating diagnostic performance using multiple complementary metrics rather than accuracy alone.
The relationship between reported accuracy and sensitivity is further illustrated in Figure 2. A positive association can be observed across studies; however, the figure also demonstrates that high accuracy does not always correspond to proportionally higher sensitivity. This finding is clinically important because ASD screening applications require the identification of as many true ASD cases as possible, making sensitivity a critical evaluation criterion alongside overall classification accuracy.
To assess the reliability of reported findings, Figure 3 presents the distribution of validation maturity and risk-of-bias signals across major modality groups. Neuroimaging-based studies, particularly those using fMRI and sMRI datasets, constituted the largest evidence group but frequently relied on repeated benchmark datasets and cross-validation frameworks. Federated and multimodal approaches demonstrated comparatively stronger validation maturity because they incorporated multi-site data sources and more diverse validation strategies. In contrast, XR-based intervention platforms and several behavioral approaches showed limited external validation, restricting current evidence for large-scale clinical adoption.
Figure 4 extends the analysis beyond predictive performance by summarizing modality-level clinical readiness. The synthesis incorporates six dimensions: dataset maturity, validation maturity, external validation, interpretability, deployment feasibility, and overall clinical readiness. Federated and multimodal AI approaches achieved the highest overall readiness scores because they demonstrated stronger validation practices and improved scalability. Behavioral video and gaze-based methods also showed promising readiness due to their relatively high interpretability and practical deployment potential. Conversely, modalities such as voice-based assessment and XR platforms exhibited lower readiness because of limited validation evidence and smaller study populations.
The relationship between clinical readiness and reported diagnostic performance is visualized in Figure 5. The figure demonstrates that very high accuracy values do not necessarily translate into strong clinical readiness. Several studies reporting accuracies above 95% remained positioned at moderate readiness levels because they relied on small datasets, single-site cohorts, or lacked independent external validation. In contrast, studies with slightly lower diagnostic performance frequently demonstrated stronger translational potential when accompanied by improved interpretability, broader validation, and clearer deployment pathways.
Methodological limitations identified across modality groups are summarized in Figure 6. Common concerns included small sample sizes, dependence on single-site datasets such as ABIDE, insufficient external validation, potential data leakage, limited calibration analysis, and uncertainty regarding deployment feasibility. The evidence map highlights that methodological rigor and validation quality remain major barriers to clinical translation despite encouraging diagnostic performance. Consequently, evidence credibility should be considered alongside reported accuracy when evaluating ASD diagnostic systems.
Extended result analysis
A broader synthesis was conducted to compare the influence of data modalities, algorithmic strategies, dataset characteristics, interpretability frameworks, computational efficiency, and emerging technologies on ASD diagnostic performance and clinical translation.
Figure 7 provides an integrated overview of factors associated with ASD diagnostic performance. The modality comparison indicates that multimodal approaches achieved the highest overall performance, followed by neuroimaging-based methods. Algorithm-level comparisons suggest that advanced architectures, including graph neural networks and federated learning frameworks, consistently demonstrated strong performance while simultaneously addressing challenges related to data heterogeneity and privacy preservation. Dataset-related analysis indicates that larger and multi-site datasets generally produced more stable and generalizable results than smaller cohorts. Furthermore, methods incorporating explainability mechanisms achieved improved clinical acceptance because they provided greater transparency in model decision-making.
The evidence-quality matrices presented in Table 4 and Table 5 reinforce these observations. Studies supported by larger datasets, stronger validation procedures, and external testing generally demonstrated more reliable performance than studies relying solely on internal validation. Across both tables, evidence quality increased when studies incorporated multi-site cohorts, independent testing populations, or federated validation frameworks. These findings suggest that future ASD diagnostic systems should prioritize validation rigor and reproducibility in addition to predictive accuracy.
Table 6 summarizes the relative importance of evaluation criteria for ASD diagnostic AI systems. Sensitivity, specificity, external validation, interpretability, fairness assessment, and deployment feasibility emerged as critical factors for clinical translation. While diagnostic accuracy remains important, the analysis demonstrates that no single metric is sufficient to establish clinical utility. Rather, a combination of performance, reliability, transparency, and validation evidence is required for meaningful adoption in healthcare settings.
The modality-level synthesis presented in Table 7 further illustrates differences in evidence maturity across ASD diagnostic approaches. Neuroimaging-based methods demonstrated strong diagnostic performance but frequently exhibited limitations related to dataset dependence and external validation. Behavioral and gaze-based approaches offered improved interpretability and deployment feasibility, whereas multimodal and federated learning frameworks showed the strongest overall balance between performance, validation maturity, and clinical readiness. Emerging approaches involving microbiome analysis, voice-acoustic biomarkers, and XR-based systems remain promising but require larger validation studies before widespread clinical implementation.
Collectively, these findings demonstrate that the future of ASD diagnostic AI is likely to depend on multimodal, externally validated, interpretable, and privacy-preserving frameworks rather than performance optimization alone. The synthesis emphasizes that clinical readiness is determined not only by predictive accuracy but also by validation quality, transparency, scalability, and real-world applicability.