Review Article

A Scoping Review of Machine Learning and Deep Learning Methods for Autism Spectrum Disorder Diagnosis and Analysis

July 17th, 2026

In This Article

Summary

Loading...
$$\rightleftharpoonup{xx}$$ $$\longleftharp{xx}$$, $$\longrightharp{xx}$$,

This review evaluates 50 recent studies applying machine learning and deep learning to autism spectrum disorder (ASD) diagnosis. Multimodal approaches integrating neuroimaging, behavioral, and biological data demonstrated improved performance. Key challenges include interpretability, scalability, and data heterogeneity. Future efforts should focus on explainable, privacy-preserving, and clinically translatable AI systems.

Abstract

Loading...
$$\rightleftharpoonup{xx}$$ $$\longleftharp{xx}$$, $$\longrightharp{xx}$$,

Autism spectrum disorder (ASD) is a heterogeneous neurodevelopmental condition characterized by diverse behavioral, cognitive, sensory, and communication profiles, making early diagnosis and personalized intervention challenging. Recent advances in machine learning (ML) and deep learning (DL) have enabled the development of computational tools for ASD screening, classification, severity assessment, and intervention monitoring. This review synthesizes findings from 50 recent studies that applied ML and DL techniques to ASD-related datasets, including electroencephalography (EEG), eye-tracking, behavioral video, microbiome, voice acoustic, demographic, and multimodal data. The review addresses three key questions: (i) which data modalities and computational approaches are most frequently used, (ii) how diagnostic performance is evaluated across different study designs, and (iii) what methodological challenges limit clinical translation. The literature is organized according to data modality, algorithmic approach, and clinical readiness. Approaches examined include conventional ML methods, convolutional neural networks, graph neural networks, hybrid deep learning architectures, federated learning, explainable artificial intelligence, topological data analysis, and multimodal fusion. The findings suggest that multimodal and graph-based approaches provide a more comprehensive representation of ASD phenotypes than single-modality methods. Explainability and privacy-preserving learning have also emerged as important considerations for clinical deployment. However, many reported high-performance models are based on small sample sizes, repeated use of the ABIDE dataset, class imbalance, single-site validation, or limited external testing, raising concerns regarding generalizability. Beyond diagnostic accuracy, this review evaluates model interpretability, calibration, scalability, validation rigor, and clinical applicability. Overall, the analysis highlights the need for standardized benchmarks, externally validated multimodal datasets, clinically relevant evaluation metrics, and decision-support systems that complement rather than replace expert clinical assessment in ASD diagnosis and management.

Introduction

Loading...
$$\rightleftharpoonup{xx}$$ $$\longleftharp{xx}$$, $$\longrightharp{xx}$$,

Autism spectrum disorder (ASD) is a complex neurodevelopmental condition characterized by impairments in social communication, restricted interests, and repetitive behaviors. The heterogeneous nature of ASD and its overlap with other neurodevelopmental disorders make diagnosis and clinical assessment particularly challenging. Traditional diagnostic approaches rely primarily on behavioral observations and clinical evaluations, which may contribute to delayed diagnosis and variability in clinical outcomes. Early and accurate identification of ASD is essential because timely intervention has been shown to improve cognitive, social, and adaptive functioning significantly. Recent advances in artificial intelligence (AI), particularly machine learning (ML) and deep learning (DL), have expanded the opportunities for developing objective and data-driven approaches to ASD diagnosis and analysis1,2,3.

ML and DL techniques enable the analysis of high-dimensional datasets derived from functional magnetic resonance imaging (fMRI), electroencephalography (EEG), eye-tracking, behavioral assessments, speech signals, and other biological markers. These approaches have demonstrated considerable potential for ASD screening, classification, severity prediction, and intervention monitoring. Recent studies have reported encouraging results using neuroimaging-based models, behavioral analytics, multimodal learning frameworks, and privacy-preserving architectures for ASD detection and assessment4,5,6. However, the rapidly expanding body of literature remains highly fragmented because many studies focus on individual data modalities, specific algorithms, or isolated clinical objectives, making direct comparison of methodologies and findings difficult.

Despite substantial progress, several challenges continue to limit the clinical translation of AI-based ASD diagnostic systems. Many published models are developed using relatively small datasets, single-site cohorts, or limited validation procedures, raising concerns regarding generalizability and real-world applicability. In addition, issues related to interpretability, computational efficiency, scalability, fairness, and data privacy remain important barriers to clinical adoption. Existing reviews have typically focused on specific methodologies or individual data sources, often overlooking the broader relationships among data modality, algorithmic design, validation rigor, and clinical readiness.

To address these gaps, this scoping review synthesizes evidence from 50 recent studies that applied machine learning (ML) and deep learning (DL) techniques to autism spectrum disorder (ASD) diagnosis and analysis. The reviewed literature encompasses a broad spectrum of computational approaches and data modalities, including neuroimaging-based methods using functional and structural magnetic resonance imaging (fMRI/sMRI), electroencephalography (EEG), behavioral and gaze analysis, physiological signal processing, speech and voice-acoustic assessment, microbiome-based biomarkers, and multimodal learning frameworks. Recent studies have explored conventional machine learning approaches such as support vector machines and random forests, as well as advanced deep learning architectures including convolutional neural networks, graph neural networks, transfer learning, attention-based models, federated learning frameworks, explainable artificial intelligence, and topological data analysis approaches1,2,3,4,5,6,7,8,9,10,11,12,13,14,15,16.

The reviewed studies further demonstrate the increasing diversity of ASD-related data sources and analytical objectives. Investigations have examined neuroanatomical heterogeneity, functional connectivity patterns, behavioral assessment, virtual reality–based screening, physiological monitoring, EEG analysis, electroretinogram-based classification, microbiome profiling, metabolomics, speech acoustics, and multimodal data fusion for diagnosis, severity prediction, intervention monitoring, and subtype identification17,18,19,20,21,22,23,24,25,26,27,28,29,30,31,32,33,34,35. These developments reflect a shift from single-modality diagnostic models toward integrative frameworks capable of capturing the heterogeneous clinical manifestations of ASD.

Beyond summarizing diagnostic performance, this review evaluates the relationship between data modality, algorithmic design, validation rigor, interpretability, scalability, and clinical applicability. Particular attention is given to emerging approaches such as federated learning, graph-based modeling, multimodal fusion, explainable artificial intelligence, privacy-preserving learning frameworks, extended reality technologies, and biomarker-driven precision diagnostics36,37,38,39,40,41,42,43,44,45,46,47,48,49,50. By integrating evidence across modalities, algorithms, and validation strategies, this review provides a structured benchmark of current ASD diagnostic technologies while highlighting methodological strengths, limitations, and evidence quality considerations. The analysis extends beyond reported accuracy to include interpretability, external validation, clinical readiness, and deployment feasibility, thereby offering a framework for evaluating the translational potential of AI-driven ASD diagnostic systems. Ultimately, this review aims to identify current research trends, highlight existing gaps, and outline future directions for the development of scalable, interpretable, and clinically translatable AI solutions for ASD assessment and management.

Review methodology

This study was conducted as a scoping review to provide a comprehensive overview of recent machine learning (ML) and deep learning (DL) approaches applied to autism spectrum disorder (ASD) diagnosis and analysis. A scoping review methodology was selected because the available literature is highly heterogeneous across data modalities, clinical objectives, algorithmic frameworks, datasets, validation strategies, and reported outcome measures.

The reviewed studies encompassed a wide range of data sources, including neuroimaging, electroencephalography (EEG), eye-tracking, behavioral assessments, speech signals, microbiome profiles, and multimodal datasets. Because substantial differences existed across study populations, reference standards, performance metrics, and experimental designs, a quantitative meta-analysis was deemed inappropriate. Instead, the objective of this review was to map the current research landscape, identify methodological trends, compare computational approaches, evaluate the quality of the evidence, and assess the clinical readiness of ML and DL systems for ASD diagnosis and monitoring.

The included studies were analyzed according to data modality, machine learning methodology, validation strategy, diagnostic performance, interpretability, scalability, and potential for clinical translation. Particular attention was given to emerging areas such as multimodal learning, explainable artificial intelligence, federated learning, graph-based modeling, and privacy-preserving diagnostic frameworks.

Research questions

This scoping review was guided by three overarching questions. First, it examined which ASD-related data modalities and machine learning/deep learning (ML/DL) approaches have been most frequently employed for screening, diagnosis, severity prediction, and intervention monitoring. Second, it evaluated how model performance, validation strategies, interpretability, and scalability have been reported across recent ASD artificial intelligence studies. Third, it investigated the methodological limitations, evidence-quality concerns, and barriers to clinical translation that continue to affect the adoption of ML/DL-based systems for ASD diagnosis and management. Collectively, these questions provide a framework for synthesizing current research trends, assessing the strengths and weaknesses of existing approaches, and identifying priorities for future development of clinically applicable AI solutions for ASD.

Search strategy and information sources

A literature search was conducted across IEEE Xplore, PubMed/MEDLINE, Scopus, Web of Science, ScienceDirect, SpringerLink, and Google Scholar. The search focused on studies published between 2022 and 2025 to capture recent advances in machine learning (ML), deep learning (DL), federated learning, graph neural networks, multimodal fusion, and explainable artificial intelligence (XAI) for autism spectrum disorder (ASD) diagnosis and analysis.

The search strategy combined terms related to ASD, artificial intelligence methodologies, diagnostic objectives, and data modalities. The primary search query included the terms: (“autism spectrum disorder” OR ASD OR autism) AND (“machine learning” OR “deep learning” OR “artificial intelligence” OR CNN OR GNN OR SVM OR “random forest” OR “federated learning” OR “explainable AI”) AND (diagnosis OR screening OR classification OR prediction OR severity OR intervention) AND (fMRI OR MRI OR EEG OR “eye tracking” OR video OR microbiome OR voice OR multimodal). The search syntax was adapted as necessary for individual databases.

To improve study coverage and reduce the risk of missing relevant publications, the reference lists of eligible review articles and primary research studies were manually screened for additional records. The final set of studies was selected based on the predefined eligibility criteria described below.

Eligibility criteria

Studies were selected according to predefined inclusion and exclusion criteria to ensure relevance to ASD-focused machine learning and deep learning applications, methodological quality, and clinical applicability. The eligibility criteria used for study selection are summarized in Table 1.

Screening procedure

Titles and abstracts of the retrieved records were initially screened to assess their relevance to autism spectrum disorder (ASD), artificial intelligence methodologies, and diagnostic or analytical objectives. Full-text articles were subsequently reviewed to confirm eligibility and to extract information regarding study populations, data sources, computational approaches, validation strategies, and reported outcome measures. Duplicate records and studies with overlapping content were excluded prior to the final synthesis. When multiple studies utilized the same dataset (e.g., ABIDE), only studies with distinct model architectures, feature representations, validation frameworks, or clinical objectives were retained to avoid redundancy and ensure a comprehensive evaluation of methodological diversity.

Data extraction and synthesis

Data were extracted from each included study using a structured approach. The extracted information included author and publication year, data modality, dataset or source, sample size when available, target population or age group, ASD-related task, algorithm or model type, feature extraction method, validation strategy, reported performance metrics, interpretability approach, external validation status, stated limitations, and clinical applicability. The studies were then synthesized according to data modality, algorithmic category, clinical purpose, validation design, and evidence quality rather than being ranked solely by reported performance.

Quality assessment and risk-of-bias appraisal

A methodological appraisal was conducted to evaluate potential bias, evidence quality, and clinical readiness. Each study was assessed based on sample size adequacy, class balance, reliance on repeated benchmark datasets such as ABIDE, risk of data leakage between preprocessing and validation stages, type of cross-validation or external validation, calibration reporting, demographic fairness across age and gender, interpretability, reproducibility, and deployment feasibility. Studies reporting very high accuracy were interpreted cautiously unless supported by adequate sample size, independent testing, leakage-control procedures, and transparent validation methods.

Performance metric reporting

The revised comparison included only metrics that were explicitly reported in the original studies. When a metric was unavailable, it was marked as NR, indicating “not reported,” rather than being estimated or approximated. Cross-study comparisons were interpreted qualitatively because the included studies differed substantially in dataset size, clinical objective, validation strategy, methodology, age group, and reference standard. Therefore, reported accuracy values were not treated as directly comparable indicators of model superiority. The updated tables distinguish predictive performance from evidence quality by including target population, dataset size, validation method, external validation status, and key limitations.

Review and Perspective

Loading...
$$\rightleftharpoonup{xx}$$ $$\longleftharp{xx}$$, $$\longrightharp{xx}$$,

Autism spectrum disorder (ASD) is a heterogeneous neurodevelopmental condition characterized by impairments in social communication and restricted or repetitive patterns of behavior1. Recent advances in machine learning (ML) and deep learning (DL) have enabled the analysis of diverse ASD-related data modalities, including neuroimaging, behavioral assessments, physiological signals, and biological markers. The following sections synthesize key developments in ML/DL-based ASD research according to data modality, methodological approach, and clinical application1.

Neuroimaging-based approaches

Functional magnetic resonance imaging (fMRI) and electroencephalography (EEG) are among the most widely used modalities for ASD classification using machine learning (ML) and deep learning (DL) approaches. Gupta et al.1 developed a quantized convolutional neural network (Q-CNN) model to classify fMRI data from the ABIDE-1 dataset, achieving an accuracy of 98%. The study employed tangent space embedding for feature extraction and integrated federated learning (FL) to preserve data privacy, demonstrating the potential of lightweight and privacy-preserving diagnostic frameworks for real-world applications. Similarly, Zhao et al.4introduced a cluster-based high-order functional connectivity network (Ho-FCN) framework that captured dynamic interactions among multiple brain regions and achieved an accuracy of 86.2%.

Explainability has emerged as an important consideration for the clinical adoption of neuroimaging-based diagnostic models. Jung et al.10 presented an explainability-guided region-of-interest (ROI) selection framework that identified high-order functional associations between brain regions and outperformed conventional low-order functional connectivity models. In addition, Han et al.9investigated a multimodal framework that combined EEG and eye-tracking (ET) data using stacked denoising autoencoders. By integrating complementary neurophysiological and behavioral information, the proposed approach achieved improved diagnostic performance compared with single-modality methods.

Behavioral and gaze analysis

Behavioral data, particularly patterns of eye movements and facial expressions, have been widely investigated as objective indicators of ASD-related traits. Artiran et al.3 developed a virtual reality (VR)-based assessment system to evaluate gaze behavior during simulated job interviews. Using advanced signal-processing and clustering techniques, the study provided insights into the social modulation of gaze among neurotypical and neurodivergent individuals. Similarly, Zhang et al.5 employed facial dynamics and few-shot learning techniques to classify ASD traits from Autism Diagnostic Observation Schedule (ADOS) interview videos, achieving an accuracy of 91.72%.

Early screening approaches have also demonstrated considerable potential for ASD identification. Liu et al.7 validated a vision-based Response-to-Instructions (RTI) protocol for toddlers and reported 95% agreement with clinical diagnoses. Collectively, these studies highlight the potential of machine learning-based behavioral analysis for improving the scalability, objectivity, and automation of ASD screening and assessment.

Multimodal and hybrid methods

The integration of heterogeneous data modalities enhances the robustness and generalizability of ASD classification models. Han et al.9 developed a multimodal framework that combined EEG and eye-tracking (ET) data, demonstrating improved diagnostic performance by leveraging complementary neurophysiological and behavioral information. Similarly, Xiao et al.14 utilized structural magnetic resonance imaging (MRI) to extract ASD-related multi-level flux features. Their margin-maximized norm-mixed representation learning framework addressed challenges associated with data heterogeneity and high-dimensional feature spaces, achieving an area under the curve (AUC) of 0.907.

Hybrid approaches that integrate biological and environmental information have also emerged as promising strategies for ASD diagnosis. Peralta-Marzal et al.24 analyzed gut microbiome profiles using recursive ensemble feature selection (REFS), identifying bacterial taxa associated with ASD and achieving an AUC of 0.816. Collectively, these studies highlight the potential of multimodal and hybrid frameworks to improve diagnostic accuracy by integrating complementary sources of information beyond conventional neuroimaging data alone.

Computational innovations

Several studies have focused on improving the computational efficiency, scalability, and robustness of ASD diagnostic models. Sravani and Kuppusamy12 developed an optimized deep convolutional neural network (DCNN) integrated with the Dipper-Throated Particle Swarm Optimization (DTPSO) algorithm, achieving an accuracy of 95.9%. The proposed framework enhanced feature optimization and classification performance, demonstrating the potential of hybrid optimization techniques for ASD detection.

Efforts have also been directed toward the development of privacy-preserving and resource-efficient diagnostic systems. Farooq et al.20 applied federated learning (FL) for ASD detection, reporting accuracies of 98% for children and 81% for adults. By enabling decentralized model training without direct data sharing, the approach demonstrated the feasibility of deploying AI-based ASD diagnostic systems in resource-constrained and privacy-sensitive environments.

Emerging technologies

Emerging technologies, including wearable devices, assistive systems, and immersive virtual environments, have expanded the scope of ASD assessment and intervention. Jayanthi et al.23proposed a machine learning-based monitoring system that utilized physiological sensor data to assess stress levels in individuals with ASD, highlighting the potential of wearable technologies to support mental health monitoring and personalized care. Similarly, Robles et al.6 developed a virtual reality (VR)-based platform for ASD screening and classification, demonstrating the value of immersive environments for behavioral assessment and diagnostic support.

Despite these advances, several challenges continue to limit the widespread clinical adoption of AI-based ASD diagnostic systems. As summarized in Table 2, issues related to data heterogeneity, model generalizability, and interpretability remain significant concerns. To address these limitations, Kunda et al.11 emphasized the importance of large multi-site datasets, while Fu et al.17proposed age-specific subtyping approaches to better capture the developmental heterogeneity of ASD. Furthermore, the successful translation of ML-based diagnostic tools into clinical practice will require standardized validation protocols, transparent reporting frameworks, and careful consideration of ethical, privacy, and regulatory requirements.

Behavioral and physiological signal analysis

Behavioral and physiological signal analyses have emerged as promising approaches for ASD diagnosis and screening. Jin et al.27 conducted a meta-analysis of home video-based machine learning techniques and reported pooled sensitivity and specificity values of 0.90 and 0.87, respectively, highlighting the potential of scalable and accessible approaches for early ASD detection. In another study, Manjur et al.29 utilized electroretinogram signals to differentiate among ASD, attention-deficit/hyperactivity disorder (ADHD), and typically developing individuals, demonstrating the utility of physiological biomarkers for neurodevelopmental disorder classification.

EEG-based diagnostic approaches also showed encouraging results. Singh and Kakkar38 proposed the Chronological Sewing Training Optimization–Deep Residual Network (CSTO-DRN) model for ASD detection using EEG signals, achieving an accuracy of 88.6% while addressing challenges related to noise reduction and feature optimization. Similarly, Radhakrishnan et al.41 developed a graph attention network framework that incorporated graph neural networks and attention mechanisms for EEG-based ASD classification, reporting an accuracy of 96.2%. Collectively, these studies demonstrate the growing potential of behavioral and physiological signal analysis for developing objective, non-invasive, and scalable ASD diagnostic tools.

Multimodal and integrative methods

Multimodal fusion techniques have significantly advanced ASD analysis by integrating heterogeneous sources of information. Wang et al.40developed the Federated Hypergraph Neural Network (FedHNN) framework for ASD diagnosis using multimodal data from distributed medical datasets. The proposed framework preserved data privacy while outperforming local models in diagnostic accuracy, demonstrating the feasibility of privacy-preserving and collaborative diagnostic systems.

Gut microbiome analysis has also emerged as a promising modality for ASD research. Olaguez-Gonzalez et al.49 applied machine learning techniques to microbiome composition data and achieved an accuracy of 94.7% using a limited number of bacterial strains. Extending this line of research, Su et al.42 incorporated multi-kingdom microbiota data and reported an area under the curve (AUC) of 0.91, identifying functional biomarkers associated with ASD, including ubiquinol-7 biosynthesis pathways.

In addition to microbiome-based approaches, biomarkers derived from speech and voice acoustics have been investigated as non-invasive diagnostic tools. Briend et al.48 analyzed acoustic features extracted from speech samples of children and achieved 91% classification accuracy in distinguishing individuals with ASD from typically developing controls. These findings highlight the potential of voice-based biomarkers to complement traditional diagnostic assessments and support more accessible ASD screening strategies.

Advancements in computational methods 

As summarized in Table 3, recent advances in computational methodologies have substantially improved ASD diagnosis and classification. Vidyadhari et al.46 developed a deep quantized neural network (DQNN) optimized using the Fractional Social Driving Training-Based Optimization (FSDTBO) algorithm and reported an accuracy of 90%. The proposed framework demonstrated the potential of optimization-enhanced deep learning models for improving diagnostic performance.

Similarly, Zhang et al.32 employed a topological data analysis (TDA) approach to extract features from regional homogeneity measures and achieved an accuracy of 96.5%, outperforming several conventional methodologies. The study highlighted the value of advanced feature representation techniques for capturing complex neurobiological patterns associated with ASD.

Hybrid machine learning frameworks have also been explored to improve model selection and classification performance. Alqaysi et al.39 evaluated 72 hybrid machine learning models within a fuzzy multicriteria decision-making framework that integrated feature selection and classification strategies. Their findings indicated that decision tree- and gradient boosting-based approaches were among the most stable classifiers, while effective feature selection played a critical role in optimizing overall model performance.

Clinical significance of multimodal ASD frameworks

Multimodal learning has emerged as an important strategy for improving both the diagnostic performance and clinical relevance of ASD classification systems. Because ASD manifests heterogeneously across individuals through differences in social communication, sensory processing, cognitive functioning, behavioral regulation, and biological characteristics, models trained on a single data modality may capture only a limited portion of the disorder's complexity. By integrating information from neuroimaging, electroencephalography (EEG), eye-tracking, behavioral video, microbiome profiles, voice acoustics, and demographic data, multimodal frameworks provide a more comprehensive representation of ASD-related characteristics.

Beyond improving classification accuracy, multimodal systems support a more phenotype-informed approach to diagnosis by identifying convergent evidence across multiple data sources while preserving clinically meaningful differences among individuals. Findings reported by Han et al.9, Wang et al.40, Su et al.42, Briend et al.48, and Olaguez-Gonzalez et al.49 collectively demonstrate the value of integrating complementary modalities to enhance diagnostic robustness, interpretability, and clinical applicability. Consequently, multimodal ASD frameworks should be viewed not only as tools for improving predictive performance but also as approaches that facilitate more personalized and clinically relevant diagnostic decision-making.

Emerging technologies and applications

Emerging technologies have expanded the scope of ASD diagnosis, intervention, and patient monitoring. Zhao et al.45 reviewed the application of extended reality (XR) technologies for ASD interventions and highlighted their potential to support cognitive, behavioral, and social skill development through immersive and interactive environments. Similarly, Nogay and Adeli47 investigated the influence of age and gender on ASD classification and demonstrated that incorporating demographic information could improve the performance of deep learning models.

Despite substantial progress, several challenges continue to limit the widespread clinical adoption of AI-based ASD diagnostic systems. Data heterogeneity, limited sample sizes, insufficient external validation, and concerns regarding model generalizability remain common limitations across many studies. To address these challenges, Quillet et al.34 explored metabolomic biomarkers associated with ASD, while Thapa et al.50 utilized retrospective clinical data to improve diagnostic classification and model robustness. These studies illustrate the importance of expanding data diversity and incorporating novel biomarkers to enhance model reliability and clinical relevance.

Overall, the reviewed literature demonstrates the transformative potential of machine learning and deep learning for ASD diagnosis and analysis across multiple data modalities. Advances in neuroimaging, behavioral assessment, multimodal learning, physiological signal analysis, microbiome research, and voice-based biomarkers have substantially improved diagnostic accuracy, scalability, and interpretability. Future research should focus on large-scale, multi-cohort validation, standardized evaluation frameworks, enhanced interpretability, and the integration of diverse data modalities to support the development of robust, clinically translatable, and personalized ASD diagnostic systems.

Revised comparative result analysis

The comparative analysis was revised to emphasize evidence quality, validation rigor, and clinical applicability rather than diagnostic accuracy alone. Because the reviewed studies differed substantially in datasets, age groups, target populations, validation strategies, and methodological objectives, direct numerical comparison of reported performance metrics was interpreted cautiously. Instead, the revised synthesis integrates predictive performance with validation maturity, external testing, interpretability, deployment feasibility, and clinical readiness.

Figure 1 summarizes the explicitly reported diagnostic performance metrics extracted from the reviewed studies, including accuracy, sensitivity, and specificity. Although many studies reported high diagnostic performance, substantial variability was observed across modalities and study designs. Several investigations achieved accuracies exceeding 95%, but sensitivity and specificity values were often less consistently reported, highlighting the importance of evaluating diagnostic performance using multiple complementary metrics rather than accuracy alone.

The relationship between reported accuracy and sensitivity is further illustrated in Figure 2. A positive association can be observed across studies; however, the figure also demonstrates that high accuracy does not always correspond to proportionally higher sensitivity. This finding is clinically important because ASD screening applications require the identification of as many true ASD cases as possible, making sensitivity a critical evaluation criterion alongside overall classification accuracy.

To assess the reliability of reported findings, Figure 3 presents the distribution of validation maturity and risk-of-bias signals across major modality groups. Neuroimaging-based studies, particularly those using fMRI and sMRI datasets, constituted the largest evidence group but frequently relied on repeated benchmark datasets and cross-validation frameworks. Federated and multimodal approaches demonstrated comparatively stronger validation maturity because they incorporated multi-site data sources and more diverse validation strategies. In contrast, XR-based intervention platforms and several behavioral approaches showed limited external validation, restricting current evidence for large-scale clinical adoption.

Figure 4 extends the analysis beyond predictive performance by summarizing modality-level clinical readiness. The synthesis incorporates six dimensions: dataset maturity, validation maturity, external validation, interpretability, deployment feasibility, and overall clinical readiness. Federated and multimodal AI approaches achieved the highest overall readiness scores because they demonstrated stronger validation practices and improved scalability. Behavioral video and gaze-based methods also showed promising readiness due to their relatively high interpretability and practical deployment potential. Conversely, modalities such as voice-based assessment and XR platforms exhibited lower readiness because of limited validation evidence and smaller study populations.

The relationship between clinical readiness and reported diagnostic performance is visualized in Figure 5. The figure demonstrates that very high accuracy values do not necessarily translate into strong clinical readiness. Several studies reporting accuracies above 95% remained positioned at moderate readiness levels because they relied on small datasets, single-site cohorts, or lacked independent external validation. In contrast, studies with slightly lower diagnostic performance frequently demonstrated stronger translational potential when accompanied by improved interpretability, broader validation, and clearer deployment pathways.

Methodological limitations identified across modality groups are summarized in Figure 6. Common concerns included small sample sizes, dependence on single-site datasets such as ABIDE, insufficient external validation, potential data leakage, limited calibration analysis, and uncertainty regarding deployment feasibility. The evidence map highlights that methodological rigor and validation quality remain major barriers to clinical translation despite encouraging diagnostic performance. Consequently, evidence credibility should be considered alongside reported accuracy when evaluating ASD diagnostic systems.

Extended result analysis

A broader synthesis was conducted to compare the influence of data modalities, algorithmic strategies, dataset characteristics, interpretability frameworks, computational efficiency, and emerging technologies on ASD diagnostic performance and clinical translation.

Figure 7 provides an integrated overview of factors associated with ASD diagnostic performance. The modality comparison indicates that multimodal approaches achieved the highest overall performance, followed by neuroimaging-based methods. Algorithm-level comparisons suggest that advanced architectures, including graph neural networks and federated learning frameworks, consistently demonstrated strong performance while simultaneously addressing challenges related to data heterogeneity and privacy preservation. Dataset-related analysis indicates that larger and multi-site datasets generally produced more stable and generalizable results than smaller cohorts. Furthermore, methods incorporating explainability mechanisms achieved improved clinical acceptance because they provided greater transparency in model decision-making.

The evidence-quality matrices presented in Table 4 and Table 5 reinforce these observations. Studies supported by larger datasets, stronger validation procedures, and external testing generally demonstrated more reliable performance than studies relying solely on internal validation. Across both tables, evidence quality increased when studies incorporated multi-site cohorts, independent testing populations, or federated validation frameworks. These findings suggest that future ASD diagnostic systems should prioritize validation rigor and reproducibility in addition to predictive accuracy.

Table 6 summarizes the relative importance of evaluation criteria for ASD diagnostic AI systems. Sensitivity, specificity, external validation, interpretability, fairness assessment, and deployment feasibility emerged as critical factors for clinical translation. While diagnostic accuracy remains important, the analysis demonstrates that no single metric is sufficient to establish clinical utility. Rather, a combination of performance, reliability, transparency, and validation evidence is required for meaningful adoption in healthcare settings.

The modality-level synthesis presented in Table 7 further illustrates differences in evidence maturity across ASD diagnostic approaches. Neuroimaging-based methods demonstrated strong diagnostic performance but frequently exhibited limitations related to dataset dependence and external validation. Behavioral and gaze-based approaches offered improved interpretability and deployment feasibility, whereas multimodal and federated learning frameworks showed the strongest overall balance between performance, validation maturity, and clinical readiness. Emerging approaches involving microbiome analysis, voice-acoustic biomarkers, and XR-based systems remain promising but require larger validation studies before widespread clinical implementation.

Collectively, these findings demonstrate that the future of ASD diagnostic AI is likely to depend on multimodal, externally validated, interpretable, and privacy-preserving frameworks rather than performance optimization alone. The synthesis emphasizes that clinical readiness is determined not only by predictive accuracy but also by validation quality, transparency, scalability, and real-world applicability.

Conclusions

Loading...
$$\rightleftharpoonup{xx}$$ $$\longleftharp{xx}$$, $$\longrightharp{xx}$$,

This review demonstrates the expanding role of machine learning (ML) and deep learning (DL) in autism spectrum disorder (ASD) diagnosis and analysis. Across the reviewed studies, computational approaches were applied to diverse data modalities, including functional magnetic resonance imaging (fMRI), electroencephalography (EEG), behavioral video, gut microbiome profiles, and voice-acoustic signals. Traditional ML classifiers, particularly support vector machines (SVMs) and random forests (RFs), remained widely used because of their interpretability and effectiveness on structured datasets. In parallel, advanced DL architectures, including convolutional neural networks (CNNs), graph neural networks (GNNs), and hybrid frameworks, demonstrated strong performance in high-dimensional neuroimaging and multimodal applications.

Among the reviewed studies, deep learning architectures reported by Toranjsimin et al.37 and Jain et al.33 achieved strong diagnostic performance on neuroimaging and physiological datasets. Federated learning frameworks proposed by Wang et al.40 and explainable AI approaches reported by Jung et al.10 and Qing et al.30 further highlighted the growing importance of privacy preservation, transparency, and clinical trust in ASD diagnostic systems. In addition, multimodal approaches integrating EEG, eye-tracking, microbiome, and behavioral data, as demonstrated by Han et al.9, Su et al.42, and Olaguez-Gonzalez et al.49, emphasized the value of combining complementary sources of information to address the heterogeneous nature of ASD.

The findings indicate that ML/DL systems should be viewed as clinical decision-support tools rather than stand-alone diagnostic replacements. Although many studies reported high classification performance, accuracy alone is insufficient for evaluating clinical utility. Sensitivity, specificity, calibration, fairness, interpretability, and external validation are equally important when assessing the translational potential of ASD diagnostic systems. The evidence synthesized in this review further demonstrates that strong reported performance does not necessarily indicate clinical readiness when validation procedures, external testing, and deployment feasibility remain limited.

Future research should prioritize externally validated multimodal datasets, standardized evaluation frameworks, interpretable model architectures, and privacy-preserving collaborative learning approaches. Emerging technologies, including extended reality platforms investigated by Zhao et al.45 and biomarker-driven precision approaches explored by Quillet et al.34, offer promising opportunities for personalized assessment and intervention. Ultimately, the successful integration of AI into routine ASD care will depend on rigorous validation, transparent reporting, fairness assessment, and the development of scalable systems that are accurate, interpretable, and clinically deployable.

Performance metrics graph; accuracy, sensitivity, specificity vs reference number, data analysis.
Figure 1: Comparative performance metrics across the reviewed ASD AI studies. The figure presents the reported accuracy, sensitivity, and specificity values for individual studies included in the review, organized by reference number. Please click here to view a larger version of this figure.

Accuracy vs Sensitivity scatter plot, graph analysis, comparing diagnostic test metrics
Figure 2: Relationship between reported accuracy and sensitivity in ASD AI studies. The scatter plot illustrates the association between reported diagnostic accuracy and sensitivity across the reviewed ASD machine learning and deep learning studies. Each marker represents an individual study for which both accuracy and sensitivity values were explicitly reported. Please click here to view a larger version of this figure.

Validation maturity chart; experimental setup comparison; risk-of-bias signal analysis in studies.
Figure 3: Validation maturity and risk-of-bias distribution across ASD AI modalities. The figure summarizes the validation maturity of studies across major ASD-related data modalities, including fMRI/sMRI, EEG/ERG, behavioral video/gaze/VR, microbiome and biomarker-based approaches, voice-acoustic methods, federated or multimodal AI frameworks, and XR-based intervention platforms. The horizontal position of each marker reflects the predominant validation maturity level, ranging from internal validation to cross-validation/repeated validation, multi-site or federated learning (FL) validation, and external or cross-cohort validation. Bubble size indicates the number of studies in each modality group. Marker shape indicates risk-of-bias signal. EV+ indicates stronger external or cross-cohort evidence; EV limited indicates absent or unclear external validation. Please click here to view a larger version of this figure.

Heatmap chart of data analysis illustrating synthesis scores for medical and AI technologies.
Figure 4: Clinical readiness synthesis matrix across ASD AI modalities. The heatmap summarizes modality-level synthesis scores for six dimensions relevant to clinical translation: dataset maturity, validation maturity, external validation, interpretability, deployment feasibility, and overall clinical readiness. Scores range from 1 to 3, where 1 indicates limited evidence, 2 indicates moderate evidence, and 3 indicates strong evidence. The matrix separates model performance from clinical maturity. A modality can report high accuracy while still receiving a moderate readiness score if evidence is single-site, externally unvalidated, or difficult to deploy. Please click here to view a larger version of this figure.

Clinical readiness vs. reported accuracy scatter plot showing various AI methodologies with different scores.
Figure 5: Clinical readiness synthesis score versus reported diagnostic accuracy across ASD AI studies. Each point represents an individual study positioned according to its clinical readiness synthesis score and explicitly reported accuracy. Colors indicate the primary data modality or methodological category, including fMRI/sMRI, behavioral video/gaze/VR, federated or multimodal AI, EEG/ERG, voice/acoustic analysis, and microbiome/biomarker/gene-based approaches. The horizontal dashed line denotes the 90% accuracy threshold and is included for visual reference. Only explicitly reported accuracy values are plotted; studies with NR accuracy are not represented on this axis. The figure shows that very high accuracy does not automatically imply clinical maturity when validation, external testing, and deployment evidence are weak. Please click here to view a larger version of this figure.

Risk factors rating matrix; heatmap; AI, neuroimaging, biomarkers; data, deployment challenges.
Figure 6: Evidence map of methodological risk signals across ASD AI modalities. The heatmap summarizes methodological risk levels across major ASD-related data modalities, including fMRI/sMRI, EEG/ERG, behavioral video and gaze analysis, microbiome-based approaches, voice-acoustic methods, federated/multimodal AI frameworks, and XR-based intervention platforms. Risk levels were categorized as low (1), moderate (2), or high (3) for sample size uncertainty, single-site or ABIDE reliance, lack of external validation, data leakage or overfitting concerns, limited fairness/calibration assessment, and deployment barriers. Please click here to view a larger version of this figure.

Modality and algorithm type vs accuracy charts; dataset size variance diagram; efficiency metrics.
Figure 7: Summary of factors influencing ASD AI diagnostic performance. The panels compare reported accuracy across data modalities, algorithm categories, dataset sizes, interpretability approaches, computational efficiency methods, and emerging AI techniques. Collectively, the results highlight the influence of multimodal learning, advanced deep learning architectures, dataset characteristics, and methodological innovations on ASD diagnostic performance. Please click here to view a larger version of this figure.

Table 1: Eligibility criteria used for study selection. This table summarizes the inclusion and exclusion criteria applied during literature screening and study selection. The criteria were designed to ensure relevance to autism spectrum disorder (ASD) diagnosis, and analysis using machine learning (ML) and deep learning (DL) approaches. Abbreviations: ASD = Autism Spectrum Disorder; ML = Machine Learning; DL = Deep Learning; AI = Artificial Intelligence; EEG = Electroencephalography; fMRI = Functional Magnetic Resonance Imaging; MRI = Magnetic Resonance Imaging. Please click here to download this Table.

Table 2: Comparative Analysis of Machine Learning and Deep Learning Approaches for ASD Diagnosis and Analysis (Studies [1]–[25]). This table summarizes studies [1]–[25], including the methods used, key findings, strengths, and limitations of machine learning and deep learning approaches for autism spectrum disorder (ASD) diagnosis and analysis. The studies encompass neuroimaging, EEG, eye-tracking, virtual reality, microbiome, wearable sensing, multimodal, and federated learning frameworks. Abbreviations: ASD = Autism Spectrum Disorder; ML = Machine Learning; DL = Deep Learning; Q-CNN = Quantized Convolutional Neural Network; FL = Federated Learning; ET = Eye Tracking; EEG = Electroencephalography; VR = Virtual Reality; Ho-FCN = High-Order Functional Connectivity Network; FC = Functional Connectivity; ADOS = Autism Diagnostic Observation Schedule; ROI = Region of Interest; CNN = Convolutional Neural Network; DTPSO = Dipper-Throated Particle Swarm Optimization; Grad-CAM = Gradient-Weighted Class Activation Mapping; sMRI = Structural Magnetic Resonance Imaging; RAGNN = Regional-Asymmetric Adaptive Graph Neural Network; rs-EEG = Resting-State Electroencephalography; AUC = Area Under the Curve; HRV = Heart Rate Variability; REFS = Recursive Ensemble Feature Selection; RF = Random Forest; SMOTE = Synthetic Minority Oversampling Technique; ABIDE = Autism Brain Imaging Data Exchange. Please click here to download this Table.

Table 3: Comparative Analysis of Emerging and Advanced ASD Diagnostic Frameworks (Studies [26]–[50]). This table summarizes studies [26]–[50], including the methods used, key findings, strengths, and limitations of advanced ASD diagnostic approaches. The studies encompass multimodal learning, biomarkers, EEG/fMRI analysis, microbiome profiling, federated learning, graph neural networks, extended reality (XR), and other emerging AI frameworks. Abbreviations: ASD = Autism Spectrum Disorder; ML = Machine Learning; MRI = Magnetic Resonance Imaging; fMRI = Functional Magnetic Resonance Imaging; EEG = Electroencephalography; ERG = Electroretinogram; ADHD = Attention-Deficit/Hyperactivity Disorder; TFS = Time-Frequency Spectrum Analysis; TDA = Topological Data Analysis; ReHo = Regional Homogeneity; DNN = Deep Neural Network; TD = Typically Developing; ISAA = Indian Scale for Assessment of Autism; DCD = Developmental Coordination Disorder; XWT = Cross Wavelet Transform; CSTO-DRN = Chronological Sewing Training Optimization–Deep Residual Network; MCDM = Multi-Criteria Decision-Making; FedHNN = Federated Hypergraph Neural Network; FL = Federated Learning; GNN = Graph Neural Network; AUC = Area Under the Curve; CNN = Convolutional Neural Network; LSTM = Long Short-Term Memory; XR = Extended Reality; MR = Mixed Reality; DQNN = Deep Quantized Neural Network; FSDTBO = Fractional Social Driving Training-Based Optimization; AUROC = Area Under the Receiver Operating Characteristic Curve; DSM = Diagnostic and Statistical Manual of Mental Disorders; ABIDE = Autism Brain Imaging Data Exchange. Please click here to download this Table.

Table 4: Revised Comparative Performance and Evidence Quality Matrix for Studies [1]–[25]. This table separates predictive performance from evidence quality by incorporating validation maturity, external validation status, clinical applicability, and major methodological limitations. The framework supports a more balanced assessment of study reliability and translational potential. Abbreviations: ASD = Autism Spectrum Disorder; AUC = Area Under the Curve; AUROC = Area Under the Receiver Operating Characteristic Curve; EEG = Electroencephalography; fMRI = Functional Magnetic Resonance Imaging; MRI = Magnetic Resonance Imaging; FL = Federated Learning; XAI = Explainable Artificial Intelligence; NR = Not Reported. Please click here to download this Table.

Table 5: Revised Comparative Performance and Evidence Quality Matrix for Studies [26]–[50]. This table extends the evidence-quality assessment to studies 26–50 and includes methodological limitations, validation rigor, interpretability, and deployment feasibility alongside reported diagnostic performance. Abbreviations: ASD = Autism Spectrum Disorder; ML = Machine Learning; DL = Deep Learning; EEG = Electroencephalography; ERG = Electroretinogram; AUC = Area Under the Curve; AUROC = Area Under the Receiver Operating Characteristic Curve; XR = Extended Reality; FL = Federated Learning; NR = Not Reported. Please click here to download this Table.

Table 6: Evaluation Criteria Priority for ASD Diagnostic AI Studies. This table ranks the relative importance of evaluation criteria for ASD diagnostic artificial intelligence systems, emphasizing the distinction between predictive performance and clinical applicability. The framework highlights factors necessary for reliable translation into real-world healthcare settings. Abbreviations: ASD = Autism Spectrum Disorder; AI = Artificial Intelligence; AUC = Area Under the Curve; AUROC = Area Under the Receiver Operating Characteristic Curve; XAI = Explainable Artificial Intelligence; ML = Machine Learning; DL = Deep Learning. Please click here to download this Table.

Table 7: Condensed Structured Synthesis Table Separating Performance from Evidence Quality. This table summarizes modality-level evidence across the reviewed studies, including reported diagnostic performance, validation maturity, clinical readiness, and major methodological limitations. The synthesis highlights differences in translational potential across ASD diagnostic approaches. Abbreviations: ASD = Autism Spectrum Disorder; fMRI = Functional Magnetic Resonance Imaging; sMRI = Structural Magnetic Resonance Imaging; EEG = Electroencephalography; ERG = Electroretinogram; VR = Virtual Reality; XR = Extended Reality; AI = Artificial Intelligence; AUC = Area Under the Curve; AUROC = Area Under the Receiver Operating Characteristic Curve; NR = Not Reported. Please click here to download this Table.

Disclosures

Loading...
$$\rightleftharpoonup{xx}$$ $$\longleftharp{xx}$$, $$\longrightharp{xx}$$,

The authors declare that they have no conflicts of interest related to this work.

Acknowledgements

Loading...
$$\rightleftharpoonup{xx}$$ $$\longleftharp{xx}$$, $$\longrightharp{xx}$$,

The authors would like to thank VIT-AP University for providing the academic environment and resources that supported this work. The authors also appreciate the valuable contributions of researchers whose studies were included in this review. This research received no external funding.

Reprints and Permissions

Request permission to reuse the text or figures of this JoVE article

Request Permission

Tags

EngineeringMachine learningDeep LearningAutism spectrum disorderMulti Modal DataNeuroimaging Analysis

Related Articles