$$\rightleftharpoonup{xx}$$
$$\longleftharp{xx}$$,
$$\longrightharp{xx}$$,
To review demonstrates the expanding role of machine learning (ML) and deep learning (DL) in autism spectrum disorder (ASD) diagnosis and analysis. Across the reviewed studies, computational approaches were applied to diverse data modalities, including functional magnetic resonance imaging (fMRI), electroencephalography (EEG), behavioral video, gut microbiome profiles, and voice-acoustic signals. Traditional ML classifiers, particularly support vector machines (SVMs) and random forests (RFs), remained widely used because of their interpretability and effectiveness on structured datasets. In parallel, advanced DL architectures, including convolutional neural networks (CNNs), graph neural networks (GNNs), and hybrid frameworks, demonstrated strong performance in high-dimensional neuroimaging and multimodal applications.
Among the reviewed studies, deep learning architectures reported by Toranjsimin et al.37 and Jain et al.33 achieved strong diagnostic performance on neuroimaging and physiological datasets. Federated learning frameworks proposed by Wang et al.40 and explainable AI approaches reported by Jung et al.10 and Qing et al.30 further highlighted the growing importance of privacy preservation, transparency, and clinical trust in ASD diagnostic systems. In addition, multimodal approaches integrating EEG, eye-tracking, microbiome, and behavioral data, as demonstrated by Han et al.9, Su et al.42, and Olaguez-Gonzalez et al.49, emphasized the value of combining complementary sources of information to address the heterogeneous nature of ASD.
The findings indicate that ML/DL systems should be viewed as clinical decision-support tools rather than stand-alone diagnostic replacements. Although many studies reported high classification performance, accuracy alone is insufficient for evaluating clinical utility. Sensitivity, specificity, calibration, fairness, interpretability, and external validation are equally important when assessing the translational potential of ASD diagnostic systems. The evidence synthesized in this review further demonstrates that strong reported performance does not necessarily indicate clinical readiness when validation procedures, external testing, and deployment feasibility remain limited.
Future research should prioritize externally validated multimodal datasets, standardized evaluation frameworks, interpretable model architectures, and privacy-preserving collaborative learning approaches. Emerging technologies, including extended reality platforms investigated by Zhao et al.45 and biomarker-driven precision approaches explored by Quillet et al.34, offer promising opportunities for personalized assessment and intervention. Ultimately, the successful integration of AI into routine ASD care will depend on rigorous validation, transparent reporting, fairness assessment, and the development of scalable systems that are accurate, interpretable, and clinically deployable.

Figure 1: Comparative performance metrics across the reviewed ASD AI studies. The figure presents the reported accuracy, sensitivity, and specificity values for individual studies included in the review, organized by reference number. Please click here to view a larger version of this figure.

Figure 2: Relationship between reported accuracy and sensitivity in ASD AI studies. The scatter plot illustrates the association between reported diagnostic accuracy and sensitivity across the reviewed ASD machine learning and deep learning studies. Each marker represents an individual study for which both accuracy and sensitivity values were explicitly reported. Please click here to view a larger version of this figure.

Figure 3: Validation maturity and risk-of-bias distribution across ASD AI modalities. The figure summarizes the validation maturity of studies across major ASD-related data modalities, including fMRI/sMRI, EEG/ERG, behavioral video/gaze/VR, microbiome and biomarker-based approaches, voice-acoustic methods, federated or multimodal AI frameworks, and XR-based intervention platforms. The horizontal position of each marker reflects the predominant validation maturity level, ranging from internal validation to cross-validation/repeated validation, multi-site or federated learning (FL) validation, and external or cross-cohort validation. Bubble size indicates the number of studies in each modality group. Marker shape indicates risk-of-bias signal. EV+ indicates stronger external or cross-cohort evidence; EV limited indicates absent or unclear external validation. Please click here to view a larger version of this figure.

Figure 4: Clinical readiness synthesis matrix across ASD AI modalities. The heatmap summarizes modality-level synthesis scores for six dimensions relevant to clinical translation: dataset maturity, validation maturity, external validation, interpretability, deployment feasibility, and overall clinical readiness. Scores range from 1 to 3, where 1 indicates limited evidence, 2 indicates moderate evidence, and 3 indicates strong evidence. The matrix separates model performance from clinical maturity. A modality can report high accuracy while still receiving a moderate readiness score if evidence is single-site, externally unvalidated, or difficult to deploy. Please click here to view a larger version of this figure.

Figure 5: Clinical readiness synthesis score versus reported diagnostic accuracy across ASD AI studies. Each point represents an individual study positioned according to its clinical readiness synthesis score and explicitly reported accuracy. Colors indicate the primary data modality or methodological category, including fMRI/sMRI, behavioral video/gaze/VR, federated or multimodal AI, EEG/ERG, voice/acoustic analysis, and microbiome/biomarker/gene-based approaches. The horizontal dashed line denotes the 90% accuracy threshold and is included for visual reference. Only explicitly reported accuracy values are plotted; studies with NR accuracy are not represented on this axis. The figure shows that very high accuracy does not automatically imply clinical maturity when validation, external testing, and deployment evidence are weak. Please click here to view a larger version of this figure.

Figure 6: Evidence map of methodological risk signals across ASD AI modalities. The heatmap summarizes methodological risk levels across major ASD-related data modalities, including fMRI/sMRI, EEG/ERG, behavioral video and gaze analysis, microbiome-based approaches, voice-acoustic methods, federated/multimodal AI frameworks, and XR-