BrainMoE performance across disorders and modality states
The protocol generated five disease-specific healthy-versus-disease BrainMoE classifiers and produced prediction tables for full EEG-fMRI, EEG-only, and fMRI-only inference states. The full EEG-fMRI state produced consistently high discrimination across the five tasks, with AUC values ranging from 84.4 ± 3.2% for RI to 88.4 ± 3.8% for ASD (Table 1). The macro-averaged full-state performance across the five tasks reached 86.9 ± 3.0% AUC, 81.4 ± 3.1% accuracy, 81.4 ± 2.5% balanced accuracy, and 81.2 ± 3.0% F1 score.
Under the simulated missing-modality conditions, BrainMoE maintained usable performance when one input modality was masked. In the EEG-only state, the model achieved a macro-averaged AUC of 83.1 ± 3.7%, with the strongest EEG-only AUC observed for ADHD at 87.9 ± 2.8%. In the fMRI-only state, the macro-averaged AUC was 80.0 ± 3.9%. These performance ranges and the expected full-state performance advantage serve as practical benchmarks for successful implementation, indicating that the trained BrainMoE framework can perform full, EEG-only, and fMRI-only inference without separate models for each modality.
A representative suboptimal outcome is failure to reproduce the expected full-state performance advantage, for example, when the full EEG-fMRI AUC is lower than the EEG-only or fMRI-only AUC. In contrast, successful implementation should reproduce the full-state advantage and the benchmark performance ranges reported in Table 1. When a suboptimal pattern is observed, verify the H5 input dimensions, the DK region order, the modality-availability mask assignment, and the saved cross-validation partitions before interpreting the model output.
Benchmark comparison
The proposed BrainMoE model was compared with classical machine learning methods (support vector machine (SVM) and multilayer perceptron (MLP)), general deep learning methods (Transformer7, 3D-CNN22 and ResNet6), advanced deep learning methods (BrainNetCNN8, BNT10, BrainGNN9, MultiEpilepsyNet11, SZAtt-Net12), and MoE-based deep learning methods (dFCExpert17, EvoMoE18, and NeuroMoE++19) (Table 2). BrainMoE achieved the highest mean AUC among all methods compared, at 86.9 ± 3.0%.
Classical machine learning methods showed lower mean performance, with SVM achieving 66.7 ± 4.5% mean AUC and MLP 64.6 ± 6.3%. General deep learning methods showed variable performance, with ResNet achieving a mean AUC of 71.0 ± 3.3% and Transformer achieving 63.8 ± 4.7%. Among the advanced deep learning baselines, BNT, BrainGNN, MultiEpilepsyNet, and SZAtt-Net outperformed most classical and general deep learning methods, yet their mean AUCs remained lower than BrainMoE's.
To provide statistical support for the benchmark comparison, BrainMoE was compared with the strongest baseline for each performance metric (Table 3). BrainMoE outperformed NeuroMoE++ in AUC and balanced accuracy (BA) and outperformed SZAtt-Net in F1 score and accuracy. All comparisons remained significant after Holm adjustment.
Subgroup analysis by sex and acquisition site
Task-specific cohort characteristics, including the numbers of retained multimodal records, disease-to-control ratios, unique participant counts, age summaries, sex distributions, and acquisition-site distributions, are summarized in the cohort characteristics table (Table 4). The disease-to-control ratios are calculated from multimodal record counts, whereas demographic and acquisition-site characteristics are summarized at the unique-participant level.
To assess the potential effects of sex and acquisition site, BrainMoE performance was stratified by these factors (Table 5). The male subgroup showed higher mean AUC, BA, F1 score, and accuracy than the female subgroup, while the corresponding unadjusted Welch p values ranged from 0.089 to 0.321. Similarly, the RUBIC subgroup showed higher mean performance than the Staten Island subgroup, with p values ranging from 0.055 to 0.309. No statistically significant subgroup difference was detected in these analyses.
Ablation analysis of BrainMoE components
Ablation experiments were performed to assess the contribution of the modality availability mask, convolutional router, shared expert, and MoE expert fusion design (Table 6). Removing the mask embedding reduced the mean AUC to 79.2 ± 3.1%, and replacing the convolutional router with an MLP router reduced it to 79.3 ± 3.7%. Removing the shared expert reduced the fMRI-only AUC to 73.5 ± 3.3%, the lowest among the tested variants. Removing all MoE experts also reduced the overall balanced accuracy to 74.2 ± 3.0%. These ablation results showed that the full BrainMoE design achieved the strongest overall performance across both full- and missing-modality states, while the mask embedding, convolutional router, shared expert, and MoE expert fusion each contributed to the final model behavior.
ROI-level interpretability results for five brain disorders
To examine the regional contributions underlying BrainMoE predictions, node-occlusion attribution was performed on correctly classified disease-positive samples in the full EEG-fMRI state. EEG-derived and fMRI-derived ROI contributions were ranked separately by measuring the decrease in target-disease probability after occluding each DK atlas region. The analysis identified modality-specific contribution patterns across the five disease tasks (Figure 2). For fMRI-derived attribution, the top-ranked regions were the left posterior cingulate in MDD, the right pericalcarine cortex in ANX, the left parahippocampal cortex in RI, the right pars triangularis in ASD, and the left entorhinal cortex in ADHD. For EEG-derived attribution, the top-ranked regions were the left banks of the superior temporal sulcus in MDD, the left precentral cortex in ANX, the left cuneus in RI, the left lingual cortex in ASD, and the right insula in ADHD. The highest-ranked EEG- and fMRI-derived ROIs for each disorder are visualized on cortical surfaces in Figure 3. These results showed that BrainMoE provided ROI-level interpretability while preserving separate attribution profiles for EEG-derived and fMRI-derived representations. For correctly classified disease-positive samples, cross-fold stability was assessed using the Top-1 ROI frequency across 50 held-out folds (Table 7). The observed frequencies ranged from 36% to 68%, exceeding the theoretical random-selection reference of 1/68 (1.47%) and supporting interpretation based on relative rankings rather than absolute contribution magnitudes. Comparisons with previous neuroimaging findings were conducted post hoc and used only to contextualize the attribution results, not as independent validation.

Figure 1: Overview of the BrainMoE framework with adaptive EEG-fMRI fusion for computer-assisted brain disorder diagnosis. Source-aligned EEG and fMRI ROI features are processed by separate Graph Encoders. The resulting representations are passed to the EEG and fMRI Experts and, together with the Modality Availability Mask, to the shared Soft-Routing Module. The modality representations and routing representation are fused and processed by the Shared neural-state expert. The three expert outputs are then combined in the Fuse Module and passed to the Diagnosis Classification Head to produce the HC and target-disease probabilities. Please click here to view a larger version of this figure.

Figure 2: Node occlusion-based ROI contribution analysis. Top 10 EEG-derived and fMRI-derived ROI contributions were visualized for each disease task under the full EEG-fMRI inference state. Panels (A–E) show the fMRI-derived results for MDD, ANX, RI, ASD, and ADHD, respectively, and panels (F–J) show the corresponding EEG-derived results in the same order. Only the highest-ranked ROI was labeled in each plot. ROI contribution was defined as the decrease in target-disease probability after occluding the corresponding DK atlas region. Please click here to view a larger version of this figure.

Figure 3: Cortical ROI-level interpretability maps across five brain disorders. Panels (A–E) show MDD, ANX, RI, ASD, and ADHD, respectively. Each panel displays the highest-ranked EEG-derived ROI in red and the highest-ranked fMRI-derived ROI in orange on the DK cortical surface. Colors indicate modality rather than contribution magnitude; therefore, no quantitative color scale is applied. The prefixes lh and rh denote the left and right hemispheres, respectively. Please click here to view a larger version of this figure.
| Disease task | State | AUC | Accuracy | BA | F1 | Sensitivity | Specificity |
| MDD | Full | 88.0 ± 3.2% | 84.8 ± 3.7% | 82.0 ± 2.3% | 76.2 ± 4.1% | 72.7 ± 3.0% | 91.3 ± 4.7% |
| MDD | EEG-only | 83.4 ± 3.4% | 81.8 ± 4.5% | 77.9 ± 3.0% | 73.7 ± 3.9% | 68.2 ± 2.8% | 87.6 ± 4.2% |
| MDD | fMRI-only | 82.1 ± 3.3% | 78.5 ± 4.7% | 74.0 ± 4.1% | 71.8 ± 4.9% | 66.6 ± 3.5% | 81.3 ± 5.6% |
| ANX | Full | 85.5 ± 2.5% | 76.1 ± 2.4% | 78.2 ± 2.0% | 79.4 ± 3.3% | 78.1 ± 3.6% | 78.2 ± 3.1% |
| ANX | EEG-only | 81.7 ± 3.9% | 73.2 ± 3.2% | 75.6 ± 3.1% | 78.1 ± 3.9% | 76.6 ± 3.8% | 74.6 ± 4.0% |
| ANX | fMRI-only | 74.8 ± 4.1% | 74.4 ± 3.7% | 72.8 ± 3.6% | 76.4 ± 4.2% | 72.6 ± 3.9% | 73.0 ± 4.4% |
| RI | Full | 84.4 ± 3.2% | 76.5 ± 3.0% | 77.3 ± 2.9% | 76.7 ± 2.8% | 77.9 ± 3.1% | 76.7 ± 3.3% |
| RI | EEG-only | 80.2 ± 4.0% | 71.5 ± 4.4% | 73.4 ± 3.5% | 73.0 ± 3.4% | 73.8 ± 3.9% | 72.9 ± 3.3% |
| RI | fMRI-only | 81.3 ± 4.5% | 71.1 ± 3.9% | 67.9 ± 3.7% | 68.5 ± 3.2% | 69.4 ± 4.5% | 66.3 ± 4.2% |
| ASD | Full | 88.4 ± 3.8% | 79.5 ± 3.5% | 81.4 ± 3.0% | 81.0 ± 2.8% | 83.4 ± 2.9% | 79.4 ± 3.0% |
| ASD | EEG-only | 82.5 ± 4.3% | 77.3 ± 4.1% | 78.2 ± 3.7% | 78.6 ± 3.5% | 79.3 ± 3.4% | 77.1 ± 4.1% |
| ASD | fMRI-only | 81.3 ± 4.4% | 73.8 ± 3.7% | 74.4 ± 3.4% | 75.0 ± 3.2% | 76.1 ± 4.2% | 72.6 ± 4.5% |
| ADHD | Full | 88.2 ± 2.2% | 90.3 ± 2.7% | 88.4 ± 2.5% | 92.8 ± 1.9% | 87.0 ± 2.4% | 89.7 ± 2.5% |
| ADHD | EEG-only | 87.9 ± 2.8% | 88.1 ± 2.6% | 86.4 ± 3.1% | 90.2 ± 3.4% | 85.1 ± 3.1% | 87.7 ± 3.0% |
| ADHD | fMRI-only | 80.5 ± 3.2% | 85.7 ± 3.0% | 84.7 ± 3.5% | 87.9 ± 4.1% | 83.2 ± 3.6% | 86.2 ± 3.4% |
Table 1: BrainMoE classification performance across disease tasks and modality-availability states. Performance of BrainMoE for five healthy-versus-disease classification tasks under full EEG-fMRI, EEG-only, and fMRI-only inference states. Metrics are reported as mean ± standard deviation and include the area under the receiver operating characteristic curve (AUC), accuracy, balanced accuracy (BA), F1 score, sensitivity, and specificity.
| Method | Method group | Mean AUC | Mean BA | Mean F1 | Mean accuracy |
| SVM | Classical ML | 66.7 ± 4.5% | 70.6 ± 3.4% | 61.7 ± 5.8% | 70.7 ± 4.6% |
| MLP | Classical ML | 64.6 ± 6.3% | 63.5 ± 8.2% | 64.5 ± 6.1% | 64.2 ± 6.2% |
| Transformer | Generic DL | 63.8 ± 4.7% | 64.4 ± 4.8% | 65.2 ± 5.1% | 66.1 ± 4.9% |
| 3D-CNN | Generic DL | 65.3 ± 4.3% | 65.2 ± 4.0% | 64.8 ± 4.5% | 60.0 ± 4.2% |
| ResNet | Generic DL | 71.0 ± 3.3% | 70.2 ± 3.6% | 67.5 ± 3.8% | 69.3 ± 3.0% |
| BrainNetCNN | Advanced DL | 70.2 ± 3.6% | 70.9 ± 3.1% | 69.2 ± 3.6% | 71.2 ± 3.1% |
| BNT | Advanced DL | 73.6 ± 3.1% | 76.4 ± 2.7% | 74.7 ± 2.3% | 76.6 ± 2.6% |
| BrainGNN | Advanced DL | 72.4 ± 2.8% | 72.7 ± 2.3% | 71.0 ± 3.4% | 72.3 ± 2.5% |
| MultiEpilepsyNet | Advanced DL | 78.3 ± 3.7% | 75.2 ± 3.5% | 76.5 ± 3.6% | 77.2 ± 3.3% |
| SZAtt-Net | Advanced DL | 78.7 ± 3.2% | 76.1 ± 3.8% | 77.4 ± 3.9% | 78.1 ± 3.4% |
| dFCExpert | MoE-based DL | 80.1 ± 3.2% | 77.3 ± 2.9% | 76.7 ± 3.1% | 77.6 ± 2.9% |
| EvoMoE | MoE-based DL | 79.6 ± 3.6% | 76.2 ± 3.1% | 76.2 ± 3.3% | 76.5 ± 3.2% |
| NeuroMoE++ | MoE-based DL | 81.5 ± 2.8% | 77.9 ± 2.7% | 77.1 ± 3.4% | 78.0 ± 2.7% |
| BrainMoE (ours) | MoE-based DL | 86.9 ± 3.0% | 81.4 ± 2.5% | 81.2 ± 3.0% | 81.4 ± 3.1% |
Table 2: Mean classification performance of BrainMoE compared with classical machine learning methods, generic deep learning models, and advanced neuroimaging deep learning architectures. Results are aggregated across the evaluated disease-classification tasks and reported as mean ± standard deviation for AUC, BA, F1 score, and accuracy.
| Comparison | Mean AUC | Mean BA | Mean F1 | Mean accuracy |
| Strongest baseline | NeuroMoE++ | NeuroMoE++ | SZAtt-Net | SZAtt-Net |
| Strongest-baseline performance | 81.5 ± 2.8% | 77.9 ± 2.7% | 77.4 ± 3.9% | 78.1 ± 3.4% |
| BrainMoE | 86.9 ± 3.0% | 81.4 ± 2.5% | 81.2 ± 3.0% | 81.4 ± 3.1% |
| Difference | +5.4 | +3.5 | +3.8 | +3.3 |
| t-test p | 0.0006 | 0.0065 | 0.0187 | 0.0276 |
| Holm-adjusted p* | p < 0.01 | p < 0.05 | p < 0.05 | p < 0.05 |
| Cohen's dz | 1.63 | 1.11 | 0.91 | 0.83 |
Table 3: Comparison of BrainMoE with the strongest baseline for each performance metric. P values were calculated using two-sided paired t-tests and adjusted using the Holm procedure. Cohen's dz denotes the standardized paired difference.
| Cohort | Multimodal records (n) | Disease-to-control ratio | Unique participants (n) | Age range (years) | Sex (Male/Female) | Acquisition site (Staten Island/RUBIC) |
| HC | 115 | -- | 75 | 5.02–21.90 | 35/40 | 26/49 |
| MDD | 52 | 0.45:1 | 33 | 8.36–19.73 | 14/19 | 15/18 |
| ANX | 156 | 1.36:1 | 98 | 5.53–21.00 | 46/52 | 45/53 |
| RI | 141 | 1.23:1 | 85 | 5.75–19.66 | 47/38 | 39/46 |
| ASD | 82 | 0.71:1 | 51 | 5.66–19.79 | 45/6 | 27/24 |
| ADHD | 555 | 4.83:1 | 338 | 5.04–21.72 | 241/97 | 136/202 |
Table 4: Characteristics of the study cohorts used in the five disease-specific classification tasks. The table reports the number of retained multimodal records, disease-to-control ratios, unique participant counts, demographic characteristics, and acquisition-site distributions.
| Subgroup | Number | Mean AUC | Mean BA | Mean F1 | Mean accuracy |
| Sex |
| Male | 705 | 87.2 ± 3.7 | 83.4 ± 3.2 | 82.1 ± 3.5 | 82.5 ± 3.8 |
| Female | 396 | 85.6 ± 3.3 | 81.2 ± 4.0 | 79.7 ± 3.8 | 79.6 ± 3.4 |
| Difference | | +1.6 | +2.2 | +2.4 | +2.9 |
| Welch p | -- | 0.321 | 0.192 | 0.159 | 0.089 |
| Acquisition site |
| RUBIC | 674 | 87.6 ± 4.0 | 83.3 ± 3.1 | 82.3 ± 3.7 | 83.4 ± 3.5 |
| Staten Island | 427 | 85.2 ± 3.6 | 81.7 ± 3.7 | 79.4 ± 3.4 | 80.1 ± 3.7 |
| Difference | | +2.4 | +1.6 | +2.9 | +3.3 |
| Welch p | -- | 0.176 | 0.309 | 0.085 | 0.055 |
Table 5: Sex- and acquisition-site-stratified performance of BrainMoE. Results are reported as mean ± standard deviation across 10 repetitions of 5-fold cross-validation. The difference represents the first subgroup minus the second, and P values were obtained using two-sided Welch t-tests.
| Variant | Full AUC | EEG-only AUC | fMRI-only AUC | Mean AUC | Mean BA |
| w/o mask embedding | 83.2 ± 2.9% | 79.5 ± 3.9% | 76.4 ± 3.7% | 79.2 ± 3.1% | 76.1 ± 3.4% |
| w/o Conv Router | 84.6 ± 3.6% | 80.7 ± 3.5% | 78.2 ± 4.2% | 81.6 ± 3.3% | 75.7 ± 3.1% |
| MLP Router | 82.8 ± 2.7% | 80.1 ± 4.1% | 76.6 ± 4.4% | 79.3 ± 3.7% | 74.1 ± 3.8% |
| w/o Shared Expert | 83.3 ± 3.6% | 80.8 ± 4.0% | 73.5 ± 3.3% | 80.1 ± 3.6% | 75.5 ± 3.2% |
| w/o MoE Experts | 81.9 ± 3.3% | 78.7 ± 3.6% | 78.0 ± 3.8% | 80.8 ± 4.1% | 74.2 ± 3.0% |
| BrainMoE (ours) | 86.9 ± 3.0% | 83.1 ± 3.7% | 80.0 ± 3.9% | 86.9 ± 3.0% | 81.4 ± 2.5% |
Table 6: Ablation analysis of key BrainMoE components. Ablation results showing the contribution of the modality availability mask, convolutional router, shared expert, and MoE expert-fusion design. Each variant is evaluated under full EEG-fMRI, EEG-only, and fMRI-only inference conditions, with mean AUC and mean BA summarizing overall performance.
| Disease | EEG: top-ranked ROI | EEG: Top-1 frequency, n/N (%) | fMRI: top-ranked ROI | fMRI: Top-1 frequency, n/N (%) | Chance reference (%) |
| MDD | left banks of the superior temporal sulcus | 22/50 (44%) | left posterior cingulate | 25/50 (50%) | 1.47 |
| ANX | left precentral cortex | 18/50 (36%) | right pericalcarine cortex | 21/50 (42%) | 1.47 |
| RI | left cuneus | 26/50 (52%) | left parahippocampal cortex | 29/50 (58%) | 1.47 |
| ASD | left lingual cortex | 27/50 (54%) | right pars triangularis | 28/50 (56%) | 1.47 |
| ADHD | right insula | 31/50 (62%) | left entorhinal cortex | 34/50 (68%) | 1.47 |
Table 7: Cross-fold stability of the top-ranked EEG- and fMRI-derived ROIs. Top-1 frequency denotes the number and percentage of 50-fold-level analyses in which the reported ROI ranked first. The theoretical random-selection reference was 1/68 (1.47%).
Supplemental File 1: H5 input-checking script. Python script for verifying the required H5 input keys, data types, diagnostic labels, and DK-atlas-compatible input dimensions before BrainMoE training. Please click here to download this file.
Supplemental File 2: fMRI preprocessing script. C-PAC configuration and execution files used for fMRI preprocessing, including initial-volume removal, motion and distortion correction, registration and normalization, nuisance regression, temporal filtering, and spatial smoothing. Please click here to download this file.
Supplemental File 3: FreeSurfer processing scripts. Scripts for processing structural MRI data, coregistering the Desikan-Killiany cortical parcellation to native fMRI space, and extracting ROI-level fMRI signals. Please click here to download this file.
Supplemental File 4: EEG preprocessing scripts. MATLAB/EEGLAB scripts used for EEG preprocessing, including filtering, artifact-component identification and removal, and rereferencing. Please click here to download this file.
Supplemental File 5: MNE-Python source-localization and feature-extraction script. Python script for EEG source localization, DK-atlas ROI extraction, and generation of ROI-level EEG features used as BrainMoE inputs. Please click here to download this file.
Supplemental File 6: BrainMoE implementation code. Python code and configuration files for the BrainMoE architecture, graph encoders, modality-state handling, expert routing and fusion, model training, evaluation, and ablation variants. Please click here to download this file.
Supplemental File 7: Node-occlusion attribution code. Python code for modality-specific node-occlusion analysis, calculation of ROI contribution scores, ranking of EEG- and fMRI-derived ROIs, and generation of attribution outputs. Please click here to download this file.