$$\rightleftharpoonup{xx}$$
$$\longleftharp{xx}$$,
$$\longrightharp{xx}$$,
Sample study
The sample included 881 participants from Spain (N = 325), Mexico (N = 169), Guatemala (N = 227), and Chile (N = 160), all of whom were native Spanish speakers. The sample was divided into two groups: 451 in the reading disability (RD) group and 430 in the normally achieving readers (NAR) group. Children with special educational needs-those requiring support and specific educational attention due to sensory impairments, neurological issues, or other conditions-were excluded because these factors are typically used as exclusionary criteria for learning disabilities or severe behavioral disorders, either temporarily or for the duration of their schooling. Participants in the RD group were selected based on criteria such as an IQ equal to or greater than 80, a word reading time measurement percentile above the 75th percentile, a pseudoword reading time measurement percentile above the 75th percentile, or a pseudoword reading accuracy measurement percentile below the 25th percentile. Similarly, the NAR group was selected based on comparable criteria, with the addition of a PROLEC36 comprehension task percentile above the 25th percentile. All percentiles were determined according to participants' grade levels. The classification of participants into these groups relied on tasks of word reading and pseudoword reading, each comprising separate blocks. In both blocks, participants were instructed to read aloud the presented verbal stimuli as quickly as possible. The word reading block included 32 stimuli, while the pseudoword reading block contained 48 stimuli. The familiarity of the words was controlled using criteria outlined by Guzman and Jiménez37. Measures of hits, errors, and latency times were recorded, with latency times measured from the appearance of each item to the start of the student's vocal response. The word reading block demonstrated high internal consistency (α = .89), while the pseudoword reading block exhibited even greater internal consistency (α = .91). The distribution of students with RD and NAR across different grade levels was as follows: In the second grade, there were 30 NAR and 27 RD students. Transitioning to the third grade, 100 NARs were observed alongside 88 RD students. The fourth grade included 100 NAR and 116 RD students. In the fifth grade, there were 100 NAR and 116 RD students, while in the sixth grade, there were 100 NAR and 104 RD students. To assess age differences in the total sample and within each grade level, one-way ANOVA tests were conducted. The results revealed no significant differences, with all p-values exceeding 0.05. Similarly, the distribution of the gender variable was examined using the chi-square test, with all p values also exceeding 0.05.
Descriptive statistics
Table 1 presents the descriptive statistics (mean and standard deviation) for age and the assessment measures. The data are categorized by gender, distinguishing between normally achieving readers and those with reading disabilities across various countries. The results showed positive and statistically significant correlations among a large number of assessment measures (see Table 2). The table displays correlations among the measures, spanning from weak to strong. Correlations less than 0.3 are considered weak, while those greater than 0.5 are categorized as strong. Specific correlation values are provided alongside their respective strength classifications. For example, weak correlations (r < 0.3) included those of blending-segmentation (r = 0.354) and deletion-homophone comprehension (r = 0.270). Moderate correlations (0.3 ≤ r < 0.5) were observed between the grammatical structure-number (r = 0.463) and the functional words-grammatical structure (r = 0.512). Strong correlations (r ≥ 0.5) are evident for number-gender (r = 0.642) and picture RAN-Color RAN (r = 0.442).
Table 1: Descriptive statistics of the measures. This table provides the mean, standard deviation, and other descriptive statistics for the various measures included in the digital tool. AGE: Age; NAR: Normally Achieving Readers; RD: Reading disability ; SEG: Segmentation; BLN: Blending; ISO: Phoneme isolation; DEL: Phoneme deletion; HOM: Homophone comprehension; R&S: Root & Suffixes; GEN: Gender; NUM: Number; FW: Functional words; GRS: Grammatical Structure; EXT: Expositive Text ; NAT: Narrative Text ; VOC: Voicing of articulation; MOA: Manner of articulation; POA= Place of articulation; DIG: Digit RAN; LET: Letter Ran; COL: Color RAN; PIC: Picture RAN Please click here to download this Table.
Table 2: Correlation coefficients of the measures. This table shows the correlation coefficients between different measures of the digital tool, indicating the strength and direction of the relationships between these measures. SEG: Segmentation; BLN: Blending; ISO: Phoneme isolation; DEL: Phoneme deletion; HOM: Homophone comprehension; R&S: Root & Suffixes; GEN: Gender; NUM: Number; GRS: Grammatical Structure; FW: Functional Words; EXT: Expositive Text ; NAT: Narrative Text ; VOC: Voicing of articulation; MOA: Manner of articulation; POA= Place of articulation; DIG: Digit RAN; LET: Letter Ran; COL: Color RAN; PIC: Picture RAN. Please click here to download this Table.
Confirmatory factor analysis
CFA was conducted to assess the proposed factor structure of the multimedia battery. The model includes one second-order factor and six latent variables, each representing distinct modules of the multimedia battery. These constructs include phonological awareness (including phonemic segmentation, isolation, blending, and deletion tasks), morpho-orthographic module (involving root and suffixes and homophone comprehension tasks), syntax (comprising gender, number, grammatical structure, and functional word tasks), speech perception (involving voicing, manner of articulation, and place of articulation tasks), reading comprehension (encompassing expositive text and narrative text comprehension tasks), and rapid automated naming (letters, colors, digits, and pictures tasks from the RAN task). To ensure consistency across all tasks in the multimedia battery (ranging from 0 to 1), we used the proportion of maximum scaling (POMS) estimation. POMS scores were computed for the time-based tasks (roots and suffixes)38.
Statistical analyses and graphical presentations were performed using R version 4.3.139, employing the lavaan40, semTools41, and ggplot242 packages. Model fit was assessed using various goodness-of-fit indices. Although the chi-square goodness of fit was significant, χ2(df) = 632.01, p < .001, which suggests a discrepancy between the hypothesized model and the observed data, it is important to note that the chi-square test is sensitive to sample size43. Therefore, additional fit indices were considered.
The comparative fit index (CFI) yielded a value of .961, exceeding the commonly accepted threshold of .95 and indicating a good fit to the data. The root mean square error of approximation (RMSEA) was .038, indicating a close fit of the model to the data. The standardized root mean square residual (SRMR) was .034, which was below the recommended value of .05, suggesting an ideal fit. Factor loadings for the items ranged from .36 to .81, indicating significant loadings on their respective factors. The total average variance extracted (AVE) exceeded .50, suggesting adequate convergent validity, and the total composite reliability (CR) was above .80, indicating good internal consistency reliability.
In conclusion, the results of the CFA supported the proposed factor structure, demonstrating good model fit, convergent validity, and reliability. Please refer to the graphical representation of the model in Figure 1.

Figure 1: Confirmatory factor analysis. This figure illustrates the results of the confirmatory factor analysis performed on the digital tool, highlighting the factor loadings and the relationships between different cognitive processing measures. Abbreviations: SIC= Sicole-R; PA= Phonological awareness; SEG= Segmentation; BLN= Blending; ISO= Isolation; DEL= Deletion; MO= Morphological processing; HOM= Homophone comprehension; R&S= Root and suffixes; SYN= Syntactic processing; GEN= Gender; NUM= Number; GRS= Grammatical structure; FW= Functional words; RC= Reading comprehension; EXT= Expositive text; NAT= Narrative text; SP= Speech perception; VOC= Voicing of articulation; MOA= Manner of articulation; POA= Place of articulation; RAN= Rapid Automatized Naming; NUM= Number RAN; LET= Letter RAN; COL= Color RAN; PIC= Picture RAN. Please click here to view a larger version of this figure.
Measurement invariance analyses were used to test whether the factor structure of the multimedia battery was stable across genders. Testing for measurement invariance consists of a series of model comparisons that define increasingly stringent equality constraints44. We carried out the statistical analyses and followed the same fit model criteria of the previous CFA. We successively constrained parameters representing the configural, metric (loadings), scalar (intercepts), and strict (residuals) structures45. A poor fit in any of these models suggests that the aspect being constrained does not operate consistently for the different groups. The degree of invariance was determined jointly when χ2D was >0.05 and ΔCFI was <0.01. A summary of the results of the indices of the models and the differences between them are shown in Table 3.
Table 3: Invariance model fit indices. This table presents the fit indices for testing the gender invariance of the model, assessing how well the digital tool maintains consistency in its structure and measurement properties across male and female groups. These indices ensure the tool's reliability and validity by confirming its ability to measure cognitive processes consistently across diverse gender groups. Please click here to download this Table.
The configural model, which constrained only the relative configuration of variables in the model to be the same in both groups, had an adequate fit to the data: χ2(292) = 745.970, p<.001, CFI = .964, SRMR = 0.034, RMSEA = .036, 90% CI (.033, .038). The metric invariance model constrained the configuration of variables and all factor loadings to be constant across groups. The fit indices were comparable to those of the configural model: χ2(310) = 768.56, p<.001, CFI = .963, SRMR = .037, RMSEA = .036, 90% CI (.033, .038). The invariance of the factor loadings was supported by the nonsignificant difference tests that assessed model similarity: χ2D (18) = 26.30, p =.09; ΔCFI = .001. In the scalar invariance model, the configuration, factor loadings, and indicator means/intercepts were constrained to be the same for each group. The fit indices were less than ideal: χ2(322) = 787,50, p<.001, CFI = .963, SRMS = .038, RMSEA = .035 (90% CI = .032, .038). The difference tests that evaluated model similarity suggested that there was factorial invariance: χ2D(12)=13.00, p = .369; ΔCFI = .001. Finally, in the strict invariance model, the configuration, factor loadings, indicator means/intercepts, and residuals were constrained to be the same for each group. The fit indices were less than ideal: χ2(341) =798.20, p<.001, CFI = .961, SRMS = .039, RMSEA = .0035 (90% CI = .032, .038). Strict invariance was supported by the nonsignificant difference tests that assessed model similarity: χ2D(11) = 25.81, p<.001; ΔCFI =.002
Diagnostic accuracy
To evaluate the accuracy and discriminative capacity of the multimedia battery, receiver operating characteristic (ROC) analyses were conducted. ROC analysis aids in determining the optimal threshold (cutoff value) for a continuous-scale assessment test, balancing sensitivity (ability to correctly identify true positives) and specificity (ability to correctly identify true negatives). Additionally, ROC analysis was used to assess the ability of the test to discriminate between the RD and NAR groups. To carry out the analysis, Z scores were calculated per grade for each task of the multimedia battery. The omnibus score consisted of the sum of the zeta scores. Statistical analyses and graphical presentations were carried out using R version 4.3.141. The pROC46 and ggplot242 packages were used. In terms of diagnostic accuracy, the multimedia battery exhibited an area under the curve (AUC) of 9439.8 [95% CI: 93.31%-96.24% (DeLong)] and a sensitivity of 91.0 (Figure 2). The ROC curves by grade showed the following indices: 2° grade, AUC= 96.195% [CI: 91.54%-100% (DeLong)], Se= .96, Sp=.90; 3° grade, AUC= 95.3, [95% CI: 92.43%-98.18% (DeLong)], Se=75.70, Sp= 72.34; 4° grade, AUC=93.4 [95% CI: 90.2%-96.66% (DeLong)], Se=.92, Sp=.84; 5° grade, AUC=95.9 [95% CI: 93.11%-98.75% (DeLong)], Se=.90, Sp=.95; 6° grade, AUC=94.4 [95% CI: 91.11%-97.69% (DeLong)], Se=.92, Sp=.91.

Figure 2: Curve ROC analysis. This figure presents the ROC curve analysis, which shows the diagnostic accuracy of the digital tool by plotting the true positive rate against the false positive rate at various threshold settings. The units for the x-axis of the ROC curve are specificity, and the units for the y-axis are sensitivity. Abbreviations: ROC= receiver operating characteristic curve; AUC= area under the curve. Please click here to view a larger version of this figure.