Participant flow and baseline characteristics
An initial cohort of 186 students from six intact classes was recruited for the study. Following eligibility screening, attendance verification, and response-quality checks, 174 students were retained for the final quantitative analysis. These participants were evenly distributed, with 87 assigned to the AI-assisted student-centered instruction group and 87 assigned to the usual student-centered instruction group. Figure 1 details the recruitment flow, class assignment, three-wave assessment timeline, exclusion criteria, and qualitative subsampling.
Table 2 summarizes the descriptive primary outcome scores by instructional group and measurement wave. Before the intervention, academic self-concept levels were similar between the two cohorts (AI-assisted: 3.00 ± 0.37; comparison: 3.05 ± 0.43). However, the AI-assisted group had lower baseline communication competence (2.85 ± 0.40) than the usual instruction group (3.04 ± 0.38). Because assignment occurred at the intact-class level rather than through individual randomization, these baseline differences were treated as potential pre-existing class differences rather than intervention effects. Baseline outcome scores were retained as covariates in all primary models, and communication competence findings were interpreted with attention to the risk of regression to the mean. To address the hierarchical data structure, unconditional intraclass correlation coefficients (ICC[1]s) were estimated from one-way random-effects ANOVA models using class identifiers as the grouping variable. The formula used was:
(1)
where MSbetween represents the mean square between classes, MSwithin represents the mean square within classes, and m denotes the class size. The final dataset included six classes of equal size (k = 6; m = 29). ICCs for primary observed scores ranged from 0.007 to 0.105, and ICCs for primary change scores were ASC T2–T1 = 0.152, ASC T3–T1 = 0.160, CC T2–T1 = 0.105, and CC T3–T1 = 0.165. These values indicated non-negligible classroom-level dependence. Because only six level-2 clusters were available, fitting a full multilevel model was considered unstable; therefore, ICCs, class-level summaries, and cluster-robust sensitivity checks were used as diagnostic evidence rather than as confirmatory multilevel estimates.
Implementation monitoring records were reviewed alongside the analytic dataset. All six classes completed the six-session unit, and attendance exposure was comparable across conditions (mean attendance rate: AI-assisted = 0.934; usual instruction = 0.936). No retained participant fell below the prespecified exposure threshold. To monitor implementation fidelity, session-level checklists and technical incident logs were utilized, with a second independent observer evaluating at least 20% of the sessions. These observations served as a formative, on-site quality assurance mechanism for providing immediate feedback to the instructor, rather than a summative measurement. Because the primary goal was real-time pedagogical correction, aggregate statistical metrics (such as mean fidelity scores or inter-rater reliability coefficients) were not permanently archived for retrospective analysis. Nevertheless, these formative checks confirmed that no major protocol deviations occurred, and the core instructional sequence was consistently maintained.
Internal consistency of the study measures
Table 3 reports the internal consistency estimates for all measured constructs using the item counts defined in the locked scoring map. Academic self-concept demonstrated acceptable reliability during both the post-intervention and follow-up phases, yielding Cronbach’s alpha values of 0.828 and 0.876, respectively. Communication competence similarly maintained acceptable reliability across these later assessment waves, with alpha values of 0.877 and 0.897.
Baseline reliability estimates for academic self-concept (α = 0.676) and communication competence (α = 0.683) were below the preferred 0.70 threshold. However, these estimates were retained to preserve the prespecified scoring structure, but they indicate greater baseline measurement error and may attenuate or destabilize baseline-adjusted estimates. The post-intervention process-related scales exhibited acceptable internal consistency, with alpha values of 0.760 for interaction quality, 0.808 for self-evaluation, and 0.806 for learner engagement.
Academic self-concept outcomes
Both groups experienced an increase in academic self-concept over time, although the AI-assisted group demonstrated a distinctly larger gain. From baseline to the immediate post-intervention assessment, the AI-assisted group's mean score rose from 3.00 ± 0.37 to 3.51 ± 0.48. In contrast, the comparison group showed a modest increase from 3.05 ± 0.43 to 3.19 ± 0.60. This resulted in a mean T2–T1 change of 0.51 ± 0.44 for the AI-assisted cohort and 0.14 ± 0.46 for the comparison cohort. The resulting between-group difference in the change score was 0.37 (95% CI [0.24, 0.50]), representing a statistically significant effect with a large magnitude (Welch t = 5.28, p < 0.001, Cohen’s d = 0.80).
This advantage was sustained at the follow-up assessment. The academic self-concept score reached 3.57 ± 0.60 in the AI-assisted group, compared with 3.20 ± 0.63 in the usual instruction group. The mean T3–T1 change was 0.57 ± 0.54 for the AI-assisted group and 0.15 ± 0.48 for the usual instruction group. The corresponding between-group difference in change increased slightly to 0.42 (95% CI [0.27, 0.57]), maintaining strong statistical significance (Welch t = 5.46, p < 0.001, Cohen’s d = 0.83). These three-wave longitudinal trajectories are shown in Figure 2A, while the complete change-score comparisons are presented in Table 4.
Table 5 presents the baseline-adjusted ANCOVA results. After controlling for baseline scores and relevant covariates, instructional group assignment was associated with higher post-intervention academic self-concept (B = 0.328, 95% CI [0.193, 0.463], SE = 0.069, t = 4.77, p < 0.001, model R2 = 0.395). This association was also observed at follow-up (B = 0.401, 95% CI [0.246, 0.556], SE = 0.079, t = 5.07, p < 0.001, model R2 = 0.384). Baseline academic self-concept was a significant covariate in both analytical models.
Communication competence outcomes
Students in the AI-assisted condition also exhibited a more pronounced improvement in communication competence. Between the baseline and post-intervention measurements, the AI-assisted group improved from 2.85 ± 0.40 to 3.22 ± 0.63, whereas the comparison group showed a modest increase from 3.04 ± 0.38 to 3.12 ± 0.62. The mean T2–T1 change was 0.37 ± 0.45 for the AI-assisted participants and 0.08 ± 0.46 for the comparison participants. This yielded a statistically significant between-group difference in change of 0.29 (95% CI [0.15, 0.43]), corresponding to a medium-to-large effect size (Welch t = 4.14, p < 0.001, Cohen’s d = 0.63).
During the follow-up phase, communication competence scores increased to 3.41 ± 0.71 in the AI-assisted group and 3.12 ± 0.69 in the comparison group. The mean T3–T1 change was recorded as 0.57 ± 0.56 in the intervention condition versus 0.09 ± 0.57 in the comparison condition. The difference in these change scores was 0.48 (95% CI [0.31, 0.65]), representing a statistically significant outcome with large effect size (Welch t = 5.61, p < 0.001, Cohen’s d = 0.85). Figure 2B illustrates these trajectories, and Table 4 provides the comprehensive change-score data.
Consistent with the raw change scores, the baseline-adjusted ANCOVA showed that group assignment was associated with post-intervention communication competence (B = 0.301, 95% CI [0.158, 0.444], SE = 0.073, t = 4.14, p < 0.001, model R2 = 0.488). This association persisted at follow-up (B = 0.516, 95% CI [0.340, 0.692], SE = 0.090, t = 5.72, p < 0.001, model R2 = 0.388). Baseline communication competence remained as a significant predictor in both adjusted models, as detailed in Table 5.
Exploratory associations with interaction quality, self-evaluation, and learner engagement
Post-intervention measures revealed that interaction quality, self-evaluation, and learner engagement were higher in the AI-assisted group. Linear modeling indicated that group assignment was positively associated with interaction quality (B = 0.344, SE = 0.075, t = 4.62, p < 0.001, model R2 = 0.133), self-evaluation (B = 0.291, SE = 0.082, t = 3.57, p < 0.001, model R2 = 0.118), and learner engagement (B = 0.361, SE = 0.081, t = 4.43, p < 0.001, model R2 = 0.129).
Table 6 summarizes exploratory regression models that examined process-related associations with outcome changes. When interaction quality, self-evaluation, and learner engagement were entered simultaneously alongside group assignment, the instructional group remained associated with changes in academic self-concept (B = 0.299, SE = 0.078, t = 3.85, p < 0.001, model R2 = 0.191). However, interaction quality showed no independent association with academic self-concept change in this model (B = 0.014, SE = 0.073, t = -0.20, p = 0.844). Both self-evaluation (B = 0.108, SE = 0.065, t = 1.66, p = 0.099) and learner engagement (B = 0.061, SE = 0.066, t = 0.93, p = 0.355) exhibited positive but statistically nonsignificant trends. These models were not interpreted as formal mediation tests.
A parallel pattern emerged for the changes in communication competence. Group assignment remained significant in the multipredictor model (B = 0.252, SE = 0.082, t = 3.09, p = 0.002, model R2 = 0.141). Interaction quality (B = 0.131, SE = 0.075, t = 1.75, p = 0.082) and self-evaluation (B = 0.111, SE = 0.066, t = 1.68, p = 0.094) displayed positive but nonsignificant associations. Learner engagement was similarly nonsignificant (B = -0.083, SE = 0.066, t = -1.25, p = 0.214).
Qualitative coding and mixed-methods integration
To complement the quantitative findings, a purposive subsample of 36 students underwent qualitative analysis. Coded themes were structured around four primary domains: interaction quality, self-evaluation, learner engagement, and implementation barriers. Figure 3A shows the distribution of these themes by instructional condition.
Within the AI-assisted cohort, the most frequent coded themes were AI-supported peer explanation (15 coded cases), active questioning (14), checking AI-supported information (14), gap awareness (13), revision reasoning (12), and task persistence (11). Implementation barriers were less frequent but still observed, including shallow use (5) and technical barriers (4). In the comparison group, AI-supported peer explanation was absent by design, whereas peer explanation related to the matched prompt routine appeared in 10 coded cases; active questioning appeared in 9 cases, gap awareness and task persistence in 8 cases each, revision reasoning in 7 cases, and shallow use in 2 cases. Technical barriers and checking AI-supported information were coded as 0 in the comparison condition. These counts describe the number of coded cases in the interview subsample rather than population.
The integration of quantitative and qualitative data (Figure 3B) revealed convergence in two areas. First, the higher interaction quality scores recorded in the AI-assisted group corresponded with the frequent qualitative codes for AI-supported peer explanation and checking AI-supported information. Second, higher self-evaluation scores in the intervention group aligned with qualitative narratives emphasizing gap awareness and revision reasoning. The qualitative themes also expanded the engagement findings by showing that active participation depended on perceived task authenticity, teacher moderation, and avoidance of shallow AI use.
Summary of findings
In summary, the integrated results suggest that embedding AI-supported prompts within a structured, student-centered learning cycle was associated with higher self-reported academic self-concept and communication competence. Exploratory process data and qualitative narratives were consistent with a classroom routine in which peer explanation, verification of AI outputs, and structured self-evaluation supported students' learning experiences. These findings should be interpreted as convergent mixed-methods evidence for a pedagogically structured AI routine rather than as proof that AI exposure alone produced the observed gains.
DATA AVAILABILITY:
A comprehensive reproducibility package was prepared prior to journal submission, containing the fully anonymized raw dataset, a variable dictionary (Supplementary Table 11), scoring specifications (Supplementary Table 12), data cleaning specifications (Supplementary Table 13), statistical analysis scripts and specifications (Supplementary Tables 14 and 15), and source data for tables and figures (Supplementary Table 16). To ensure privacy compliance, all directly or indirectly identifiable student records, raw consent forms, and private chat histories were excluded. Structured tables were prepared as separate editable spreadsheet files, and figures were provided as distinct files, with their respective legends retained within the main manuscript. The underlying de-identified dataset and reproducibility materials, including the scored item matrix and qualitative code frequency matrix, are publicly available in the Zenodo repository (https://zenodo.org/records/20609465). The analysis dataset contains class identifiers, group coding, attendance rates, scale scores, and outcome change scores used for the reported analyses.

Figure 1: Classroom study flow and participant retention. This figure illustrates the recruitment process, class assignment, three-wave assessment schedule, exclusion criteria, and qualitative interview sampling. The final quantitative analytic sample comprised 174 students from six intact classes. Please click here to view a larger version of this figure.

Figure 2: Three-wave trajectories of academic self-concept and communication competence. (A) Longitudinal changes in academic self-concept across baseline, post-intervention, and follow-up assessments. (B) Longitudinal changes in communication competence across the corresponding assessment waves. Error bars represent the standard error of the mean; the plotted labels show mean values and SEM calculated as SD/sqrt(n), with n = 87 per group at each wave. Abbreviations: SEM = standard error of the mean; SD = standard deviation. Please click here to view a larger version of this figure.

Figure 3: Qualitative theme distribution and mixed-methods integration. (A) Qualitative theme frequencies by instructional group. (B) Structural alignment of qualitative themes with the quantitative process domains. The corrected theme counts are based on the qualitative code frequency matrix in Supplementary Table 10, including “checking AI-supported information” (AI-assisted = 14; comparison = 0). Please click here to view a larger version of this figure.
| Construct / record | Item code range | Items / records | Measurement wave | Response range | Score direction | Reverse-scored items | Missing-data rule | Final variable | Analytic use |
| Academic self-concept | ASC1–ASC6 | 6 items | T1, T2, T3 | 1–5 | Higher = stronger academic self-concept | ASC3, ASC5 | Compute scale score when ≥5 items are valid; otherwise set missing for that wave | ASC_T1, ASC_T2, ASC_T3 | Primary outcome; baseline-adjusted ANCOVA; T2–T1 and T3–T1 change-score analysis |
Communi-
cation competence | CC1–CC6 | 6 items | T1, T2, T3 | 1–5 | Higher = stronger communication competence | CC4 | Compute scale score when ≥5 items are valid; otherwise set missing for that wave | CC_T1, CC_T2, CC_T3 | Primary outcome; baseline-adjusted ANCOVA; T2–T1 and T3–T1 change-score analysis |
| Interaction quality | IQ1–IQ5 | 5 items | T2 | 1–5 | Higher = stronger perceived classroom interaction quality | None | Compute scale score when ≥4 items are valid; otherwise set missing | IQ_T2 | Process variable; exploratory regression; mixed-methods integration |
| Self-evaluation | SE1–SE5 | 5 items | T2 | 1–5 | Higher = stronger self-evaluation during learning | SE2 | Compute scale score when ≥4 items are valid; otherwise set missing | SE_T2 | Process variable; exploratory regression; mixed-methods integration |
| Learner engagement | LE1–LE6 | 6 items | T2 | 1–5 | Higher = stronger learner engagement | LE5 | Compute scale score when ≥5 items are valid; otherwise set missing | LE_T2 | Process variable; exploratory regression; mixed-methods integration |
| Attendance record | ATT1–ATT6 | 6 session records | Each session | Present / absent | Higher attendance = greater exposure | Not applicable | Exclude from final analytic sample if more than 25% of sessions are missed | attendance
_rate | Exposure check; exclusion screening |
| Fidelity checklist | FID1–FID6 | 6 session-level ratings | Each session | 0–2 per component | Higher = stronger protocol fidelity | Not applicable | Score each observed session; mark session-level deviation if fidelity <70% | fidelity_
percentage | Protocol deviation record; sensitivity interpretation |
| Technical incident log | TECH1–TECH6 | Session-level record | Each session | Categorical log | Not applicable | Not applicable | Record if incident affects more than one-third of students in a class | technical_
deviation | Technical deviation record; sensitivity interpretation |
| Qualitative interview | INT1–INT5 | Semi-structured domains | Post-intervention | Text / audio notes | Theme frequency by coded domain | Not applicable | Use only transcripts or structured notes with valid consent and anonymized ID | qualitative_
code_matrix | Mixed-methods explanation of interaction quality, self-evaluation, and learner engagement |
Table 1: Questionnaire structure, scoring rules, and analytic use. This table defines the locked scoring map established prior to data cleaning and analysis. T1 refers to the baseline assessment, T2 denotes the immediate post-intervention assessment, and T3 indicates the follow-up assessment. Scale scores are calculated as the arithmetic mean of valid items following reverse coding procedures. Entirely missing assessment waves are not imputed.
| Measure | Wave | Group | N | Mean | SD |
| Academic self-concept | T1 baseline | Usual student-centered | 87 | 3.05 | 0.43 |
| Academic self-concept | T1 baseline | AI-assisted student-centered instruction | 87 | 3 | 0.37 |
| Academic self-concept | T2 post-intervention | Usual student-centered | 87 | 3.19 | 0.6 |
| Academic self-concept | T2 post-intervention | AI-assisted student-centered instruction | 87 | 3.51 | 0.48 |
| Academic self-concept | T3 follow-up | Usual student-centered | 87 | 3.2 | 0.63 |
| Academic self-concept | T3 follow-up | AI-assisted student-centered instruction | 87 | 3.57 | 0.6 |
| Communication competence | T1 baseline | Usual student-centered | 87 | 3.04 | 0.38 |
| Communication competence | T1 baseline | AI-assisted student-centered instruction | 87 | 2.85 | 0.4 |
| Communication competence | T2 post-intervention | Usual student-centered | 87 | 3.12 | 0.62 |
| Communication competence | T2 post-intervention | AI-assisted student-centered instruction | 87 | 3.22 | 0.63 |
| Communication competence | T3 follow-up | Usual student-centered | 87 | 3.12 | 0.69 |
| Communication competence | T3 follow-up | AI-assisted student-centered instruction | 87 | 3.41 | 0.71 |
Table 2: Descriptive primary outcome scores by group and wave. This table presents descriptive metrics for academic self-concept and communication competence in both the AI-assisted student-centered instruction group and the usual student-centered instruction group across the three assessment waves. Abbreviations: SD = standard deviation.
| Construct | Items | Cronbach’s alpha | Decision |
| Academic self-concept, T1 | 6 | 0.676 | Borderline; retained according to the prespecified scoring map |
| Academic self-concept, T2 | 6 | 0.828 | Acceptable |
| Academic self-concept, T3 | 6 | 0.876 | Acceptable |
| Communication competence, T1 | 6 | 0.683 | Borderline; retained according to the prespecified scoring map |
| Communication competence, T2 | 6 | 0.877 | Acceptable |
| Communication competence, T3 | 6 | 0.897 | Acceptable |
| Interaction quality, T2 | 5 | 0.76 | Acceptable |
| Self-evaluation, T2 | 5 | 0.808 | Acceptable |
| Learner engagement, T2 | 6 | 0.806 | Acceptable |
Table 3: Internal consistency estimates for study measures. This table reports Cronbach’s alpha values for the primary outcome scales and process-related scales. Borderline baseline reliability estimates were retained to strictly adhere to the prespecified scoring map.
| AI-assisted change mean | AI-assisted change SD | Comparison change mean | Comparison change SD | Mean difference | 95% CI for difference | Welch t | p value | Cohen’s d |
| 0.51 | 0.44 | 0.14 | 0.46 | 0.37 | [0.24, 0.50] | 5.28 | <0.001 | 0.8 |
| 0.57 | 0.54 | 0.15 | 0.48 | 0.42 | [0.27, 0.57] | 5.46 | <0.001 | 0.83 |
| 0.37 | 0.45 | 0.08 | 0.46 | 0.29 | [0.15, 0.43] | 4.14 | <0.001 | 0.63 |
| 0.57 | 0.56 | 0.09 | 0.57 | 0.48 | [0.31, 0.65] | 5.61 | <0.001 | 0.85 |
Table 4: Between-group differences in outcome change scores. This table reports the T2–T1 and T3–T1 change scores for academic self-concept and communication competence. Between-group differences were evaluated using Welch’s t tests, with 95% confidence intervals and Cohen’s d reported for transparency. Abbreviations: CI = confidence interval; SD = standard deviation.
| Outcome | N | Group B | 95% CI for Group B | SE | t | p value | Baseline outcome B | Baseline p value | Model R² |
| Academic self-concept, T2 | 174 | 0.328 | [0.193, 0.463] | 0.1 | 5 | <0.001 | 0.748 | <0.001 | 0.395 |
| Academic self-concept, T3 | 174 | 0.401 | [0.246, 0.556] | 0.1 | 5 | <0.001 | 0.827 | <0.001 | 0.384 |
| Communication competence, T2 | 174 | 0.301 | [0.158, 0.444] | 0.1 | 4 | <0.001 | 1.116 | <0.001 | 0.488 |
| Communication competence, T3 | 174 | 0.516 | [0.340, 0.692] | 0.1 | 6 | <0.001 | 1.095 | <0.001 | 0.388 |
Table 5: Baseline-adjusted ANCOVA results for primary outcomes. This table presents the baseline-adjusted ANCOVA models for post-intervention and follow-up academic self-concept and communication competence, including group coefficients, 95% confidence intervals, standard errors, p values, and model R2 values. The instructional group variable was coded as 0 for usual student-centered instruction and 1 for AI-assisted student-centered instruction. Abbreviations: CI= confidence interval; SE= standard error.
| Outcome change | Predictor | B | SE | t | p value | Model R2 | Interpretation |
| Academic self-concept change | Instructional group | 0.299 | 0.078 | 3.85 | <0.001 | 0.191 | Group association retained |
| Academic self-concept change | Interaction quality | -0.014 | 0.073 | -0.20 | 0.844 | | Not independently significant |
| Academic self-concept change | Self-evaluation | 0.108 | 0.065 | 1.66 | 0.099 | | Positive, non-significant trend |
| Academic self-concept change | Learner engagement | 0.061 | 0.066 | 0.93 | 0.355 | | Not independently significant |
| Communication competence change | Instructional group | 0.252 | 0.082 | 3.09 | 0.002 | 0.141 | Group association retained |
| Communication competence change | Interaction quality | 0.131 | 0.075 | 1.75 | 0.082 | | Positive, non-significant trend |
| Communication competence change | Self-evaluation | 0.111 | 0.066 | 1.68 | 0.094 | | Positive, non-significant trend |
| Communication competence change | Learner engagement | -0.083 | 0.066 | -1.25 | 0.214 | | Not independently significant |
Table 6: Exploratory regression and class-level diagnostic results for outcome changes. This table reports exploratory linear associations among instructional group, interaction quality, self-evaluation, learner engagement, and outcome change scores, together with class-level ICC diagnostics. These models were not interpreted as confirmatory mediation tests. Abbreviations: SE = standard error.
Supplementary Table 1: Six-session classroom lesson guide. It summarizes the learning objectives, pedagogical phase allocations, and evidence retained for analysis.Please click here to download this file.
Supplementary Table 2: AI prompt sheet for the intervention group. It outlines permitted functions, prohibited uses, and required student records.Please click here to download this file.
Supplementary Table 3: Teacher prompt sheet for the comparison group. It outlines the standardized teacher-guided prompts used to match the instructional structure of the AI-assisted group.Please click here to download this file.
Supplementary Table 4: Student worksheet template. It provides the required fields, formats, and data collection status.Please click here to download this file.
Supplementary Table 5: Self-evaluation form. It summarizes the reflection items, response formats, and respective coding rules.Please click here to download this file.
Supplementary Table 6: Classroom observation and fidelity checklist. It provides the component-specific scoring rubrics from 0 to 2.Please click here to download this file.
Supplementary Table 7: Technical incident log. It documents incident types, affected student counts, and protocol deviation codes.Please click here to download this file.
Supplementary Table 8: Semi-structured interview guide. It details the qualitative domains, core questions, optional probes, and linked constructs.Please click here to download this file.
Supplementary Table 9: Qualitative codebook. It defines the domains, inclusion and exclusion rules, and specific evidence sources.Please click here to download this file.
Supplementary Table 10: Qualitative code frequency matrix. It reports the coded case counts across both instructional groups.Please click here to download this file.
Supplementary Table 11: Data dictionary. It defines variable names, labels, permitted values, missing value rules, and analysis uses.Please click here to download this file.
Supplementary Table 12: Scoring specifications. It details reverse-coded items, minimum valid item requirements, scale formulas, and output variables for each construct.Please click here to download this file.
Supplementary Table 13: Data cleaning specifications. It outlines the cleaning steps, rules, and generated output logs for dataset verification.Please click here to download this file.
Supplementary Table 14: Statistical analysis specifications. It summarizes the analytical procedures, models, and decision thresholds for primary and secondary analyses.Please click here to download this file.
Supplementary Table 15: Cluster-sensitive analysis specifications. It describes the procedures and interpretation rules for class-level aggregation and protocol deviation reviews.Please click here to download this file.
Supplementary Table 16: Source data index. It links the final output tables and figures to their required statistics, source files, and manuscript locations.Please click here to download this file.