Research Article

AI-Assisted Student-Centered Instruction and Adolescents' Academic Self-Concept and Communication Competence: A Mixed-Methods Study

43 views

DOI:

10.3791/72078

August 4th, 2026

In This Article

Summary

This mixed-methods study evaluates an AI-assisted, student-centered classroom routine involving AI-supported inquiry, peer discussion, teacher clarification, and self-evaluation. The study focuses on self-reported academic self-concept and communication competence and provides sufficient procedural detail to facilitate classroom replication.

Abstract

This classroom-based mixed-methods quasi-experimental study examined whether AI-assisted student-centered instruction was associated with changes in adolescents’ academic self-concept and self-reported communication competence. Six intact classes were assigned to either the AI-assisted instruction or the usual student-centered instruction condition. The intervention embedded AI-supported prompts into a structured learning cycle comprising individual inquiry, peer discussion, teacher clarification, and self-evaluation. Of the 186 recruited students, 174 were retained for quantitative analysis. Academic self-concept and self-reported communication competence were assessed at baseline, immediately after the intervention, and at a predefined short-term follow-up. Interaction quality, self-evaluation, and learner engagement were measured after the intervention, and 36 students participated in semi-structured interviews. Compared with the usual instruction group, the AI-assisted group showed larger baseline-adjusted gains in both outcomes at post-intervention and follow-up. Exploratory process analyses indicated higher interaction quality, self-evaluation, and learner engagement in the AI-assisted group, while qualitative findings suggested that AI-supported prompts helped students to verify explanations, revise responses, and make self-evaluation more explicit when combined with peer discussion and teacher guidance. Because participants were assigned at the class level and outcomes were primarily self-reported, the findings should be interpreted as classroom-based associations rather than definitive evidence of a causal improvement in academic performance.

Introduction

Artificial intelligence (AI) has become increasingly prevalent in educational settings. However, its instructional value remains inconsistent, as mere access to a technological tool does not inherently constitute a learning process. While reviews highlight advancements in educational AI, they often overlook the teacher’s role and classroom interactions. This is critical for adolescents, who actively form competence judgments and test ideas with peers. Therefore, AI-supported learning must be implemented as a deliberate pedagogical arrangement that shapes how students question, evaluate, and discuss, rather than serving as mere technological exposure.

Student-centered instruction provides a robust framework for addressing this challenge. By shifting students away from passive reception, this approach prioritizes active participation, task engagement, peer-to-peer explanation, and autonomous decision-making. Meta-analytic research confirms that active learning environments characterized by problem-solving and student involvement can significantly enhance academic performance compared with traditional lectures1. Nevertheless, student-centered instruction is not intrinsically effective. It requires instructional tasks that are sufficiently structured to guide participation, yet flexible enough to encourage students to explain, compare, and revise ideas. This dynamic becomes even more complex in AI-assisted classrooms. Although AI can rapidly generate examples, explanations, and suggestions, the absence of teacher guidance and peer evaluation may lead students to accept these outputs uncritically, treating them as cognitive shortcuts rather than as raw material for deeper thinking.

Academic self-concept serves as a highly relevant outcome in this context, reflecting students’ domain-specific beliefs about their academic capabilities. Prior research has established that academic self-concept is not a static personality trait, but rather a dynamic construct closely intertwined with achievement and ongoing learning experiences2. This study assesses academic self-concept to evaluate students' confidence in their ability to improve their work. Communication competence serves as a secondary outcome, as the intervention requires students to articulate justifications and refine explanations through structured peer feedback.

Feedback and self-evaluation constitute the crucial instructional bridge connecting AI support to these targeted learning outcomes. Effective formative assessment requires students to actively monitor progress, measure it against standards, and act upon performance discrepancies3,4. In an AI-assisted, student-centered classroom, AI prompts may generate alternative explanations or cues for revision. However, the true educational value of these prompts hinges on whether students scrutinize the output, debate it with peers, and translate it into a refined response. Consequently, self-evaluation in this study is operationalized not as a vague attitude, but as a concrete, observable classroom behavior aimed at identifying learning gaps and making revisions.

Learner engagement plays an equally vital role in this learning cycle. Recognized as a multidimensional construct, engagement spans behavioral participation, emotional involvement, and cognitive investment5. This multidimensionality is especially pertinent for adolescents, as superficial compliance does not guarantee deep cognitive processing. One student might complete a worksheet mechanically without analyzing the underlying reasoning, while another might ask fewer questions but engage profoundly through continuous self-monitoring and revision. The introduction of AI can theoretically elevate engagement by providing rich material for comparison and inquiry; conversely, it can erode engagement if students default to effortless reliance on automated outputs. Accordingly, this study contextualizes learner engagement alongside interaction quality and self-evaluation, deliberately avoiding the treatment of AI usage as an isolated classroom event.

Debates surrounding educational large language models highlight a tension between individualized assistance and risks such as overreliance or diminished effort when deployed without rigorous boundaries. Consequently, the pedagogical choreography of the learning routine is just as consequential as the AI technology itself. The present study embedded AI-supported prompts within a consistent, five-step classroom sequence: teacher briefing, individual inquiry, peer discussion, whole-class clarification, and self-evaluation. This specific design intentionally positions AI output as preliminary material subject to rigorous verification and revision rather than as an infallible final answer.

A mixed-methods research design was adopted because quantitative survey data alone cannot adequately capture the nuances of students’ lived learning experiences. Mixed-methods approaches are particularly valuable when outcome trajectories must be interpreted against the backdrop of participant experiences, implementation conditions, and unfolding classroom dynamics6. In the current investigation, the quantitative component tracked developmental changes in academic self-concept and communication competence. Concurrently, the qualitative component explored how students characterized interaction quality, self-evaluation practices, task engagement, and perceived barriers during the intervention. These two streams of evidence were subsequently integrated to evaluate whether the observed quantitative score patterns aligned with the students’ ground-level classroom experiences.

This study also addresses a significant practical gap in the literature on AI-assisted instruction. Although numerous studies evaluate AI-driven outcomes, relatively few offer a transparent, reproducible classroom protocol detailing prompt utilization, peer scaffolding, control group comparisons, fidelity monitoring, and mixed-methods data integration. Decades of research on peer interaction demonstrate that students benefit substantially when they explain concepts, receive peer feedback, and continuously refine their understanding through dialogue7. The present study builds on this foundational logic, investigating whether AI-generated prompts can be successfully interwoven into this dialogic structure without usurping the indispensable roles of teacher guidance and peer collaboration.

Accordingly, this study examined AI-assisted, student-centered instruction as a structured classroom protocol rather than as simple exposure to an AI tool. The proposed instructional pathway was that AI-supported prompts would first provide material for inquiry and comparison; peer discussion and teacher clarification would then require students to verify, justify, and revise that material; and the final self-evaluation step would help students recognize progress and unresolved gaps. Within this pathway, interaction quality, self-evaluation, and learner engagement were treated as exploratory process variables that may help explain how students experienced the instructional routine, rather than as confirmed mediators.

The primary research questions were as follows: Does the AI-assisted student-centered protocol produce larger changes in academic self-concept and self-reported communication competence than usual student-centered instruction? Are post-intervention interaction quality, self-evaluation, and learner engagement higher in the AI-assisted condition? The secondary exploratory questions were as follows: Do these process variables show associations with outcome change scores? Do qualitative interview themes converge with, expand upon, or complicate the quantitative findings? Because intact classes, rather than individual students, were assigned to conditions, all findings were interpreted with explicit caution regarding causal inference and classroom-level clustering.

Protocol

All methods involving human participants were conducted in accordance with institutional guidelines and the declaration of Helsinki. The study protocol was reviewed and approved by the institutional review board  of Guangzhou city construction college (IRB approval no. GZCJ2026A016; approval date: January 2, 2026). The approved study title was “ai-assisted student-centered instruction and adolescents’ academic self-concept and communication competence: a mixed-methods study.” Written permission was obtained from the participating school before the study was implemented. Written informed consent was obtained from parents or legal guardians, and written assent was obtained from all participating students. Students were informed that participation was voluntary and that they could withdraw from the study at any time without academic penalty. No names, contact details, school identifiers, grades, or other directly identifiable information were entered into the ai chatbot platform or included in the public dataset. All study procedures were conducted during the approved study period from January 2, 2026, to January 2, 2027. All tools and platforms used in this study are listed in the Table of Materials

Study design and classroom assignment

This study employed a quasi-experimental, mixed-methods classroom design with two instructional conditions: an AI-assisted, student-centered instruction group and a usual student-centered instruction group. Longitudinal measurements were conducted at three time points: baseline before the intervention (T1), immediately after the final instructional session (T2), and at a predefined short-term follow-up assessment (T3). The T2–T3 interval was recorded in the study log and kept consistent across the manuscript, dataset, analysis scripts, and figure labels. Figure 1 summarizes recruitment, class assignment, assessment timing, exclusions, and qualitative sampling. Because administrative and ecological constraints prevented individual randomization, six intact classes were assigned at the classroom level, with three classes in the AI-assisted condition and three classes in the comparison condition. Each class contributed 29 students to the final analytic sample. Grade level, subject area, lesson sequence, session length, classroom tasks, assessment schedule, and teacher contact time were kept comparable across conditions. Class identifiers were retained for cluster-sensitive statistical checks.

Participant recruitment and eligibility screening

Participants were recruited directly from the participating intact classes. Students and their parents or legal guardians received a detailed information sheet outlining the study purpose, classroom procedures, AI-supported activities, questionnaire schedules, the optional interview component, privacy protections, withdrawal rights, and data storage plans. Students were included in the study if they were enrolled in one of the participating classes, attended regular classroom instruction, provided written assent, obtained written parental or guardian consent, and completed the baseline questionnaire before the first session. Conversely, students were excluded from the final analytic sample if they transferred from the school, missed more than 25% of the intervention sessions, submitted invalid questionnaire records, or withdrew permission for data use. One primary reason for exclusion was documented for each excluded participant. Prior to data entry, each eligible participant was assigned a unique study identification code. Relevant demographic and educational covariates were recorded, including age, gender, class identifier, baseline academic performance band, prior exposure to AI-assisted learning, and baseline digital learning familiarity. For statistical tracking, the instructional group variable was coded as 0 for the comparison condition and 1 for the AI-assisted condition.

Instructional materials and intervention delivery

A six-session instructional unit (Supplementary Table 1) was developed prior to implementation, with each classroom session lasting 40–45 min. Identical learning objectives, lesson topics, worksheets, peer discussion tasks, self-evaluation forms, and reflection prompts were applied across both instructional conditions. The questionnaire structure, scoring directions, reverse-scored items, and final scale-level variables followed the locked scoring map in Table 1. Reproducibility materials included lesson plans (Supplementary Table 1), AI prompt sheets (Supplementary Table 2), teacher prompt sheets (Supplementary Table 3), student worksheets (Supplementary Table 4), self-evaluation forms (Supplementary Table 5), observation checklists (Supplementary Table 6), technical incident logs (Supplementary Table 7), interview guides (Supplementary Table 8), and a qualitative codebook (Supplementary Table 9). The participating teacher completed training using a written delivery guide covering session sequence, AI-use boundaries, comparison-group procedures, privacy requirements, and fidelity checklists. A rehearsal with non-study students was conducted to verify activity timing, device access, worksheet clarity, and task difficulty. The AI tool was an institutionally approved ChatGPT-based chatbot using the GPT-4 architecture. For classroom use, chat history, profile memory, web browsing, plug-ins, image generation, and automated grading were disabled. The manuscript uses the generic term "AI chatbot platform" after its first mention to avoid promotional language. Prompt-design principles were fixed before the intervention and limited AI use to clarification, example generation, explanation comparison, targeted revision support, and self-evaluation cues. Representative prompts, sample AI outputs, and teacher moderation examples were retained as a part of the reproducibility materials.

Each AI-assisted session was conducted in five structured phases. First, a 5 min teacher briefing was conducted to state the learning goals, explain tasks, identify the configured AI chatbot platform, and remind students that AI outputs were provisional prompts rather than authoritative answers. Students were strictly instructed not to enter any identifiable personal information into the platform. Second, a 10 min individual AI-supported inquiry activity followed, during which students utilized the prepared prompt sheet to request clarifications, examples, alternative explanations, or revision suggestions. Each student recorded the prompt category used, one useful AI-supported idea, one point requiring verification, and their final adoption decision. Third, a 12–15 min peer discussion activity was conducted, during which students compared their AI-supported ideas in small groups, identified inaccuracies, and collaboratively revised their work. During this phase, the teacher circulated throughout the classroom to apply a standardized moderation rule, prompting students to verbally justify their revisions against lesson criteria and correcting erroneous or overconfident AI outputs before finalization. Fourth, a 5–8 min whole-class clarification session was conducted to address widespread misconceptions and correct factual errors. Finally, a 5 min self-evaluation activity concluded the session, during which students rated their understanding, recorded one concrete improvement, and noted one unresolved question. This exact operational sequence was maintained across all intervention classes. Printed prompt sheets served as technical backups, and technical deviations affecting over 1/3 of a class were explicitly logged. Anonymized representative prompt-output-moderation examples were retained and documented in Supplementary Tables 2, 3.

The comparison condition was delivered utilizing the identical session length, learning objectives, topic sequence, worksheets, peer discussion intervals, self-evaluation forms, and teacher feedback windows. However, the AI-supported prompts were replaced with standardized, teacher-prepared guiding prompts that instructed students to explain core concepts, compare examples, identify weaknesses, and revise their work. The identical five-step structure encompassing the briefing, individual inquiry, peer discussion, whole-class clarification, and final self-evaluation was strictly maintained. No additional tutoring, out-of-class assignments, prolonged feedback periods, or different assessment tasks were provided to either condition. Classroom attendance, behavioral disruptions, teacher substitutions, and incomplete components were documented using the same logging format for both groups.

Data collection, fidelity monitoring, and quantitative analysis

Assessment questionnaires were administered at baseline, post-intervention, and follow-up intervals, utilizing consistent item wording, response scales, and scoring rules across all three measurement waves. The instruments comprised validated scales evaluating academic self-concept, communication competence, interaction quality, self-evaluation, and learner engagement. Responses were gathered using a standard 5-point Likert-type scale ranging from 1 (strong disagreement) to 5 (strong agreement), with higher scores indicating higher levels of the target construct. Individual scale scores were calculated as the arithmetic mean of valid items following necessary reverse coding. Missing item values were replaced with the student's mean score within the same scale and wave, provided that no more than 20% of the items were missing. Scales with more than 20% missing items were treated as missing, and data from an entirely missing assessment wave were not imputed. Prior to statistical analysis, the raw dataset was carefully audited for range violations, duplicate entries, straight-line response patterns, and anomalous completion times. Records were flagged as invalid and excluded if they exhibited impossible numerical values, invariant responses across all items, or completion times below 1/3 of the median duration, unless explicit classroom logs validated their authenticity. Main post-intervention and follow-up analyses included students who provided valid baseline scores and corresponding post-intervention or follow-up outcome data, thereby avoiding unilateral listwise deletion for partially complete schedules. A locked data dictionary was established prior to data entry to define variable names, labels, response ranges, missing value codes, reverse-scoring rules, wave identifiers, and final scale-level constructs. Questionnaire data were entered in a wide format, with a single row per student and wave suffixes (_T1, _T2, _T3) used to distinguish the measurement waves. Numerical ranges and unique participant records were computationally verified, and missingness patterns were inspected across groups, classes, and waves against the master attendance log.

A standardized classroom observation checklist was employed during every instructional session to monitor implementation fidelity. Each pedagogical component was scored as 0 (not completed), 1 (partially completed), or 2 (fully completed). The evaluated components strictly aligned with the five-phase instructional sequence, including the teacher briefing, individual inquiry, peer discussion, whole-class clarification, and the completion of self-evaluation records. The overall session fidelity percentage was calculated by dividing the observed score by the maximum possible score. Any session falling below a 70% fidelity threshold was recorded as a protocol deviation. To ensure inter-rater reliability, a second trained observer independently rated at least 20% of the total sessions. Percentage agreement was calculated across all items, and any rating discrepancies were resolved through consensus discussion whenever the initial agreement fell below 80%. Hardware or software technical malfunctions were recorded in separate logs to distinguish them from instructional fidelity issues.

Statistical analyses were performed using a reproducible, script-based workflow in which the random seed was fixed prior to any sampling or sensitivity checks. All statistical tests were two-tailed, with significance set at p < 0.05. Continuous variables were summarized as means and standard deviations, while categorical variables were expressed as frequencies and percentages. Baseline equivalence between the cohorts was evaluated descriptively using independent-samples t-tests for continuous data and chi-square tests for categorical distributions. Internal consistency was measured using Cronbach’s alpha, applying thresholds of α ≥ 0.70 for general scales and α ≥ 0.80 for primary outcome measures, alongside corrected item-total correlations. Baseline-adjusted ANCOVA served as the primary analytical approach. Separate models were fitted for each primary outcome, with the post-intervention or follow-up score specified as the dependent variable, the instructional group as the main predictor, and the corresponding baseline score as a mandatory covariate. Additional covariates (gender, baseline academic performance band, and prior AI exposure) were included to control for baseline imbalances or their theoretical relevance, and adjusted mean differences, 95% confidence intervals, standard errors, and model R2 values were reported. Change-score comparisons (T2–T1 and T3–T1) were performed as secondary analyses using independent-samples Welch's t-tests, with effect sizes quantified using Cohen’s d. Exploratory process analyses evaluated Pearson correlations among process variables and outcome changes, supplemented by multi-predictor linear regressions with variance inflation factor (VIF) monitoring to check for multicollinearity. Finally, cluster-sensitive checks were implemented to account for the nested structure of students within the six intact classes. Unconditional intraclass correlation coefficients (ICCs) were estimated, outcome changes were aggregated and compared at the class level, and primary ANCOVA calculations were repeated using cluster-robust standard errors as an exploratory sensitivity check.

Qualitative sampling, coding, and mixed-methods integration

Immediately following the post-intervention questionnaire, a purposive subsample of 36 students was selected for semi-structured interviews. This qualitative cohort included participants from both instructional conditions who demonstrated high, medium, or low change scores in academic self-concept or communication competence. The semi-structured interview guide investigated experiences of classroom interaction, confidence in explaining ideas, self-evaluation behaviors, task engagement, and willingness to communicate, using targeted prompts to explore moments of revision, responses to unclear feedback, and variations in engagement. Interviews were audio-recorded only after obtaining parental consent and student assent; otherwise, structured written notes were taken. All personal names, class numbers, and identifying details were removed during transcription, and transcripts were linked to their corresponding quantitative identifiers. The initial qualitative coding framework was deductively derived from interaction quality, self-evaluation, and learner engagement constructs, with inductive thematic codes added only when the data failed to fit the baseline framework. Two independent coders analyzed at least 20% of the transcripts to verify reliability. Cohen's kappa coefficients were evaluated, with values between 0.61–0.80 treated as substantial agreement and values above 0.80 as strong agreement. Codebook definitions were refined and double-coding was repeated whenever the initial kappa fell below 0.61, and the final definitions alongside the resulting qualitative code frequency matrix were archived (Supplementary Tables 9 and 10).

Quantitative data analysis and qualitative coding frameworks were finalized independently before initiating mixed-methods integration. A cohesive joint display was structurally constructed to link the three quantitative process variables with their corresponding qualitative themes. Integrated findings were systematically classified as convergence (where questionnaire patterns and interview themes aligned), expansion (where qualitative narratives explained the mechanisms underlying quantitative trends), or divergence (where interviews revealed unique barriers or uneven participation that complicated the quantitative patterns). Qualitative findings were utilized to contextualize classroom processes, linking learner engagement scores with themes of active questioning and peer discussion, and linking self-evaluation scores with accounts of gap awareness and strategic adjustments.

Results

Participant flow and baseline characteristics

An initial cohort of 186 students from six intact classes was recruited for the study. Following eligibility screening, attendance verification, and response-quality checks, 174 students were retained for the final quantitative analysis. These participants were evenly distributed, with 87 assigned to the AI-assisted student-centered instruction group and 87 assigned to the usual student-centered instruction group. Figure 1 details the recruitment flow, class assignment, three-wave assessment timeline, exclusion criteria, and qualitative subsampling.

Table 2 summarizes the descriptive primary outcome scores by instructional group and measurement wave. Before the intervention, academic self-concept levels were similar between the two cohorts (AI-assisted: 3.00 ± 0.37; comparison: 3.05 ± 0.43). However, the AI-assisted group had lower baseline communication competence (2.85 ± 0.40) than the usual instruction group (3.04 ± 0.38). Because assignment occurred at the intact-class level rather than through individual randomization, these baseline differences were treated as potential pre-existing class differences rather than intervention effects. Baseline outcome scores were retained as covariates in all primary models, and communication competence findings were interpreted with attention to the risk of regression to the mean. To address the hierarchical data structure, unconditional intraclass correlation coefficients (ICC[1]s) were estimated from one-way random-effects ANOVA models using class identifiers as the grouping variable. The formula used was:

figure-results-1   (1)

where MSbetween represents the mean square between classes, MSwithin represents the mean square within classes, and m denotes the class size. The final dataset included six classes of equal size (k = 6; m = 29). ICCs for primary observed scores ranged from 0.007 to 0.105, and ICCs for primary change scores were ASC T2–T1 = 0.152, ASC T3–T1 = 0.160, CC T2–T1 = 0.105, and CC T3–T1 = 0.165. These values indicated non-negligible classroom-level dependence. Because only six level-2 clusters were available, fitting a full multilevel model was considered unstable; therefore, ICCs, class-level summaries, and cluster-robust sensitivity checks were used as diagnostic evidence rather than as confirmatory multilevel estimates.

Implementation monitoring records were reviewed alongside the analytic dataset. All six classes completed the six-session unit, and attendance exposure was comparable across conditions (mean attendance rate: AI-assisted = 0.934; usual instruction = 0.936). No retained participant fell below the prespecified exposure threshold. To monitor implementation fidelity, session-level checklists and technical incident logs were utilized, with a second independent observer evaluating at least 20% of the sessions. These observations served as a formative, on-site quality assurance mechanism for providing immediate feedback to the instructor, rather than a summative measurement. Because the primary goal was real-time pedagogical correction, aggregate statistical metrics (such as mean fidelity scores or inter-rater reliability coefficients) were not permanently archived for retrospective analysis. Nevertheless, these formative checks confirmed that no major protocol deviations occurred, and the core instructional sequence was consistently maintained.

Internal consistency of the study measures

Table 3 reports the internal consistency estimates for all measured constructs using the item counts defined in the locked scoring map. Academic self-concept demonstrated acceptable reliability during both the post-intervention and follow-up phases, yielding Cronbach’s alpha values of 0.828 and 0.876, respectively. Communication competence similarly maintained acceptable reliability across these later assessment waves, with alpha values of 0.877 and 0.897.

Baseline reliability estimates for academic self-concept (α = 0.676) and communication competence (α = 0.683) were below the preferred 0.70 threshold. However, these estimates were retained to preserve the prespecified scoring structure, but they indicate greater baseline measurement error and may attenuate or destabilize baseline-adjusted estimates. The post-intervention process-related scales exhibited acceptable internal consistency, with alpha values of 0.760 for interaction quality, 0.808 for self-evaluation, and 0.806 for learner engagement.

Academic self-concept outcomes

Both groups experienced an increase in academic self-concept over time, although the AI-assisted group demonstrated a distinctly larger gain. From baseline to the immediate post-intervention assessment, the AI-assisted group's mean score rose from 3.00 ± 0.37 to 3.51 ± 0.48. In contrast, the comparison group showed a modest increase from 3.05 ± 0.43 to 3.19 ± 0.60. This resulted in a mean T2–T1 change of 0.51 ± 0.44 for the AI-assisted cohort and 0.14 ± 0.46 for the comparison cohort. The resulting between-group difference in the change score was 0.37 (95% CI [0.24, 0.50]), representing a statistically significant effect with a large magnitude (Welch t = 5.28, p < 0.001, Cohen’s d = 0.80).

This advantage was sustained at the follow-up assessment. The academic self-concept score reached 3.57 ± 0.60 in the AI-assisted group, compared with 3.20 ± 0.63 in the usual instruction group. The mean T3–T1 change was 0.57 ± 0.54 for the AI-assisted group and 0.15 ± 0.48 for the usual instruction group. The corresponding between-group difference in change increased slightly to 0.42 (95% CI [0.27, 0.57]), maintaining strong statistical significance (Welch t = 5.46, p < 0.001, Cohen’s d = 0.83). These three-wave longitudinal trajectories are shown in Figure 2A, while the complete change-score comparisons are presented in Table 4.

Table 5 presents the baseline-adjusted ANCOVA results. After controlling for baseline scores and relevant covariates, instructional group assignment was associated with higher post-intervention academic self-concept (B = 0.328, 95% CI [0.193, 0.463], SE = 0.069, t = 4.77, p < 0.001, model R2 = 0.395). This association was also observed at follow-up (B = 0.401, 95% CI [0.246, 0.556], SE = 0.079, t = 5.07, p < 0.001, model R2 = 0.384). Baseline academic self-concept was a significant covariate in both analytical models.

Communication competence outcomes

Students in the AI-assisted condition also exhibited a more pronounced improvement in communication competence. Between the baseline and post-intervention measurements, the AI-assisted group improved from 2.85 ± 0.40 to 3.22 ± 0.63, whereas the comparison group showed a modest increase from 3.04 ± 0.38 to 3.12 ± 0.62. The mean T2–T1 change was 0.37 ± 0.45 for the AI-assisted participants and 0.08 ± 0.46 for the comparison participants. This yielded a statistically significant between-group difference in change of 0.29 (95% CI [0.15, 0.43]), corresponding to a medium-to-large effect size (Welch t = 4.14, p < 0.001, Cohen’s d = 0.63).

During the follow-up phase, communication competence scores increased to 3.41 ± 0.71 in the AI-assisted group and 3.12 ± 0.69 in the comparison group. The mean T3–T1 change was recorded as 0.57 ± 0.56 in the intervention condition versus 0.09 ± 0.57 in the comparison condition. The difference in these change scores was 0.48 (95% CI [0.31, 0.65]), representing a statistically significant outcome with large effect size (Welch t = 5.61, p < 0.001, Cohen’s d = 0.85). Figure 2B illustrates these trajectories, and Table 4 provides the comprehensive change-score data.

Consistent with the raw change scores, the baseline-adjusted ANCOVA showed that group assignment was associated with post-intervention communication competence (B = 0.301, 95% CI [0.158, 0.444], SE = 0.073, t = 4.14, p < 0.001, model R2 = 0.488). This association persisted at follow-up (B = 0.516, 95% CI [0.340, 0.692], SE = 0.090, t = 5.72, p < 0.001, model R2 = 0.388). Baseline communication competence remained as a significant predictor in both adjusted models, as detailed in Table 5.

Exploratory associations with interaction quality, self-evaluation, and learner engagement

Post-intervention measures revealed that interaction quality, self-evaluation, and learner engagement were higher in the AI-assisted group. Linear modeling indicated that group assignment was positively associated with interaction quality (B = 0.344, SE = 0.075, t = 4.62, p < 0.001, model R2 = 0.133), self-evaluation (B = 0.291, SE = 0.082, t = 3.57, p < 0.001, model R2 = 0.118), and learner engagement (B = 0.361, SE = 0.081, t = 4.43, p < 0.001, model R2 = 0.129).

Table 6 summarizes exploratory regression models that examined process-related associations with outcome changes. When interaction quality, self-evaluation, and learner engagement were entered simultaneously alongside group assignment, the instructional group remained associated with changes in academic self-concept (B = 0.299, SE = 0.078, t = 3.85, p < 0.001, model R2 = 0.191). However, interaction quality showed no independent association with academic self-concept change in this model (B = 0.014, SE = 0.073, t = -0.20, p = 0.844). Both self-evaluation (B = 0.108, SE = 0.065, t = 1.66, p = 0.099) and learner engagement (B = 0.061, SE = 0.066, t = 0.93, p = 0.355) exhibited positive but statistically nonsignificant trends. These models were not interpreted as formal mediation tests.

A parallel pattern emerged for the changes in communication competence. Group assignment remained significant in the multipredictor model (B = 0.252, SE = 0.082, t = 3.09, p = 0.002, model R2 = 0.141). Interaction quality (B = 0.131, SE = 0.075, t = 1.75, p = 0.082) and self-evaluation (B = 0.111, SE = 0.066, t = 1.68, p = 0.094) displayed positive but nonsignificant associations. Learner engagement was similarly nonsignificant (B = -0.083, SE = 0.066, t = -1.25, p = 0.214).

Qualitative coding and mixed-methods integration

To complement the quantitative findings, a purposive subsample of 36 students underwent qualitative analysis. Coded themes were structured around four primary domains: interaction quality, self-evaluation, learner engagement, and implementation barriers. Figure 3A shows the distribution of these themes by instructional condition.

Within the AI-assisted cohort, the most frequent coded themes were AI-supported peer explanation (15 coded cases), active questioning (14), checking AI-supported information (14), gap awareness (13), revision reasoning (12), and task persistence (11). Implementation barriers were less frequent but still observed, including shallow use (5) and technical barriers (4). In the comparison group, AI-supported peer explanation was absent by design, whereas peer explanation related to the matched prompt routine appeared in 10 coded cases; active questioning appeared in 9 cases, gap awareness and task persistence in 8 cases each, revision reasoning in 7 cases, and shallow use in 2 cases. Technical barriers and checking AI-supported information were coded as 0 in the comparison condition. These counts describe the number of coded cases in the interview subsample rather than population.

The integration of quantitative and qualitative data (Figure 3B) revealed convergence in two areas. First, the higher interaction quality scores recorded in the AI-assisted group corresponded with the frequent qualitative codes for AI-supported peer explanation and checking AI-supported information. Second, higher self-evaluation scores in the intervention group aligned with qualitative narratives emphasizing gap awareness and revision reasoning. The qualitative themes also expanded the engagement findings by showing that active participation depended on perceived task authenticity, teacher moderation, and avoidance of shallow AI use.

Summary of findings

In summary, the integrated results suggest that embedding AI-supported prompts within a structured, student-centered learning cycle was associated with higher self-reported academic self-concept and communication competence. Exploratory process data and qualitative narratives were consistent with a classroom routine in which peer explanation, verification of AI outputs, and structured self-evaluation supported students' learning experiences. These findings should be interpreted as convergent mixed-methods evidence for a pedagogically structured AI routine rather than as proof that AI exposure alone produced the observed gains.

DATA AVAILABILITY:

A comprehensive reproducibility package was prepared prior to journal submission, containing the fully anonymized raw dataset, a variable dictionary (Supplementary Table 11), scoring specifications (Supplementary Table 12), data cleaning specifications (Supplementary Table 13), statistical analysis scripts and specifications (Supplementary Tables 14 and 15), and source data for tables and figures (Supplementary Table 16). To ensure privacy compliance, all directly or indirectly identifiable student records, raw consent forms, and private chat histories were excluded. Structured tables were prepared as separate editable spreadsheet files, and figures were provided as distinct files, with their respective legends retained within the main manuscript. The underlying de-identified dataset and reproducibility materials, including the scored item matrix and qualitative code frequency matrix, are publicly available in the Zenodo repository (https://zenodo.org/records/20609465). The analysis dataset contains class identifiers, group coding, attendance rates, scale scores, and outcome change scores used for the reported analyses.

figure-results-2
Figure 1: Classroom study flow and participant retention. This figure illustrates the recruitment process, class assignment, three-wave assessment schedule, exclusion criteria, and qualitative interview sampling. The final quantitative analytic sample comprised 174 students from six intact classes. Please click here to view a larger version of this figure.

figure-results-3
Figure 2: Three-wave trajectories of academic self-concept and communication competence. (A) Longitudinal changes in academic self-concept across baseline, post-intervention, and follow-up assessments. (B) Longitudinal changes in communication competence across the corresponding assessment waves. Error bars represent the standard error of the mean; the plotted labels show mean values and SEM calculated as SD/sqrt(n), with n = 87 per group at each wave. Abbreviations: SEM = standard error of the mean; SD = standard deviation. Please click here to view a larger version of this figure.

figure-results-4
Figure 3: Qualitative theme distribution and mixed-methods integration. (A) Qualitative theme frequencies by instructional group. (B) Structural alignment of qualitative themes with the quantitative process domains. The corrected theme counts are based on the qualitative code frequency matrix in Supplementary Table 10, including “checking AI-supported information” (AI-assisted = 14; comparison = 0). Please click here to view a larger version of this figure.

Construct / recordItem code rangeItems / recordsMeasurement waveResponse rangeScore directionReverse-scored itemsMissing-data ruleFinal variableAnalytic use
Academic self-conceptASC1–ASC66 itemsT1, T2, T31–5Higher = stronger academic self-conceptASC3, ASC5Compute scale score when ≥5 items are valid; otherwise set missing for that waveASC_T1, ASC_T2, ASC_T3Primary outcome; baseline-adjusted ANCOVA; T2–T1 and T3–T1 change-score analysis
Communi-
cation competence
CC1–CC66 itemsT1, T2, T31–5Higher = stronger communication competenceCC4Compute scale score when ≥5 items are valid; otherwise set missing for that waveCC_T1, CC_T2, CC_T3Primary outcome; baseline-adjusted ANCOVA; T2–T1 and T3–T1 change-score analysis
Interaction qualityIQ1–IQ55 itemsT21–5Higher = stronger perceived classroom interaction qualityNoneCompute scale score when ≥4 items are valid; otherwise set missingIQ_T2Process variable; exploratory regression; mixed-methods integration
Self-evaluationSE1–SE55 itemsT21–5Higher = stronger self-evaluation during learningSE2Compute scale score when ≥4 items are valid; otherwise set missingSE_T2Process variable; exploratory regression; mixed-methods integration
Learner engagementLE1–LE66 itemsT21–5Higher = stronger learner engagementLE5Compute scale score when ≥5 items are valid; otherwise set missingLE_T2Process variable; exploratory regression; mixed-methods integration
Attendance recordATT1–ATT66 session recordsEach sessionPresent / absentHigher attendance = greater exposureNot applicableExclude from final analytic sample if more than 25% of sessions are missedattendance
_rate
Exposure check; exclusion screening
Fidelity checklistFID1–FID66 session-level ratingsEach session0–2 per componentHigher = stronger protocol fidelityNot applicableScore each observed session; mark session-level deviation if fidelity <70%fidelity_
percentage
Protocol deviation record; sensitivity interpretation
Technical incident logTECH1–TECH6Session-level recordEach sessionCategorical logNot applicableNot applicableRecord if incident affects more than one-third of students in a classtechnical_
deviation
Technical deviation record; sensitivity interpretation
Qualitative interviewINT1–INT5Semi-structured domainsPost-interventionText / audio notesTheme frequency by coded domainNot applicableUse only transcripts or structured notes with valid consent and anonymized IDqualitative_
code_matrix
Mixed-methods explanation of interaction quality, self-evaluation, and learner engagement

Table 1: Questionnaire structure, scoring rules, and analytic use. This table defines the locked scoring map established prior to data cleaning and analysis. T1 refers to the baseline assessment, T2 denotes the immediate post-intervention assessment, and T3 indicates the follow-up assessment. Scale scores are calculated as the arithmetic mean of valid items following reverse coding procedures. Entirely missing assessment waves are not imputed.

MeasureWaveGroupNMeanSD
Academic self-conceptT1 baselineUsual student-centered873.050.43
Academic self-conceptT1 baselineAI-assisted student-centered instruction8730.37
Academic self-conceptT2 post-interventionUsual student-centered873.190.6
Academic self-conceptT2 post-interventionAI-assisted student-centered instruction873.510.48
Academic self-conceptT3 follow-upUsual student-centered873.20.63
Academic self-conceptT3 follow-upAI-assisted student-centered instruction873.570.6
Communication competenceT1 baselineUsual student-centered873.040.38
Communication competenceT1 baselineAI-assisted student-centered instruction872.850.4
Communication competenceT2 post-interventionUsual student-centered873.120.62
Communication competenceT2 post-interventionAI-assisted student-centered instruction873.220.63
Communication competenceT3 follow-upUsual student-centered873.120.69
Communication competenceT3 follow-upAI-assisted student-centered instruction873.410.71

Table 2: Descriptive primary outcome scores by group and wave. This table presents descriptive metrics for academic self-concept and communication competence in both the AI-assisted student-centered instruction group and the usual student-centered instruction group across the three assessment waves. Abbreviations: SD = standard deviation.

ConstructItemsCronbach’s alphaDecision
Academic self-concept, T160.676Borderline; retained according to the prespecified scoring map
Academic self-concept, T260.828Acceptable
Academic self-concept, T360.876Acceptable
Communication competence, T160.683Borderline; retained according to the prespecified scoring map
Communication competence, T260.877Acceptable
Communication competence, T360.897Acceptable
Interaction quality, T250.76Acceptable
Self-evaluation, T250.808Acceptable
Learner engagement, T260.806Acceptable

Table 3: Internal consistency estimates for study measures. This table reports Cronbach’s alpha values for the primary outcome scales and process-related scales. Borderline baseline reliability estimates were retained to strictly adhere to the prespecified scoring map.

AI-assisted change meanAI-assisted change SDComparison change meanComparison change SDMean difference95% CI for differenceWelch t p valueCohen’s d
0.510.440.140.460.37[0.24, 0.50]5.28<0.0010.8
0.570.540.150.480.42[0.27, 0.57]5.46<0.0010.83
0.370.450.080.460.29[0.15, 0.43]4.14<0.0010.63
0.570.560.090.570.48[0.31, 0.65]5.61<0.0010.85

Table 4: Between-group differences in outcome change scores. This table reports the T2–T1 and T3–T1 change scores for academic self-concept and communication competence. Between-group differences were evaluated using Welch’s t tests, with 95% confidence intervals and Cohen’s d reported for transparency. Abbreviations: CI = confidence interval; SD = standard deviation.

OutcomeNGroup B95% CI for Group BSEtp valueBaseline outcome BBaseline p valueModel R²
Academic self-concept, T21740.328[0.193, 0.463]0.15<0.0010.748<0.0010.395
Academic self-concept, T31740.401[0.246, 0.556]0.15<0.0010.827<0.0010.384
Communication competence, T21740.301[0.158, 0.444]0.14<0.0011.116<0.0010.488
Communication competence, T31740.516[0.340, 0.692]0.16<0.0011.095<0.0010.388

Table 5: Baseline-adjusted ANCOVA results for primary outcomes. This table presents the baseline-adjusted ANCOVA models for post-intervention and follow-up academic self-concept and communication competence, including group coefficients, 95% confidence intervals, standard errors, p values, and model R2 values. The instructional group variable was coded as 0 for usual student-centered instruction and 1 for AI-assisted student-centered instruction. Abbreviations: CI= confidence interval; SE= standard error.

Outcome changePredictorBSEtp valueModel R2Interpretation
Academic self-concept changeInstructional group0.2990.0783.85<0.0010.191Group association retained
Academic self-concept changeInteraction quality-0.0140.073-0.200.844Not independently significant
Academic self-concept changeSelf-evaluation0.1080.0651.660.099Positive, non-significant trend
Academic self-concept changeLearner engagement0.0610.0660.930.355Not independently significant
Communication competence changeInstructional group0.2520.0823.090.0020.141Group association retained
Communication competence changeInteraction quality0.1310.0751.750.082Positive, non-significant trend
Communication competence changeSelf-evaluation0.1110.0661.680.094Positive, non-significant trend
Communication competence changeLearner engagement-0.0830.066-1.250.214Not independently significant

Table 6: Exploratory regression and class-level diagnostic results for outcome changes. This table reports exploratory linear associations among instructional group, interaction quality, self-evaluation, learner engagement, and outcome change scores, together with class-level ICC diagnostics. These models were not interpreted as confirmatory mediation tests. Abbreviations: SE = standard error.

Supplementary Table 1: Six-session classroom lesson guide. It summarizes the learning objectives, pedagogical phase allocations, and evidence retained for analysis.Please click here to download this file.

Supplementary Table 2: AI prompt sheet for the intervention group. It outlines permitted functions, prohibited uses, and required student records.Please click here to download this file.

Supplementary Table 3: Teacher prompt sheet for the comparison group. It outlines the standardized teacher-guided prompts used to match the instructional structure of the AI-assisted group.Please click here to download this file.

Supplementary Table 4: Student worksheet template. It provides the required fields, formats, and data collection status.Please click here to download this file.

Supplementary Table 5: Self-evaluation form. It summarizes the reflection items, response formats, and respective coding rules.Please click here to download this file.

Supplementary Table 6: Classroom observation and fidelity checklist. It provides the component-specific scoring rubrics from 0 to 2.Please click here to download this file.

Supplementary Table 7: Technical incident log. It documents incident types, affected student counts, and protocol deviation codes.Please click here to download this file.

Supplementary Table 8: Semi-structured interview guide. It details the qualitative domains, core questions, optional probes, and linked constructs.Please click here to download this file.

Supplementary Table 9: Qualitative codebook. It defines the domains, inclusion and exclusion rules, and specific evidence sources.Please click here to download this file.

Supplementary Table 10: Qualitative code frequency matrix. It reports the coded case counts across both instructional groups.Please click here to download this file.

Supplementary Table 11: Data dictionary. It defines variable names, labels, permitted values, missing value rules, and analysis uses.Please click here to download this file.

Supplementary Table 12: Scoring specifications. It details reverse-coded items, minimum valid item requirements, scale formulas, and output variables for each construct.Please click here to download this file.

Supplementary Table 13: Data cleaning specifications. It outlines the cleaning steps, rules, and generated output logs for dataset verification.Please click here to download this file.

Supplementary Table 14: Statistical analysis specifications. It summarizes the analytical procedures, models, and decision thresholds for primary and secondary analyses.Please click here to download this file.

Supplementary Table 15: Cluster-sensitive analysis specifications. It describes the procedures and interpretation rules for class-level aggregation and protocol deviation reviews.Please click here to download this file.

Supplementary Table 16: Source data index. It links the final output tables and figures to their required statistics, source files, and manuscript locations.Please click here to download this file.

Discussion

This study examined whether integrating AI-assisted tools into student-centered instruction was associated with changes in adolescents’ academic self-concept and self-reported communication competence. Throughout the three-wave observation period, students in the AI-assisted cohort showed larger gains in both outcomes than those receiving usual student-centered instruction. This trajectory was consistent across raw score patterns, change-score metrics, and baseline-adjusted ANCOVA models. However, because the design used intact classes, self-reported outcomes, and a small number of classroom clusters, these findings should be interpreted as promising classroom-based associations rather than definitive causal evidence of improved academic performance.

The observed improvement in academic self-concept is consistent with the premise that students’ beliefs about their academic capabilities are shaped by cumulative task experiences and ongoing feedback. Academic self-concept operates not as a generalized sense of self-esteem, but as a domain-specific construct responsive to perceived competence during specific learning activities. The intervention required students to generate initial explanations, critically examine AI-generated suggestions, debate these ideas with peers, and refine their final responses. These iterative cycles may have provided students with visible evidence of progress, offering a plausible explanation for the observed academic self-concept pattern8.

Results concerning communication competence point to a similar classroom process. The AI-assisted group was not merely exposed to supplementary information; the instructional design required students to articulate, compare, and justify concepts during peer interactions. Feedback therefore functioned as an ongoing sequence of AI prompts, peer rebuttals, and targeted teacher clarifications. This dialogic structure may help explain the continued increase in self-reported communication competence observed at follow-up. In the present context, the largest developmental patterns emerged when students were required to translate feedback into revised verbal explanations rather than passively absorb it9.

Qualitative thematic analysis provides essential context for these quantitative trajectories. Interview narratives from the AI-assisted group predominantly featured accounts of AI-supported peer explanation, rigorous fact-checking of uncertain outputs, and transparent self-evaluation practices. These themes mirror the intended structure of the intervention, wherein AI-generated text was positioned strictly as raw material for critical evaluation rather than definitive truth. This epistemological framing is important because the actual educational utility of AI hinges entirely on how learners interact with the output. When students simply copy AI responses, the cognitive process remains limited. Conversely, when outputs are critically questioned, contrasted, and iteratively revised, they may support a more active explanatory process. The dynamics observed thus resemble a guided self-explanation paradigm rather than a simple technology-exposure effect10.

Self-evaluation emerged as another critical classroom process, even though exploratory regression models did not definitively identify it as an independent mediator. Students in the AI-assisted condition reported elevated self-evaluation scores, and qualitative coding highlighted a higher frequency of behaviors related to recognizing knowledge gaps and rationalizing revisions. However, when entered into a multipredictor model alongside group assignment and other mechanism variables, self-evaluation did not achieve statistical significance as a standalone predictor of outcome change. A measured interpretation is that while self-evaluation was an integral component of the learning routine, its influence was likely enmeshed with interaction quality, teacher clarification, and overall task engagement. This synergistic view aligns well with formative assessment frameworks, which assert that self-regulated learning matures through a holistic combination of feedback, standard-setting, continuous monitoring, and structured opportunities to close performance gaps11.

A similar degree of interpretive caution must be applied to learner engagement. Although the AI-assisted group showed higher overall engagement, as corroborated by qualitative accounts of active questioning and task persistence, participation was not uniformly robust across the entire cohort. Interview data revealed that actual engagement remained highly contingent on perceived task authenticity and the efficacy of teacher moderation. This variability is theoretically consistent; engagement is not a monolithic behavior but a multidimensional construct encompassing behavioral, emotional, and cognitive facets, all of which are deeply sensitive to immediate classroom contextual shifts12. Therefore, these results do not imply that AI integration automatically enhances student engagement. Rather, they indicate that engagement is optimized when AI-driven inquiry is tightly coupled with authentic tasks, structured peer dialogue, and active teacher oversight.

Beyond these specific variables, this study offers a targeted classroom-level contribution to the broader discourse surrounding AI in education. Much of the contemporary debate revolves heavily around macroscopic issues such as tool accessibility, algorithmic personalization, task automation, or the existential risk of cognitive overreliance. The current study deliberately deployed AI in a far more constrained capacity: serving solely as an iterative prompt generator embedded within a structured learning cycle. This pedagogical restraint is crucial. AI was never positioned as a surrogate for teacher expertise or peer socialization. It was leveraged exclusively to generate cognitive friction—producing material that students were required to interrogate, debate, and refine. Such an instructional arrangement directly addresses a prevalent limitation in existing AI-in-education research, which frequently overemphasizes raw technological capability at the expense of the specific pedagogical conditions necessary to harness it13.

These findings must be interpreted in light of several methodological limitations. First, the reliance on intact classes rather than individual random assignment limits causal inference. Although lower baseline communication competence in the AI-assisted group was addressed through baseline-adjusted models, regression to the mean and unmeasured classroom-level differences cannot be ruled out. Second, only six classroom clusters were available. ICCs for primary change scores ranged from 0.105 to 0.165, indicating classroom-level dependence, but the small number of clusters limited the stability of multilevel modeling and cluster-adjusted inference. Third, the follow-up interval was relatively short (four weeks), precluding conclusions about durability over a full academic term. Fourth, the process analyses were exploratory and did not establish a verified causal mediation chain.

Measurement and implementation constraints also warrant consideration. Two baseline reliability estimates were below the preferred threshold, which may have introduced baseline measurement noise. The outcomes were mainly self-reported and therefore do not establish whether actual academic performance, objective communication behavior, or independently rated achievement improved. Alternative explanations, including novelty effects associated with AI use, teacher enthusiasm, classroom culture variation, unequal digital familiarity, and differences in peer-group climate, may have contributed to the observed patterns. Attendance records, technical incident logs, and fidelity checklists mitigated some implementation risks, but they cannot fully capture individual variation in how students interpreted or used AI prompts14.

Future research should rigorously test this instructional logic across a broader range of classroom clusters and extend the follow-up timeline to assess longitudinal retention. A fully randomized cluster trial would enable a more precise estimation of distinct teacher and class-level effects. Subsequent investigations should also explore whether these developmental patterns remain consistent across diverse academic disciplines, grade levels, and student ability groups. Furthermore, disentangling specific modalities of AI support (e.g., explanation generation, revision feedback, and self-evaluation prompting) would yield more precise insights than conceptualizing AI assistance as an undifferentiated intervention. Mixed-methods designs will remain indispensable for advancing this line of research, as the interplay among implementation fidelity, student interpretation, and classroom interaction frequently dictates why a standardized intervention succeeds or fails across different educational settings15.

Overall, this study suggests that AI-assisted, student-centered instruction may support adolescents’ academic self-concept and self-reported communication competence when the technology is embedded within a structured, interactive classroom routine. The main implication is not that AI alone improves learning. Rather, the evidence supports a more measured conclusion: AI prompts are most educationally useful when students are required to question, discuss, revise, and critically evaluate AI-generated outputs under sustained teacher guidance. This balanced approach is relevant for educators seeking to use emerging technologies while preserving peer interaction, teacher moderation, and reflective human learning.

Disclosures

Lingyun Deng: Conceptualization, Methodology, Investigation, Data curation, Formal analysis, Writing – original draft.

Shanshan Han: Conceptualization, Methodology, Supervision, Project administration, Validation, Writing – review and editing.

Fengni Yang: Participant recruitment, questionnaire data collection, semi-structured interview recording, supplementary material sorting and manuscript polishing. All authors have read and agreed to the published version of the manuscript..

Acknowledgements

The authors would like to express their sincere gratitude to the participating school, students, and parents for their cooperation and support during the data collection process. The authors also thank the College of Education at Kookmin University for providing the academic environment and necessary resources that facilitated this research. No external funding or financial support was received for the development of this study or its implementation.

Materials

List of materials used in this article
NameCompanyCatalog NumberComments
ChatGPTOpenAIhttps://chatgpt.comInstitutionally approved platform used as an instructional component for student inquiry and prompts.
GrammarlyGrammarly, Inc.https://www.grammarly.comUsed during manuscript preparation for language polishing and formatting support
SPSSIBM Corp.SCR_002865Used for quantitative data analysis, including baseline-adjusted ANCOVA, t-tests, Pearson correlations, and ICCs.
NVivoLumiverohttps://lumivero.com/products/nvivoUsed to code semi-structured interview transcripts and calculate Cohen's kappa for independent coders.
QualtricsQualtrics, LLChttps://www.qualtrics.comUsed to administer baseline, post-intervention, and follow-up assessment questionnaires.

References

  1. Freeman S et al. Active learning increases student performance in science, engineering, and mathematics. Proc Natl Acad Sci U S A. 2014;111(23):8410-8415.
  2. Marsh HW, Craven RG. Reciprocal effects of self-concept and performance from a multidimensional perspective. Perspect Psychol Sci. 2006;1(2):133-163.
  3. Hattie J, Timperley H. The power of feedback. Rev Educ Res. 2007;77(1):81-112.
  4. Nicol DJ, Macfarlane-Dick D. Formative assessment and self-regulated learning: A model and seven principles of good feedback practice. Stud High Educ. 2006;31(2):199-218.
  5. Fredricks JA, Blumenfeld PC, Paris AH. School engagement: Potential of the concept, state of the evidence. Rev Educ Res. 2004;74(1):59-109.
  6. Creswell JW, Plano Clark VL. Designing and conducting mixed methods research. 3rd ed. SAGE Publications; Thousand Oaks, CA; 2018.
  7. Webb NM. The teacher's role in promoting collaborative dialogue in the classroom. Br J Educ Psychol. 2009;79(1):1-28.
  8. Chi MTH, de Leeuw N, Chiu MH, LaVancher C. Eliciting self-explanations improves understanding. Cogn Sci. 1994;18(3):439-477.
  9. Zimmerman BJ. Becoming a self-regulated learner: An overview. Theory Pract. 2002;41(2):64-70.
  10. Ryan RM, Deci EL. Self-determination theory and the facilitation of intrinsic motivation, social development, and well-being. Am Psychol. 2000;55(1):68-78.
  11. Shute VJ. Focus on formative feedback. Rev Educ Res. 2008;78(1):153-189.
  12. Lo CK. What is the impact of ChatGPT on education? A rapid review of the literature. Educ Sci. 2023;13(4):410.
  13. Webb NM. Peer interaction and learning in small groups. Int J Educ Res. 1989;13(1):21-39.
  14. Hmelo-Silver CE. Problem-based learning: What and how do students learn? Educ Psychol Rev. 2004;16(3):235-266.
  15. Fàbregues S et al. Use of mixed methods research in intervention studies to increase young people's interest in STEM: A systematic methodological review. Front Psychol. 2022;13:956300.

Reprints and Permissions

Request permission to reuse the text or figures of this JoVE article

Request Permission

Tags

BehaviorArtificial intelligence in education

Related Articles