All methods involving human participants were conducted in accordance with institutional guidelines and the declaration of Helsinki. The study protocol was reviewed and approved by the institutional review board of Guangzhou city construction college (IRB approval no. GZCJ2026A016; approval date: January 2, 2026). The approved study title was “ai-assisted student-centered instruction and adolescents’ academic self-concept and communication competence: a mixed-methods study.” Written permission was obtained from the participating school before the study was implemented. Written informed consent was obtained from parents or legal guardians, and written assent was obtained from all participating students. Students were informed that participation was voluntary and that they could withdraw from the study at any time without academic penalty. No names, contact details, school identifiers, grades, or other directly identifiable information were entered into the ai chatbot platform or included in the public dataset. All study procedures were conducted during the approved study period from January 2, 2026, to January 2, 2027. All tools and platforms used in this study are listed in the Table of Materials
Study design and classroom assignment
This study employed a quasi-experimental, mixed-methods classroom design with two instructional conditions: an AI-assisted, student-centered instruction group and a usual student-centered instruction group. Longitudinal measurements were conducted at three time points: baseline before the intervention (T1), immediately after the final instructional session (T2), and at a predefined short-term follow-up assessment (T3). The T2–T3 interval was recorded in the study log and kept consistent across the manuscript, dataset, analysis scripts, and figure labels. Figure 1 summarizes recruitment, class assignment, assessment timing, exclusions, and qualitative sampling. Because administrative and ecological constraints prevented individual randomization, six intact classes were assigned at the classroom level, with three classes in the AI-assisted condition and three classes in the comparison condition. Each class contributed 29 students to the final analytic sample. Grade level, subject area, lesson sequence, session length, classroom tasks, assessment schedule, and teacher contact time were kept comparable across conditions. Class identifiers were retained for cluster-sensitive statistical checks.
Participant recruitment and eligibility screening
Participants were recruited directly from the participating intact classes. Students and their parents or legal guardians received a detailed information sheet outlining the study purpose, classroom procedures, AI-supported activities, questionnaire schedules, the optional interview component, privacy protections, withdrawal rights, and data storage plans. Students were included in the study if they were enrolled in one of the participating classes, attended regular classroom instruction, provided written assent, obtained written parental or guardian consent, and completed the baseline questionnaire before the first session. Conversely, students were excluded from the final analytic sample if they transferred from the school, missed more than 25% of the intervention sessions, submitted invalid questionnaire records, or withdrew permission for data use. One primary reason for exclusion was documented for each excluded participant. Prior to data entry, each eligible participant was assigned a unique study identification code. Relevant demographic and educational covariates were recorded, including age, gender, class identifier, baseline academic performance band, prior exposure to AI-assisted learning, and baseline digital learning familiarity. For statistical tracking, the instructional group variable was coded as 0 for the comparison condition and 1 for the AI-assisted condition.
Instructional materials and intervention delivery
A six-session instructional unit (Supplementary Table 1) was developed prior to implementation, with each classroom session lasting 40–45 min. Identical learning objectives, lesson topics, worksheets, peer discussion tasks, self-evaluation forms, and reflection prompts were applied across both instructional conditions. The questionnaire structure, scoring directions, reverse-scored items, and final scale-level variables followed the locked scoring map in Table 1. Reproducibility materials included lesson plans (Supplementary Table 1), AI prompt sheets (Supplementary Table 2), teacher prompt sheets (Supplementary Table 3), student worksheets (Supplementary Table 4), self-evaluation forms (Supplementary Table 5), observation checklists (Supplementary Table 6), technical incident logs (Supplementary Table 7), interview guides (Supplementary Table 8), and a qualitative codebook (Supplementary Table 9). The participating teacher completed training using a written delivery guide covering session sequence, AI-use boundaries, comparison-group procedures, privacy requirements, and fidelity checklists. A rehearsal with non-study students was conducted to verify activity timing, device access, worksheet clarity, and task difficulty. The AI tool was an institutionally approved ChatGPT-based chatbot using the GPT-4 architecture. For classroom use, chat history, profile memory, web browsing, plug-ins, image generation, and automated grading were disabled. The manuscript uses the generic term "AI chatbot platform" after its first mention to avoid promotional language. Prompt-design principles were fixed before the intervention and limited AI use to clarification, example generation, explanation comparison, targeted revision support, and self-evaluation cues. Representative prompts, sample AI outputs, and teacher moderation examples were retained as a part of the reproducibility materials.
Each AI-assisted session was conducted in five structured phases. First, a 5 min teacher briefing was conducted to state the learning goals, explain tasks, identify the configured AI chatbot platform, and remind students that AI outputs were provisional prompts rather than authoritative answers. Students were strictly instructed not to enter any identifiable personal information into the platform. Second, a 10 min individual AI-supported inquiry activity followed, during which students utilized the prepared prompt sheet to request clarifications, examples, alternative explanations, or revision suggestions. Each student recorded the prompt category used, one useful AI-supported idea, one point requiring verification, and their final adoption decision. Third, a 12–15 min peer discussion activity was conducted, during which students compared their AI-supported ideas in small groups, identified inaccuracies, and collaboratively revised their work. During this phase, the teacher circulated throughout the classroom to apply a standardized moderation rule, prompting students to verbally justify their revisions against lesson criteria and correcting erroneous or overconfident AI outputs before finalization. Fourth, a 5–8 min whole-class clarification session was conducted to address widespread misconceptions and correct factual errors. Finally, a 5 min self-evaluation activity concluded the session, during which students rated their understanding, recorded one concrete improvement, and noted one unresolved question. This exact operational sequence was maintained across all intervention classes. Printed prompt sheets served as technical backups, and technical deviations affecting over 1/3 of a class were explicitly logged. Anonymized representative prompt-output-moderation examples were retained and documented in Supplementary Tables 2, 3.
The comparison condition was delivered utilizing the identical session length, learning objectives, topic sequence, worksheets, peer discussion intervals, self-evaluation forms, and teacher feedback windows. However, the AI-supported prompts were replaced with standardized, teacher-prepared guiding prompts that instructed students to explain core concepts, compare examples, identify weaknesses, and revise their work. The identical five-step structure encompassing the briefing, individual inquiry, peer discussion, whole-class clarification, and final self-evaluation was strictly maintained. No additional tutoring, out-of-class assignments, prolonged feedback periods, or different assessment tasks were provided to either condition. Classroom attendance, behavioral disruptions, teacher substitutions, and incomplete components were documented using the same logging format for both groups.
Data collection, fidelity monitoring, and quantitative analysis
Assessment questionnaires were administered at baseline, post-intervention, and follow-up intervals, utilizing consistent item wording, response scales, and scoring rules across all three measurement waves. The instruments comprised validated scales evaluating academic self-concept, communication competence, interaction quality, self-evaluation, and learner engagement. Responses were gathered using a standard 5-point Likert-type scale ranging from 1 (strong disagreement) to 5 (strong agreement), with higher scores indicating higher levels of the target construct. Individual scale scores were calculated as the arithmetic mean of valid items following necessary reverse coding. Missing item values were replaced with the student's mean score within the same scale and wave, provided that no more than 20% of the items were missing. Scales with more than 20% missing items were treated as missing, and data from an entirely missing assessment wave were not imputed. Prior to statistical analysis, the raw dataset was carefully audited for range violations, duplicate entries, straight-line response patterns, and anomalous completion times. Records were flagged as invalid and excluded if they exhibited impossible numerical values, invariant responses across all items, or completion times below 1/3 of the median duration, unless explicit classroom logs validated their authenticity. Main post-intervention and follow-up analyses included students who provided valid baseline scores and corresponding post-intervention or follow-up outcome data, thereby avoiding unilateral listwise deletion for partially complete schedules. A locked data dictionary was established prior to data entry to define variable names, labels, response ranges, missing value codes, reverse-scoring rules, wave identifiers, and final scale-level constructs. Questionnaire data were entered in a wide format, with a single row per student and wave suffixes (_T1, _T2, _T3) used to distinguish the measurement waves. Numerical ranges and unique participant records were computationally verified, and missingness patterns were inspected across groups, classes, and waves against the master attendance log.
A standardized classroom observation checklist was employed during every instructional session to monitor implementation fidelity. Each pedagogical component was scored as 0 (not completed), 1 (partially completed), or 2 (fully completed). The evaluated components strictly aligned with the five-phase instructional sequence, including the teacher briefing, individual inquiry, peer discussion, whole-class clarification, and the completion of self-evaluation records. The overall session fidelity percentage was calculated by dividing the observed score by the maximum possible score. Any session falling below a 70% fidelity threshold was recorded as a protocol deviation. To ensure inter-rater reliability, a second trained observer independently rated at least 20% of the total sessions. Percentage agreement was calculated across all items, and any rating discrepancies were resolved through consensus discussion whenever the initial agreement fell below 80%. Hardware or software technical malfunctions were recorded in separate logs to distinguish them from instructional fidelity issues.
Statistical analyses were performed using a reproducible, script-based workflow in which the random seed was fixed prior to any sampling or sensitivity checks. All statistical tests were two-tailed, with significance set at p < 0.05. Continuous variables were summarized as means and standard deviations, while categorical variables were expressed as frequencies and percentages. Baseline equivalence between the cohorts was evaluated descriptively using independent-samples t-tests for continuous data and chi-square tests for categorical distributions. Internal consistency was measured using Cronbach’s alpha, applying thresholds of α ≥ 0.70 for general scales and α ≥ 0.80 for primary outcome measures, alongside corrected item-total correlations. Baseline-adjusted ANCOVA served as the primary analytical approach. Separate models were fitted for each primary outcome, with the post-intervention or follow-up score specified as the dependent variable, the instructional group as the main predictor, and the corresponding baseline score as a mandatory covariate. Additional covariates (gender, baseline academic performance band, and prior AI exposure) were included to control for baseline imbalances or their theoretical relevance, and adjusted mean differences, 95% confidence intervals, standard errors, and model R2 values were reported. Change-score comparisons (T2–T1 and T3–T1) were performed as secondary analyses using independent-samples Welch's t-tests, with effect sizes quantified using Cohen’s d. Exploratory process analyses evaluated Pearson correlations among process variables and outcome changes, supplemented by multi-predictor linear regressions with variance inflation factor (VIF) monitoring to check for multicollinearity. Finally, cluster-sensitive checks were implemented to account for the nested structure of students within the six intact classes. Unconditional intraclass correlation coefficients (ICCs) were estimated, outcome changes were aggregated and compared at the class level, and primary ANCOVA calculations were repeated using cluster-robust standard errors as an exploratory sensitivity check.
Qualitative sampling, coding, and mixed-methods integration
Immediately following the post-intervention questionnaire, a purposive subsample of 36 students was selected for semi-structured interviews. This qualitative cohort included participants from both instructional conditions who demonstrated high, medium, or low change scores in academic self-concept or communication competence. The semi-structured interview guide investigated experiences of classroom interaction, confidence in explaining ideas, self-evaluation behaviors, task engagement, and willingness to communicate, using targeted prompts to explore moments of revision, responses to unclear feedback, and variations in engagement. Interviews were audio-recorded only after obtaining parental consent and student assent; otherwise, structured written notes were taken. All personal names, class numbers, and identifying details were removed during transcription, and transcripts were linked to their corresponding quantitative identifiers. The initial qualitative coding framework was deductively derived from interaction quality, self-evaluation, and learner engagement constructs, with inductive thematic codes added only when the data failed to fit the baseline framework. Two independent coders analyzed at least 20% of the transcripts to verify reliability. Cohen's kappa coefficients were evaluated, with values between 0.61–0.80 treated as substantial agreement and values above 0.80 as strong agreement. Codebook definitions were refined and double-coding was repeated whenever the initial kappa fell below 0.61, and the final definitions alongside the resulting qualitative code frequency matrix were archived (Supplementary Tables 9 and 10).
Quantitative data analysis and qualitative coding frameworks were finalized independently before initiating mixed-methods integration. A cohesive joint display was structurally constructed to link the three quantitative process variables with their corresponding qualitative themes. Integrated findings were systematically classified as convergence (where questionnaire patterns and interview themes aligned), expansion (where qualitative narratives explained the mechanisms underlying quantitative trends), or divergence (where interviews revealed unique barriers or uneven participation that complicated the quantitative patterns). Qualitative findings were utilized to contextualize classroom processes, linking learner engagement scores with themes of active questioning and peer discussion, and linking self-evaluation scores with accounts of gap awareness and strategic adjustments.