This study benchmarks STEP-DWCF-R (Structured, Tiered, Evidence-driven Process for Dynamic Written Corrective Feedback with Robotic AI Agent) for improving IELTS writing performance through multi-round AI and teacher-supported revisions.
Method Article
This study benchmarks STEP-DWCF-R (Structured, Tiered, Evidence-driven Process for Dynamic Written Corrective Feedback with Robotic AI Agent) for improving IELTS writing performance through multi-round AI and teacher-supported revisions.
Automated writing feedback systems are prevalent, yet most deliver static, fragmented comments that provide limited scaffolding for revision. This study evaluates the STEP-DWCF-R framework (Structured, Tiered, Evidence-driven Process for Dynamic Written Corrective Feedback with Robotic AI Agent), in which AI-generated feedback, moderated by a teacher, is delivered via a robotic AI agent across multiple iterative rounds within a one-week task cycle. In an eight-week quasi-experimental trial, 32 EFL learners were randomized to either traditional written corrective feedback (one round per task) or STEP-DWCF-R. Both groups completed IELTS Task 2 essays at baseline and post-test, which were anonymized, randomized, and scored by two independent raters (ICC = 0.86–0.93). Linear mixed-effects models demonstrated that the STEP-DWCF-R group exhibited significantly greater gains in overall band score (Δ = 1.03 vs. 0.31 bands) and across all four analytic dimensions, with the largest improvement observed in Coherence and Cohesion. Process data indicated that STEP-DWCF-R learners completed an average of 2.26 revision rounds per task, with error counts decreasing linearly across rounds. These findings suggest that the integrated STEP-DWCF-R framework, encompassing AI-generated feedback, teacher moderation, and iterative robotic AI agent delivery, is associated with greater IELTS writing improvement than traditional single-round feedback, pointing to practical applications for AI-enhanced dynamic feedback in EFL contexts.
Written corrective feedback (WCF) remains central to L2 writing pedagogy. Evidence shows that comprehensive WCF improves learners’ accuracy over time and, when aligned with classroom practice, can coexist with focused approaches that are feasible in authentic contexts1,2,3. Recent classroom studies show durable gains in accuracy and fluency from sustained, comprehensive WCF. Research on feedback scope cautions that teachers should target forms strategically rather than mark everything4,5,6,7,8,9. In parallel, automated writing evaluation (AWE) and automated written corrective feedback (AWCF) have grown rapidly. Findings are mixed: some studies show gains in task achievement and grammatical range with AWCF or Criterion-style programs, while others find no significant advantage over teacher-only feedback or note largely local, surface-level revisions. To reflect this heterogeneity, we deliberately triangulate across multiple AWCF studies rather than leaning on a single source10,11,12,13.
Learners’ beliefs about who is providing feedback also matter: quasi-experimental work on perceived source indicates that performance and trust can shift when identical feedback is framed as “teacher” versus “automated,” underscoring the need to account for perception effects in AWCF designs14,15. Extending this line, emerging studies suggest that educational robots, due to their embodied presence, can also act as feedback providers, potentially influencing learner engagement and trust in distinctive ways16.
A complementary line of research, dynamic written corrective feedback (DWCF), emphasizes frequent, manageable, comprehensive feedback cycles with coding and rapid revision. Evidence from multiple settings suggests DWCF reliably improves accuracy (with frequency effects on fluency), though results may vary by course goals and learner population17,18,19,20. Finally, early studies on generative-AI-mediated feedback report parity with teacher feedback on writing outcomes and potential benefits when paired with explicit metalinguistic guidance, motivating closer integration of human and AI feedback in L2 contexts21,22.
To address these challenges, we propose the STEP-DWCF-R (Structured, Tiered, Evidence-driven Process for Dynamic Written Corrective Feedback with Robotic AI Agent) framework, a novel multi-stage DWCF feedback system designed to provide structured, iterative, and process-oriented writing support. Our system leverages the predictive and generative power of Large Language Models (LLMs), with delivery through a robotic AI agent interface and teacher moderation, to segment complex writing issues into manageable categories, deliver feedback through multiple rounds of interaction, and evaluate its effectiveness using both outcome-based and process-based metrics. By transforming feedback into a structured and interactive process, STEP-DWCF-R advances automated writing support from static error correction toward a more engaging, embodied, and learning-oriented system.
Our contributions center on three integrated advances. First, we propose a structured feedback representation schema that categorizes errors (e.g., grammar, coherence) so that robotic AI agent-delivered LLM feedback is both actionable and pedagogically interpretable. Second, we introduce an iterative multi-stage process that sequences feedback across language, structure, and reasoning across multiple rounds to reduce cognitive load and deepen engagement. Third, we extend evaluation beyond post-scores to include process-oriented metrics such as error reduction and feedback adoption, offering a richer view of learning impact. WCF frameworks represent the earliest attempts to systematize written corrective feedback within educational technology. Originating in Automated Writing Evaluation (AWE) and Automated Essay Scoring (AES), these systems emphasized error detection and proficiency scoring13,14. Their contribution lies in scalability: thousands of learners could receive feedback without the bottleneck of human grading. However, these frameworks typically provided static and fragmented comments, such as grammar checks, lexical suggestions, or global scores that learners struggled to transform into meaningful revisions23,24,25,26.
Over time, WCF research evolved to include more technology-mediated approaches. These newer systems sought to capture broader aspects of writing by leveraging computational methods13,21. More recently, the scope has expanded further with embodied technologies, where educational robots have begun to act as feedback providers in writing or tutoring contexts. Their physical presence and interactive affordances introduce new dynamics of trust, engagement, and learner motivation. Yet, despite these developments, the dominant orientation of WCF has remained outcome-focused8,27: producing scores or isolated comments in a one-shot manner. Pedagogically, this orientation is misaligned with formative assessment, where feedback should scaffold revision across drafts and support metacognitive engagement16,28,29,30,31. Thus, the conceptual boundary of WCF reveals an enduring gap: feedback is delivered, but not structured or sequenced to guide learning processes.
In response to these limitations, scholars have begun to explore more dynamic approaches to writing feedback. Early investigations highlighted the importance of iterative feedback cycles, where learners revise multiple times under scaffolded guidance17,32. Structured error categorization has also been proposed to improve the interpretability of feedback and align it with pedagogical objectives33,34. The notion of a DWCF framework consolidates these directions by embedding three design principles: explicit error schemas, staged and iterative delivery, and process-oriented evaluation metrics19,35. Rather than overwhelming students with a bulk of corrections at once, DWCF distributes attention across layers of writing—surface accuracy, structural coherence, and argumentative depth—through successive rounds17,33. Evaluation expands beyond final scores to include process indicators such as error reduction rates, feedback adoption, and revision depth10,19,25. Despite its promise, practical applications of DWCF remain sparse, and most studies stop short of operationalizing it in large-scale, technology-mediated contexts10,36,37. This gap highlights the need for frameworks that not only theorize DWCF but also demonstrate its feasibility in real-world educational settings. Figure 1 shows the broader theoretical framework of DWCF proposed in the literature. The figure provides a general rationale for how dynamic written corrective feedback may operate at multiple levels, including both measured outcomes (e.g., accuracy, coherence) and theorized but unmeasured constructs in the present study (e.g., writing anxiety reduction, fluency, and complexity). It is included for theoretical orientation rather than as a direct model of this trial. Elements corresponding to the outcomes and process metrics analyzed here are indicated in the figure; other pathways are illustrative and were not evaluated in this study.
The emergence of LLMs and multimodal systems offers a powerful means to operationalize DWCF at scale. LLMs, with their zero-shot and few-shot capabilities, can generate context-sensitive feedback without task-specific training21,22. Recent experiments suggest their potential for directive segmentation, staged refinement, and explanation generation, which align closely with the principles of DWCF13,37,38,39. Complementary research in human–AI collaboration has shown that iterative loops of engaging with, adapting, and selectively adopting AI feedback can improve both uptake and revision quality25,30. When these processes are embodied through robots, the interaction has the theoretical potential to become multimodal—combining gesture, voice, and presence—which may further enhance learner perception and motivation40.
Grounded in the Zone of Proximal Development (ZPD)41, the Input Hypothesis42, and the Skill Acquisition Theory43, we position STEP-DWCF-R as a controller that links learner support, input processing, and skill development. Concretely, rubric-linked itemization, timely consolidated lists, and multi-round teacher-moderated AI–, human–, and robot–cycles operate at an integration layer that mediates revision and yields the observable endpoints analyzed here (IELTS Overall and TR/CC/LR/GRA; rounds, uptake, persistent errors, time). Constructs such as writing anxiety reduction are theorized but not measured in this trial.
This protocol was approved by the Ethics Committee of the School of Primary Education, Shangrao Preschool Education College. All participants were adults (≥18 years) and provided written informed consent prior to the study.
1. Study design
NOTE: Due to the constraints of the available EFL classroom (total enrollment N = 32), participants were randomized into two groups of 16 each. This sample size, while modest, was determined to be sufficient for a pilot investigation given the within-subject pre-post design and expected large effect sizes based on prior DWCF studies17,18,19,20.
2. Writing Tasks
3. Feedback Workflow
4. Outcome Measurement
5. Process Measurement
6. Data Analysis
7. Code Availability
All code, prompts, JSON schemas, and example implementation files required to reproduce the STEP-DWCF-R system are openly available at https://github.com/YifangGaoinPG/Dynamic-Writing-Correction-Feedback.
Experiments and Analysis
All 32 learners provided baseline data. Post-test data were available for 16 learners in each group (WCF and STEP-DWCF-R; Figure 2). The study timeline and measures are presented in Figure 3, and the end-to-end feedback workflow is illustrated in Figure 4. Experimental setup and implementation parameters for our AI system are shown in Table 1. Analyses followed the pre-specified plan with intention-to-treat. We first assess baseline equivalence (Table 2), then report the primary outcome (Figure 5 and Table 3), followed by secondary outcomes (Table 4) with robustness and reliability checks.
Baseline Equivalence
At baseline (Week 1), the groups were descriptively similar on pre-specified covariates, including demographics and IELTS writing performance before any feedback (Table 2). Individual IELTS component scores are assigned in 0.5-band steps, whereas the table reports group means. The Week 1 overall means (WCF = 5.625, STEP-DWCF-R = 5.500) are computed from the same arrays used to generate Figure 5. We do not perform hypothesis tests at baseline; any residual imbalance is addressed by the pre-post design and the Group × Time term in the mixed-effects model, and a post-test comparison adjusted for baseline is reported in the Protocol section.
Primary Outcome
Figure 5 summarizes pre–post changes, while Table 3 and Table 5 report model estimates and post-test attainment. At baseline (Week 1), the group means were 5.625 for WCF and 5.500 for STEP-DWCF-R, yielding a non-significant group difference of −0.125 bands (SE = 0.133, t = − .939, df = 29.90, p = .355, 95% CI [−0.397, 0.147]). The baseline in Table 3, therefore, estimates the WCF Week 1 mean at 5.625 with 95% CI [5.419, 5.831] (SE = 0.097, df = 15). From Week 1 to Week 8, WCF improved by 0.3125 bands (SE = 0.128, t = 2.44, df = 15, p = .028, 95% CI [0.039, 0.586]; standardized within-group effect dz = 0.61). STEP-DWCF-R improved by 1.0313 bands (SE = 0.140, t = 7.34, df = 15, p = 2.44 × 10−6, 95% CI [0.732, 1.331]; dz = 1.84). The linear mixed-effects contrast for the Group × Time interaction—i.e., the difference-in-differences—was 0.7188 bands (SE = 0.190, t = 3.78, df = 29.75, p = .0007, 95% CI [0.330, 1.107]; Table 3), indicating larger gains under STEP-DWCF-R. At post-test (Week 8), the estimated marginal means favored STEP-DWCF-R by 0.5938 bands (Welch test on group means: SE = 0.115, t = 5.16, df = 29.74, p = 1.50 × 10−5, 95% CI [0.359, 0.829]). Consistent with this, the attainment table (Table 5) shows that at Week 8, 13% of WCF learners reached Overall ≥ 6.5 and 0% reached ≥ 7.0 (counts 2/16 and 0/16), whereas 81% of STEP-DWCF-R learners reached ≥ 6.5 and 25% reached ≥ 7.0 (counts 13/16 and 4/16). Table 3 reports the corresponding fixed-effect estimates and their uncertainty.
Dimension-Level Outcomes
Table 4 summarizes pre–post changes across the four Task 2 dimensions. Task Response (TR) showed a gain of Δ = 0.25 bands under WCF versus Δ = 0.95 under STEP-DWCF-R, yielding ΔΔ = 0.70 with 95% CI [0.28, 1.12] and Holm-adjusted p = .006. Coherence and Cohesion (CC) exhibited the largest contrast: WCF Δ = 0.30, STEP-DWCF-R Δ = 1.15, hence ΔΔ = 0.85 (95% CI [0.44, 1.26], adj-p = .002). Lexical Resource (LR) followed the same pattern with WCF Δ = 0.35 and STEP-DWCF-R Δ = 0.95, giving ΔΔ = 0.60 (95% CI [0.18, 1.02], adj-p = .012). Grammatical Range and Accuracy (GRA) improved by Δ = 0.25 in WCF and Δ = 0.97 in STEP-DWCF-R; the resulting ΔΔ = 0.72 carried a 95% CI of [0.31, 1.13] with adj-p = .004. Across all four dimensions, the Group×Time contrasts favored STEP-DWCF-R after Holm–Bonferroni adjustment.
Process Evidence
Figure 6 and Figure 7 visualize the core process results. Across Weeks 2–7, STEP-DWCF-R completed on average 2.26 revision rounds per task (95% CI [2.17, 2.35]; n = 96), whereas WCF was 1.00 rounds for all tasks (n = 96). The mean difference was 1.26 rounds (95% CI [1.17, 1.35]); a Mann–Whitney test confirmed separation of the distributions (U = 9216, z = 11.97, p < 10−30). Within STEP-DWCF-R, mean error counts declined by round: 13.78 at Round 1 (95% CI [13.02, 14.54]; n = 96), 8.94 at Round 2 (95% CI [8.42, 9.46]; n = 96), and 6.16 at Round 3 (95% CI [5.20, 7.12]; n = 25). A linear trend fitted to per-round observations estimated a decline of 4.14 errors per round (95% CI [3.51, 4.78] in absolute magnitude; p < 10−12). Measures of feedback uptake and time-on-task are not encoded in Figure 6 and Figure 7 and are therefore reported in Table 7.
Reliability and Robustness
Inter-rater agreement for Week 1 and Week 8 essays was good to excellent. Two trained raters scored all scripts independently and were blinded to group and time; using average-measures intraclass correlation coefficient (ICC[2,2]) on the rater means, reliability ranged from 0.86–0.93 across Overall, and the four dimensions, and all lower 95% confidence limits exceeded 0.80 (Table 6). Diagnostic checks for the primary mixed-effects model indicated approximately normal residuals without material heteroscedasticity, and models fitted to task-level process summaries showed dispersion close to one. Sensitivity analyses based on the same Week 1/Week 8 dataset yielded consistent inferences: a linear regression comparing groups at post-test while adjusting for baseline scores estimated an adjusted between-group difference of 0.575 bands at Week 8 (95% CI [0.336, 0.814], p = 3.15 × 10−5), and a rater-specific mixed model that included a random effect for Rater and used un-averaged scores produced a Group × Time estimate of 0.72 bands (95% CI [0.33, 1.11], p = .0007), aligning with Table 3. Results were unchanged when using robust standard errors and when analyses were restricted to complete cases, reinforcing the conclusion that STEP-DWCF-R yields larger pre-post gains than WCF.
Qualitative Findings
For qualitative analysis, we examined STEP-DWCF-R revision traces across all six tasks and end-of-course written reflections from the 16 participants in the STEP-DWCF-R group. Two coders independently coded the materials using a focused deductive framework based on the four feedback categories (grammar, vocabulary, organization, reasoning) and uptake rationales. Inter-coder agreement was substantial with Cohen’s kappa of 0.78, and disagreements were resolved through discussion.
Analysis revealed five key themes: uptake followed a staged pattern, with learners resolving surface features before addressing organization and reasoning; span-specific, rubric-linked items were more actionable and frequently adopted than narrative suggestions; AI-teacher complementarity shaped priority, with overlapping signals increasing focus and teacher moderation resolving conflicts; benefits peaked early, with reuse of organization templates and lexical patterns noted on new prompts; and learners adapted strategies across tasks. These themes inform the process dynamics explored in the process evidence section. Feedback uptake, persistent errors, and time-on-task measures were also recorded. A summary of these additional process indicators is presented in Table 7.
Interpretation of Representative Results
Taken together, the outcome and process data indicate that the STEP-DWCF-R framework is associated with greater IELTS writing improvement than traditional single-round feedback. The primary outcome shows a between-group difference-in-differences of 0.72 bands, with the STEP-DWCF-R group improving by 1.03 bands overall compared with 0.31 bands in the WCF group. This advantage was consistent across all four analytic dimensions, with the largest contrast in Coherence and Cohesion at 0.85 bands, suggesting that structured, rubric-linked, multi-stage feedback particularly scaffolded organizational development. Process evidence further indicates that iterative engagement underlies this advantage. STEP-DWCF-R learners completed an average of 2.26 revision rounds per task, and error counts declined linearly by approximately 4.1 errors per round. For instance, a representative workflow cycle included GPT-5 generating structured JSON feedback across the four categories, followed by teacher moderation to eliminate false positives and enhance actionability, and subsequent delivery via the robotic AI agent interface for multi-round learner revision. These findings support the feasibility and potential effectiveness of the integrated multi-component workflow combining AI-generated feedback, teacher moderation, and iterative robotic AI agent delivery, but they should not be interpreted as proof of any single component in isolation. The dose imbalance of 2–3 rounds versus one round and the modest sample size of 32 mean that the observed gains reflect the combined effect of increased revision opportunities and structured feedback. These results should be viewed as promising pilot evidence awaiting confirmation in larger, component-controlled trials.

Figure 1. Theoretical framework of DWCF (dynamic written corrective feedback). The figure provides a broader rationale for how DWCF may operate; it is included for orientation rather than as a direct model of this trial. Elements corresponding to the outcomes and process metrics analyzed here are indicated in the figure; other pathways are illustrative and were not evaluated. Please click here to view a larger version of this figure.

Figure 2. CONSORT flow diagram of participants. Thirty-two learners were assessed for eligibility; none were excluded. All participants provided consent and were then randomized in a 1:1 ratio to either the WCF group (n = 16) or the DWCF group (n = 16). All randomized participants received the allocated intervention; there were no losses to follow-up or discontinuations, and all 32 were included in the final analysis. Please click here to view a larger version of this figure.

Figure 3. Study timeline and measures. The trial spanned 8 weeks: at W1, participants completed the baseline IELTS Task 2 and were scored; six weekly writing tasks were administered from W2 to W7 (Task 1–Task 6), with group-specific feedback (WCF or DWCF) and revisions after each task. Process measures, including revision rounds, feedback uptake, persistent-error rate, and time on task, were logged for each task during W2–W7. At W8, the post-test IELTS Task 2 was administered and scored. Please click here to view a larger version of this figure.

Figure 4. End-to-end feedback workflow. Learners produce a proctored initial draft, then receive group-specific feedback: WCF (one human feedback round) or STEP-DWCF-R (AI-generated feedback with teacher moderation, delivered via robotic AI agents, involving 2–3 iterative rounds). Learners complete revision round(s) accordingly. Outcome is the IELTS Task 2 band score; process measures include number of rounds, feedback uptake (adopted/modified/declined), persistent-error rate, and time on task. The system logs item-level IDs, category/severity, source (AI vs. human vs. robot), moderation actions, and uptake status. Please click here to view a larger version of this figure.

Figure 5. Change in IELTS overall band from Week 1 to Week 8 by group. Points give the estimated group means with 95% confidence intervals, and the trajectories indicate a larger gain under STEP-DWCF-R Please click here to view a larger version of this figure.

Figure 6. Revision rounds per task by group (Weeks 2–7). Data points represent individual task revision rounds for the WCF (blue) and STEP-DWCF-R (orange) groups. Horizontal lines and error bars indicate group means with 95% confidence intervals. Please click here to view a larger version of this figure.

Figure 7. Mean error counts per round within STEP-DWCF-R with 95% CIs. Points indicate mean error counts at each feedback round (Round 1, Round 2, Round 3), and vertical lines represent 95% confidence intervals (CIs). Please click here to view a larger version of this figure.
| Component | Specification |
| Hardware | PC with 12th Gen Intel(R) Core(TM) i7-12700H (2.30 GHz), 32GB RAM, NVIDIA RTX 3080Ti GPU (16GB) |
| Operating System | Windows 11 |
| Backend | Flask~2.1 (Python~3.10) |
| Frontend | Jinja2 server-rendered templates |
| AI Models | OpenAI GPT-5 API (for feedback generation), schema-based post-processing |
| Feedback Scheduler | Custom controller managing multi-stage feedback (3 cycles) |
| Error Schema | Grammar, Vocabulary, Organization, Reasoning |
| Version Control | GitHub |
| Logging | Automatic logging of submissions, feedback, and revisions for analysis |
Table 1: Experimental setup and implementation parameters. Technical specifications of the AI feedback system, including hardware configuration, operating system, backend and frontend frameworks, AI model settings, feedback scheduler, error schema categories, version control, and logging infrastructure.
| Variable | WCF (n=16) | STEP-DWCF-R (n=16) |
| Age (years), mean (SD) | 20.9 (1.5) | 21.1 (1.6) |
| Female, n (%) | 9 (56%) | 8 (50%) |
| Prior IELTS preparation (yes), n (%) | 7 (44%) | 6 (38%) |
| Years of prior writing instruction | 3.1 (1.2) | 3.2 (1.1) |
| Baseline IELTS Overall (band) | 5.625 (0.387) | 5.500 (0.365) |
| Baseline TR (band) | 5.56 (0.50) | 5.50 (0.50) |
| Baseline CC (band) | 5.62 (0.49) | 5.53 (0.49) |
| Baseline LR (band) | 5.60 (0.46) | 5.52 (0.46) |
| Baseline GRA (band) | 5.56 (0.48) | 5.49 (0.45) |
Table 2: Baseline characteristics of participants by group (Week 1). Demographic variables and pre-test International English Language Testing System (IELTS) Task 2 writing performance for the WCF and STEP-DWCF-R groups. Values are mean (standard deviation) or n (%).
| Fixed effect | Estimate | SE | df | t | p | 95% CI |
| Intercept (WCF, Week~1) | 5.625 | 0.097 | 15 | 58.1 | <.001 | [5.419, 5.831] |
| Group (STEP-DWCF-R vs. WCF) | -0.125 | 0.133 | 29.9 | -0.939 | .355 | [-0.397, 0.147] |
| Time (Week~8 vs. Week~1) | 0.3125 | 0.128 | 15 | 2.44 | .028 | [0.039, 0.586] |
| Group × Time | 0.7188 | 0.19 | 29.75 | 3.78 | .0007 | [0.330, 1.107] |
Table 3: Linear mixed-effects model for IELTS overall band from Week 1/Week 8 scores. Fixed-effect estimates, standard errors (SE), degrees of freedom (df), t-values, p-values, and 95% confidence intervals (CIs) for the Group × Time interaction and model terms.
| Dimension | WCF Δ (bands) | STEP-DWCF-R Δ (bands) | ΔΔ (95% CI) | adj-p |
| Task Response (TR) | 0.25 | 0.95 | 0.70 (0.28, 1.12) | .006 |
| Coherence and Cohesion (CC) | 0.3 | 1.15 | 0.85 (0.44, 1.26) | .002 |
| Lexical Resource (LR) | 0.35 | 0.95 | 0.60 (0.18, 1.02) | .012 |
| Grammatical Range and Accuracy (GRA) | 0.25 | 0.97 | 0.72 (0.31, 1.13) | .004 |
Table 4: Dimension-level outcomes (Task 2): pre–post changes and between-group differences. Pre–post changes (Δ) and difference-in-differences (ΔΔ) with 95% confidence intervals (CIs) and Holm-adjusted p-values for Task Response (TR), Coherence and Cohesion (CC), Lexical Resource (LR), and Grammatical Range and Accuracy (GRA).
| Threshold | WCF (%) | STEP-DWCF-R (%) |
| Overall ≥ 6.5 | 13 | 81 |
| Overall ≥ 7.0 | 0 | 25 |
Table 5: Post-test band attainment (Week 8). Proportion of learners in the WCF and STEP-DWCF-R groups reaching IELTS Overall band score thresholds of at least 6.5 and at least 7.0.
| Outcome | ICC | 95% CI |
| Overall | 0.91 | [0.86, 0.94] |
| Task Response (TR) | 0.88 | [0.82, 0.92] |
| Coherence and Cohesion (CC) | 0.9 | [0.85, 0.94] |
| Lexical Resource (LR) | 0.89 | [0.83, 0.93] |
| Grammatical Range and Accuracy (GRA) | 0.87 | [0.80, 0.91] |
Table 6: Inter-rater reliability for IELTS outcomes. Two blinded raters on N = 64 essays (Week 1 and Week 8 combined). Average-measures intraclass correlation coefficients (ICC[2,2]) with 95% confidence intervals (CIs) for Overall and dimension scores.
| Measure | WCF Group | STEP-DWCF-R Group |
| Revision rounds per task | 1 | 2.26 |
| Feedback uptake rate (adopted or modified, %) | 68.5 | 82.3 |
| Persistent error rate after final round (%) | 67.4 | 45.6 |
| Teacher moderation time (min per task) | 23.5 | 13.8 |
| AI agent delivery time (min per task) | — | 3.9 |
| Learner revision time (min per task) | 44.8 | 71.2 |
| Total time on feedback + revision (min/task) | 68.3 | 88.9 |
Table 7: Means across 96 tasks (6 tasks × 16 learners). Uptake and persistent error rates were calculated at the feedback-point level. Time measures were recorded through teacher logs and system timestamps. AI agent delivery time is not applicable for the WCF group.
This study compared STEP-DWCF-R, a multi-component dynamic written corrective feedback framework with a robotic AI agent, with single-round WCF in an eight-week IELTS writing course. STEP-DWCF-R yielded larger pre–post gains in overall band and across TR, CC, LR, and GRA. Learners under STEP-DWCF-R also completed more revision rounds and showed faster error reduction, with models estimating a decline of about 4.1 errors per round. The greatest improvement was in coherence and cohesion (ΔΔ = 0.85), indicating that structured, rubric-linked feedback fosters higher-level organization as well as accuracy. These findings align with prior DWCF research showing that iterative, comprehensive feedback improves writing accuracy over time14,15,16,17.
The STEP-DWCF-R workflow involves three critical steps. First, the AI system generates itemized feedback across four categories (Grammar, Vocabulary, Organization, Reasoning) using a constrained JSON schema and a standardized system prompt. Second, a teacher moderates the output to remove false positives, add missed critical issues aligned with the IELTS rubric, ensure actionable and encouraging phrasing, and maintain consistency with the four-category schema. This moderation step serves as the primary troubleshooting mechanism, correcting AI hallucinations or over-marking that could otherwise overwhelm learners. Notably, teacher moderation time per task was lower in STEP-DWCF-R than in WCF (13.8 min versus 23.5 min), because the AI generated the initial structured feedback and the teacher only moderated it. Third, the moderated feedback is delivered via the web-based robotic AI agent interface, which logs learner acknowledgments and resubmissions across multiple rounds.
Compared with existing approaches, STEP-DWCF-R addresses distinct limitations. Conventional single-round WCF remains constrained by instructor time and typically provides limited opportunity for sustained revision. Automated writing evaluation tools offer scalability but often deliver static, fragmented comments that learners struggle to translate into meaningful revisions20,21,22. Early DWCF studies demonstrated the efficacy of multi-round feedback cycles, yet their reliance on instructor-authored corrections raised feasibility concerns in large classes14,15,16,17. Recent generative AI feedback studies report parity with teacher feedback on writing outcomes, but without structured moderation or iterative delivery18,19. STEP-DWCF-R combines AI scalability with teacher oversight and iterative robotic AI agent delivery, achieving lower teacher workload while increasing revision frequency.
We acknowledge several limitations. The relatively small sample size (N = 32) limits the generalizability of our findings, and the study should be viewed as a pilot investigation. The two conditions differed in feedback dosage: the WCF group received only one round of feedback per task, whereas the STEP-DWCF-R group underwent 2–3 iterative rounds. Therefore, the superior performance of the STEP-DWCF-R group reflects the combined effects of multiple revision cycles, AI-generated feedback, and teacher moderation, rather than the isolated effect of the robotic AI agent or its delivery modality. The robotic AI agent is a purely virtual web-based interface, so its effects cannot be disentangled from those of the iterative structure itself. Multimodal human feedback (e.g., speech plus gesture) was not included as a control condition because it is not standard practice in conventional EFL writing classrooms and would require a separate experimental paradigm. Future studies should include dosage-matched control groups (e.g., multi-round teacher feedback) to further disentangle these components.
Despite these limitations, the STEP-DWCF-R framework offers practical directions for AI-assisted writing instruction. Its modular architecture permits integration into learning management systems for large-scale EFL courses, and the structured four-category schema could be adapted for academic writing or discipline-specific genres. Future iterations might explore multimodal delivery to leverage the theoretical engagement benefits of embodied agents, though the current pilot establishes the feasibility of text-based iterative AI-teacher collaboration. Nevertheless, the consistent pattern across outcome measures, process indicators, and strong inter-rater reliability supports the feasibility and potential effectiveness of this multi-component approach in similar EFL contexts.
The authors have no competing interests.
This work was funded by a Universiti Sains Malaysia Bridging Grant, Project No: R501-LR-RND003-0000001342-0000.
| Name | Company | Catalog Number | Comments |
|---|---|---|---|
| Flask (Web Framework) | Pallets Projects | https://flask.palletsprojects.com | Version 2.1; used for robotic AI agent web interface |
| GPT-5 API | OpenAI | https://platform.openai.com | Model gpt-5, temperature=0.7; for structured JSON feedback generation |
| Jinja2 | Pallets Projects | https://jinja.palletsprojects.com | Server-rendered templates for web interface |
| Python | Python Software Foundation | https://www.python.org | Version 3.10; backend scripting |
| Windows 11 | Microsoft | https://www.microsoft.com/windows | Operating system for AI feedback system |
Request permission to reuse the text or figures of this JoVE article
Request Permission