Method Article

A Dynamic Written Corrective Feedback Framework Integrating AI Agent Delivery for Structured and Iterative Essay Support

DOI:

10.3791/71992

July 24th, 2026

In This Article

Summary

Loading...
$$\rightleftharpoonup{xx}$$ $$\longleftharp{xx}$$, $$\longrightharp{xx}$$,

This study benchmarks STEP-DWCF-R (Structured, Tiered, Evidence-driven Process for Dynamic Written Corrective Feedback with Robotic AI Agent) for improving IELTS writing performance through multi-round AI and teacher-supported revisions.

Abstract

Loading...
$$\rightleftharpoonup{xx}$$ $$\longleftharp{xx}$$, $$\longrightharp{xx}$$,

Automated writing feedback systems are prevalent, yet most deliver static, fragmented comments that provide limited scaffolding for revision. This study evaluates the STEP-DWCF-R framework (Structured, Tiered, Evidence-driven Process for Dynamic Written Corrective Feedback with Robotic AI Agent), in which AI-generated feedback, moderated by a teacher, is delivered via a robotic AI agent across multiple iterative rounds within a one-week task cycle. In an eight-week quasi-experimental trial, 32 EFL learners were randomized to either traditional written corrective feedback (one round per task) or STEP-DWCF-R. Both groups completed IELTS Task 2 essays at baseline and post-test, which were anonymized, randomized, and scored by two independent raters (ICC = 0.86–0.93). Linear mixed-effects models demonstrated that the STEP-DWCF-R group exhibited significantly greater gains in overall band score (Δ = 1.03 vs. 0.31 bands) and across all four analytic dimensions, with the largest improvement observed in Coherence and Cohesion. Process data indicated that STEP-DWCF-R learners completed an average of 2.26 revision rounds per task, with error counts decreasing linearly across rounds. These findings suggest that the integrated STEP-DWCF-R framework, encompassing AI-generated feedback, teacher moderation, and iterative robotic AI agent delivery, is associated with greater IELTS writing improvement than traditional single-round feedback, pointing to practical applications for AI-enhanced dynamic feedback in EFL contexts.

Introduction

Loading...
$$\rightleftharpoonup{xx}$$ $$\longleftharp{xx}$$, $$\longrightharp{xx}$$,

Written corrective feedback (WCF) remains central to L2 writing pedagogy. Evidence shows that comprehensive WCF improves learners’ accuracy over time and, when aligned with classroom practice, can coexist with focused approaches that are feasible in authentic contexts1,2,3. Recent classroom studies show durable gains in accuracy and fluency from sustained, comprehensive WCF. Research on feedback scope cautions that teachers should target forms strategically rather than mark everything4,5,6,7,8,9. In parallel, automated writing evaluation (AWE) and automated written corrective feedback (AWCF) have grown rapidly. Findings are mixed: some studies show gains in task achievement and grammatical range with AWCF or Criterion-style programs, while others find no significant advantage over teacher-only feedback or note largely local, surface-level revisions. To reflect this heterogeneity, we deliberately triangulate across multiple AWCF studies rather than leaning on a single source10,11,12,13.

Learners’ beliefs about who is providing feedback also matter: quasi-experimental work on perceived source indicates that performance and trust can shift when identical feedback is framed as “teacher” versus “automated,” underscoring the need to account for perception effects in AWCF designs14,15. Extending this line, emerging studies suggest that educational robots, due to their embodied presence, can also act as feedback providers, potentially influencing learner engagement and trust in distinctive ways16.

A complementary line of research, dynamic written corrective feedback (DWCF), emphasizes frequent, manageable, comprehensive feedback cycles with coding and rapid revision. Evidence from multiple settings suggests DWCF reliably improves accuracy (with frequency effects on fluency), though results may vary by course goals and learner population17,18,19,20. Finally, early studies on generative-AI-mediated feedback report parity with teacher feedback on writing outcomes and potential benefits when paired with explicit metalinguistic guidance, motivating closer integration of human and AI feedback in L2 contexts21,22.

To address these challenges, we propose the STEP-DWCF-R (Structured, Tiered, Evidence-driven Process for Dynamic Written Corrective Feedback with Robotic AI Agent) framework, a novel multi-stage DWCF feedback system designed to provide structured, iterative, and process-oriented writing support. Our system leverages the predictive and generative power of Large Language Models (LLMs), with delivery through a robotic AI agent interface and teacher moderation, to segment complex writing issues into manageable categories, deliver feedback through multiple rounds of interaction, and evaluate its effectiveness using both outcome-based and process-based metrics. By transforming feedback into a structured and interactive process, STEP-DWCF-R advances automated writing support from static error correction toward a more engaging, embodied, and learning-oriented system.

Our contributions center on three integrated advances. First, we propose a structured feedback representation schema that categorizes errors (e.g., grammar, coherence) so that robotic AI agent-delivered LLM feedback is both actionable and pedagogically interpretable. Second, we introduce an iterative multi-stage process that sequences feedback across language, structure, and reasoning across multiple rounds to reduce cognitive load and deepen engagement. Third, we extend evaluation beyond post-scores to include process-oriented metrics such as error reduction and feedback adoption, offering a richer view of learning impact. WCF frameworks represent the earliest attempts to systematize written corrective feedback within educational technology. Originating in Automated Writing Evaluation (AWE) and Automated Essay Scoring (AES), these systems emphasized error detection and proficiency scoring13,14. Their contribution lies in scalability: thousands of learners could receive feedback without the bottleneck of human grading. However, these frameworks typically provided static and fragmented comments, such as grammar checks, lexical suggestions, or global scores that learners struggled to transform into meaningful revisions23,24,25,26.

Over time, WCF research evolved to include more technology-mediated approaches. These newer systems sought to capture broader aspects of writing by leveraging computational methods13,21. More recently, the scope has expanded further with embodied technologies, where educational robots have begun to act as feedback providers in writing or tutoring contexts. Their physical presence and interactive affordances introduce new dynamics of trust, engagement, and learner motivation. Yet, despite these developments, the dominant orientation of WCF has remained outcome-focused8,27: producing scores or isolated comments in a one-shot manner. Pedagogically, this orientation is misaligned with formative assessment, where feedback should scaffold revision across drafts and support metacognitive engagement16,28,29,30,31. Thus, the conceptual boundary of WCF reveals an enduring gap: feedback is delivered, but not structured or sequenced to guide learning processes.

In response to these limitations, scholars have begun to explore more dynamic approaches to writing feedback. Early investigations highlighted the importance of iterative feedback cycles, where learners revise multiple times under scaffolded guidance17,32. Structured error categorization has also been proposed to improve the interpretability of feedback and align it with pedagogical objectives33,34. The notion of a DWCF framework consolidates these directions by embedding three design principles: explicit error schemas, staged and iterative delivery, and process-oriented evaluation metrics19,35. Rather than overwhelming students with a bulk of corrections at once, DWCF distributes attention across layers of writing—surface accuracy, structural coherence, and argumentative depth—through successive rounds17,33. Evaluation expands beyond final scores to include process indicators such as error reduction rates, feedback adoption, and revision depth10,19,25. Despite its promise, practical applications of DWCF remain sparse, and most studies stop short of operationalizing it in large-scale, technology-mediated contexts10,36,37. This gap highlights the need for frameworks that not only theorize DWCF but also demonstrate its feasibility in real-world educational settings. Figure 1 shows the broader theoretical framework of DWCF proposed in the literature. The figure provides a general rationale for how dynamic written corrective feedback may operate at multiple levels, including both measured outcomes (e.g., accuracy, coherence) and theorized but unmeasured constructs in the present study (e.g., writing anxiety reduction, fluency, and complexity). It is included for theoretical orientation rather than as a direct model of this trial. Elements corresponding to the outcomes and process metrics analyzed here are indicated in the figure; other pathways are illustrative and were not evaluated in this study.

The emergence of LLMs and multimodal systems offers a powerful means to operationalize DWCF at scale. LLMs, with their zero-shot and few-shot capabilities, can generate context-sensitive feedback without task-specific training21,22. Recent experiments suggest their potential for directive segmentation, staged refinement, and explanation generation, which align closely with the principles of DWCF13,37,38,39. Complementary research in human–AI collaboration has shown that iterative loops of engaging with, adapting, and selectively adopting AI feedback can improve both uptake and revision quality25,30. When these processes are embodied through robots, the interaction has the theoretical potential to become multimodal—combining gesture, voice, and presence—which may further enhance learner perception and motivation40.

Grounded in the Zone of Proximal Development (ZPD)41, the Input Hypothesis42, and the Skill Acquisition Theory43, we position STEP-DWCF-R as a controller that links learner support, input processing, and skill development. Concretely, rubric-linked itemization, timely consolidated lists, and multi-round teacher-moderated AI–, human–, and robot–cycles operate at an integration layer that mediates revision and yields the observable endpoints analyzed here (IELTS Overall and TR/CC/LR/GRA; rounds, uptake, persistent errors, time). Constructs such as writing anxiety reduction are theorized but not measured in this trial.

Protocol

Loading...
$$\rightleftharpoonup{xx}$$ $$\longleftharp{xx}$$, $$\longrightharp{xx}$$,

This protocol was approved by the Ethics Committee of the School of Primary Education, Shangrao Preschool Education College. All participants were adults (≥18 years) and provided written informed consent prior to the study.

1. Study design

NOTE: Due to the constraints of the available EFL classroom (total enrollment N = 32), participants were randomized into two groups of 16 each. This sample size, while modest, was determined to be sufficient for a pilot investigation given the within-subject pre-post design and expected large effect sizes based on prior DWCF studies17,18,19,20.

  1. Randomize participants into two groups (n = 16 each) using simple randomization. Generate the allocation sequence using a computer-generated random number list. Conceal allocation from the instructor until baseline assessment is completed.
  2. Ensure both groups are taught by the same instructor using identical materials and contact hours.
  3. Deliver feedback either through direct human correction (WCF) or through the STEP-DWCF-R workflow, where AI and teacher feedback are presented via a robotic AI agent.

2. Writing Tasks

  1. Administer a baseline IELTS Task 2 essay in Week 1 under exam conditions (≥250 words, 40 min).
  2. Conduct six weekly writing tasks (Weeks 2–7). For each task:
    1. In WCF, provide one round of human feedback on the essay.
    2. In STEP-DWCF-R, generate AI feedback, moderate it with a human teacher, and instruct the robotic AI agent to present the feedback interactively to the learner.
    3. In STEP-DWCF-R, repeat the STEP-DWCF-R feedback cycle for 2–3 rounds until the revision is complete.
  3. Administer a post-test IELTS Task 2 essay in Week 8 under the same exam conditions. Follow the same one-week calendar for each task in both groups. Deliver Round 1 feedback within 24 h of draft submission in STEP-DWCF-R. Require learners to revise and resubmit within 48 h. Follow the same schedule for subsequent rounds and complete all cycles within the weekly window.

3. Feedback Workflow

  1. Process each learner draft with the GPT-5 API (model = "gpt-5", temperature = 0.7, max_output_tokens = 2048) to generate itemized feedback across four fixed categories: Grammar, Vocabulary, Organization, and Reasoning. The exact system prompt and JSON output schema are provided in the open-source repository (https://github.com/YifangGaoinPG/Dynamic-Writing-Correction-Feedback/blob/main/evaluate.py, function build_prompt() and normalize_reasoning()).
  2. Moderate the AI-generated feedback according to the following criteria: remove false positives or inaccurate comments; add missed critical issues aligned with the IELTS rubric; ensure comments are actionable and encouraging; and maintain consistency with the four-category schema. Log moderation decisions manually.
  3. Save the moderated feedback in JSON format and upload it to the robotic AI agent interface (a lightweight Flask web application).
  4. Present feedback through the web interface. Display structured tables for each category (summary, issues, and revision tips). Record learner acknowledgments and clarification requests through system logs. Instruct learners to revise and resubmit their drafts. Repeat the cycle up to three times per task.
  5. Verification and troubleshooting.
    1. Verify successful AI feedback generation by confirming that the output is valid JSON containing all four required categories: Grammar, Vocabulary, Organization, and Reasoning. Confirm completion of teacher moderation by ensuring that false positives have been removed, missing issues have been added, and the moderated JSON file has been saved.
    2. Upload the moderated JSON file to the Flask-based robotic AI agent interface through the upload endpoint. Confirm successful upload through the interface confirmation message and database log entry. Record learner resubmissions automatically through form submission logs.
    3. Process non-compliant AI outputs (e.g., malformed JSON, missing categories, or inaccurate feedback) using the ensure_json() and normalize_reasoning() functions. Regenerate the API output with adjusted prompt parameters as needed. Manually correct the JSON during teacher moderation, if necessary, to ensure that all four categories are present before upload.
      NOTE: Examples of raw AI feedback, teacher-moderated versions, and the delivered interface output are available in the repository (data/ folder and notebooks/).

4. Outcome Measurement

  1. Define the primary outcome as the change in IELTS overall band score (Week 1 to Week 8).
  2. Define secondary outcomes as changes in Task Response, Coherence and Cohesion, Lexical Resource, and Grammatical Range and Accuracy.
  3. Assign two independent raters with experience in IELTS Task 2 scoring to evaluate all essays. Remove names and group identifiers from all essays. Randomize essay order and blind raters to group assignment and time point. Use the same IELTS Task 2 prompt for both the Week 1 baseline and Week 8 post-test. Resolve scoring disagreements greater than 0.5 bands through discussion until consensus is reached. Assess inter-rater reliability using the average-measures intraclass correlation coefficient (ICC[2,2]).

5. Process Measurement

  1. Log the number of revision rounds per task.
  2. Record uptake of feedback (adopted, modified, declined) by learners.
  3. Track persistent errors between cycles.
  4. Record teacher feedback time, robot delivery time, and learner revision time.

6. Data Analysis

  1. Fit linear mixed-effects models with Group, Time, and Group × Time interaction.
  2. Apply generalized mixed-effects models for process measures.
  3. Report effect sizes, 95% confidence intervals, and adjusted p-values.
  4. Conduct robustness checks with ANCOVA, rater-specific models, and robust SEs.

7. Code Availability

All code, prompts, JSON schemas, and example implementation files required to reproduce the STEP-DWCF-R system are openly available at https://github.com/YifangGaoinPG/Dynamic-Writing-Correction-Feedback.

Results

Loading...
$$\rightleftharpoonup{xx}$$ $$\longleftharp{xx}$$, $$\longrightharp{xx}$$,

Experiments and Analysis

All 32 learners provided baseline data. Post-test data were available for 16 learners in each group (WCF and STEP-DWCF-R; Figure 2). The study timeline and measures are presented in Figure 3, and the end-to-end feedback workflow is illustrated in Figure 4. Experimental setup and implementation parameters for our AI system are shown in Table 1. Analyses followed the pre-specified plan with intention-to-treat. We first assess baseline equivalence (Table 2), then report the primary outcome (Figure 5 and Table 3), followed by secondary outcomes (Table 4) with robustness and reliability checks.

Baseline Equivalence

At baseline (Week 1), the groups were descriptively similar on pre-specified covariates, including demographics and IELTS writing performance before any feedback (Table 2). Individual IELTS component scores are assigned in 0.5-band steps, whereas the table reports group means. The Week 1 overall means (WCF = 5.625, STEP-DWCF-R = 5.500) are computed from the same arrays used to generate Figure 5. We do not perform hypothesis tests at baseline; any residual imbalance is addressed by the pre-post design and the Group × Time term in the mixed-effects model, and a post-test comparison adjusted for baseline is reported in the Protocol section.

Primary Outcome

Figure 5 summarizes pre–post changes, while Table 3 and Table 5 report model estimates and post-test attainment. At baseline (Week 1), the group means were 5.625 for WCF and 5.500 for STEP-DWCF-R, yielding a non-significant group difference of −0.125 bands (SE = 0.133, t = − .939, df = 29.90, p = .355, 95% CI [−0.397, 0.147]). The baseline in Table 3, therefore, estimates the WCF Week 1 mean at 5.625 with 95% CI [5.419, 5.831] (SE = 0.097, df = 15). From Week 1 to Week 8, WCF improved by 0.3125 bands (SE = 0.128, t = 2.44, df = 15, p = .028, 95% CI [0.039, 0.586]; standardized within-group effect dz = 0.61). STEP-DWCF-R improved by 1.0313 bands (SE = 0.140, t = 7.34, df = 15, p = 2.44 × 10−6, 95% CI [0.732, 1.331]; dz = 1.84). The linear mixed-effects contrast for the Group × Time interaction—i.e., the difference-in-differences—was 0.7188 bands (SE = 0.190, t = 3.78, df = 29.75, p = .0007, 95% CI [0.330, 1.107]; Table 3), indicating larger gains under STEP-DWCF-R. At post-test (Week 8), the estimated marginal means favored STEP-DWCF-R by 0.5938 bands (Welch test on group means: SE = 0.115, t = 5.16, df = 29.74, p = 1.50 × 10−5, 95% CI [0.359, 0.829]). Consistent with this, the attainment table (Table 5) shows that at Week 8, 13% of WCF learners reached Overall ≥ 6.5 and 0% reached ≥ 7.0 (counts 2/16 and 0/16), whereas 81% of STEP-DWCF-R learners reached ≥ 6.5 and 25% reached ≥ 7.0 (counts 13/16 and 4/16). Table 3 reports the corresponding fixed-effect estimates and their uncertainty.

Dimension-Level Outcomes

Table 4 summarizes pre–post changes across the four Task 2 dimensions. Task Response (TR) showed a gain of Δ = 0.25 bands under WCF versus Δ = 0.95 under STEP-DWCF-R, yielding ΔΔ = 0.70 with 95% CI [0.28, 1.12] and Holm-adjusted p = .006. Coherence and Cohesion (CC) exhibited the largest contrast: WCF Δ = 0.30, STEP-DWCF-R Δ = 1.15, hence ΔΔ = 0.85 (95% CI [0.44, 1.26], adj-p = .002). Lexical Resource (LR) followed the same pattern with WCF Δ = 0.35 and STEP-DWCF-R Δ = 0.95, giving ΔΔ = 0.60 (95% CI [0.18, 1.02], adj-p = .012). Grammatical Range and Accuracy (GRA) improved by Δ = 0.25 in WCF and Δ = 0.97 in STEP-DWCF-R; the resulting ΔΔ = 0.72 carried a 95% CI of [0.31, 1.13] with adj-p = .004. Across all four dimensions, the Group×Time contrasts favored STEP-DWCF-R after Holm–Bonferroni adjustment.

Process Evidence

Figure 6 and Figure 7 visualize the core process results. Across Weeks 2–7, STEP-DWCF-R completed on average 2.26 revision rounds per task (95% CI [2.17, 2.35]; n = 96), whereas WCF was 1.00 rounds for all tasks (n = 96). The mean difference was 1.26 rounds (95% CI [1.17, 1.35]); a Mann–Whitney test confirmed separation of the distributions (U = 9216, z = 11.97, p < 10−30). Within STEP-DWCF-R, mean error counts declined by round: 13.78 at Round 1 (95% CI [13.02, 14.54]; n = 96), 8.94 at Round 2 (95% CI [8.42, 9.46]; n = 96), and 6.16 at Round 3 (95% CI [5.20, 7.12]; n = 25). A linear trend fitted to per-round observations estimated a decline of 4.14 errors per round (95% CI [3.51, 4.78] in absolute magnitude; p < 10−12). Measures of feedback uptake and time-on-task are not encoded in Figure 6 and Figure 7 and are therefore reported in Table 7.

Reliability and Robustness

Inter-rater agreement for Week 1 and Week 8 essays was good to excellent. Two trained raters scored all scripts independently and were blinded to group and time; using average-measures intraclass correlation coefficient (ICC[2,2]) on the rater means, reliability ranged from 0.86–0.93 across Overall, and the four dimensions, and all lower 95% confidence limits exceeded 0.80 (Table 6). Diagnostic checks for the primary mixed-effects model indicated approximately normal residuals without material heteroscedasticity, and models fitted to task-level process summaries showed dispersion close to one. Sensitivity analyses based on the same Week 1/Week 8 dataset yielded consistent inferences: a linear regression comparing groups at post-test while adjusting for baseline scores estimated an adjusted between-group difference of 0.575 bands at Week 8 (95% CI [0.336, 0.814], p = 3.15 × 10−5), and a rater-specific mixed model that included a random effect for Rater and used un-averaged scores produced a Group × Time estimate of 0.72 bands (95% CI [0.33, 1.11], p = .0007), aligning with Table 3. Results were unchanged when using robust standard errors and when analyses were restricted to complete cases, reinforcing the conclusion that STEP-DWCF-R yields larger pre-post gains than WCF.

Qualitative Findings

For qualitative analysis, we examined STEP-DWCF-R revision traces across all six tasks and end-of-course written reflections from the 16 participants in the STEP-DWCF-R group. Two coders independently coded the materials using a focused deductive framework based on the four feedback categories (grammar, vocabulary, organization, reasoning) and uptake rationales. Inter-coder agreement was substantial with Cohen’s kappa of 0.78, and disagreements were resolved through discussion.

Analysis revealed five key themes: uptake followed a staged pattern, with learners resolving surface features before addressing organization and reasoning; span-specific, rubric-linked items were more actionable and frequently adopted than narrative suggestions; AI-teacher complementarity shaped priority, with overlapping signals increasing focus and teacher moderation resolving conflicts; benefits peaked early, with reuse of organization templates and lexical patterns noted on new prompts; and learners adapted strategies across tasks. These themes inform the process dynamics explored in the process evidence section. Feedback uptake, persistent errors, and time-on-task measures were also recorded. A summary of these additional process indicators is presented in Table 7.

Interpretation of Representative Results

Taken together, the outcome and process data indicate that the STEP-DWCF-R framework is associated with greater IELTS writing improvement than traditional single-round feedback. The primary outcome shows a between-group difference-in-differences of 0.72 bands, with the STEP-DWCF-R group improving by 1.03 bands overall compared with 0.31 bands in the WCF group. This advantage was consistent across all four analytic dimensions, with the largest contrast in Coherence and Cohesion at 0.85 bands, suggesting that structured, rubric-linked, multi-stage feedback particularly scaffolded organizational development. Process evidence further indicates that iterative engagement underlies this advantage. STEP-DWCF-R learners completed an average of 2.26 revision rounds per task, and error counts declined linearly by approximately 4.1 errors per round. For instance, a representative workflow cycle included GPT-5 generating structured JSON feedback across the four categories, followed by teacher moderation to eliminate false positives and enhance actionability, and subsequent delivery via the robotic AI agent interface for multi-round learner revision. These findings support the feasibility and potential effectiveness of the integrated multi-component workflow combining AI-generated feedback, teacher moderation, and iterative robotic AI agent delivery, but they should not be interpreted as proof of any single component in isolation. The dose imbalance of 2–3 rounds versus one round and the modest sample size of 32 mean that the observed gains reflect the combined effect of increased revision opportunities and structured feedback. These results should be viewed as promising pilot evidence awaiting confirmation in larger, component-controlled trials.

Diagram of dynamic written corrective feedback (DWCF) mechanism linked to writing development outcomes.
Figure 1. Theoretical framework of DWCF (dynamic written corrective feedback). The figure provides a broader rationale for how DWCF may operate; it is included for orientation rather than as a direct model of this trial. Elements corresponding to the outcomes and process metrics analyzed here are indicated in the figure; other pathways are illustrative and were not evaluated. Please click here to view a larger version of this figure.

Flowchart of randomized clinical trial; participant allocation and analysis; control vs. treatment groups.
Figure 2. CONSORT flow diagram of participants. Thirty-two learners were assessed for eligibility; none were excluded. All participants provided consent and were then randomized in a 1:1 ratio to either the WCF group (n = 16) or the DWCF group (n = 16). All randomized participants received the allocated intervention; there were no losses to follow-up or discontinuations, and all 32 were included in the final analysis. Please click here to view a larger version of this figure.

Timeline diagram of IELTS task intervention period with baseline and post-test scoring.
Figure 3. Study timeline and measures. The trial spanned 8 weeks: at W1, participants completed the baseline IELTS Task 2 and were scored; six weekly writing tasks were administered from W2 to W7 (Task 1–Task 6), with group-specific feedback (WCF or DWCF) and revisions after each task. Process measures, including revision rounds, feedback uptake, persistent-error rate, and time on task, were logged for each task during W2–W7. At W8, the post-test IELTS Task 2 was administered and scored. Please click here to view a larger version of this figure.

Flowchart of feedback process; AI moderation, teacher review, cycle revision.
Figure 4. End-to-end feedback workflow. Learners produce a proctored initial draft, then receive group-specific feedback: WCF (one human feedback round) or STEP-DWCF-R (AI-generated feedback with teacher moderation, delivered via robotic AI agents, involving 2–3 iterative rounds). Learners complete revision round(s) accordingly. Outcome is the IELTS Task 2 band score; process measures include number of rounds, feedback uptake (adopted/modified/declined), persistent-error rate, and time on task. The system logs item-level IDs, category/severity, source (AI vs. human vs. robot), moderation actions, and uptake status. Please click here to view a larger version of this figure.

IELTS writing band scores over 8 weeks; scatter plot comparing WCF and STEP-DWCF-R methods.
Figure 5. Change in IELTS overall band from Week 1 to Week 8 by group. Points give the estimated group means with 95% confidence intervals, and the trajectories indicate a larger gain under STEP-DWCF-R Please click here to view a larger version of this figure.

Graph comparing WCF and STEP-DWCF-R tasks; revision rounds per task analyzed statistically.
Figure 6. Revision rounds per task by group (Weeks 2–7). Data points represent individual task revision rounds for the WCF (blue) and STEP-DWCF-R (orange) groups. Horizontal lines and error bars indicate group means with 95% confidence intervals. Please click here to view a larger version of this figure.

Mean error count graph; decreasing trend over rounds; statistical data analysis; line graph.
Figure 7. Mean error counts per round within STEP-DWCF-R with 95% CIs. Points indicate mean error counts at each feedback round (Round 1, Round 2, Round 3), and vertical lines represent 95% confidence intervals (CIs). Please click here to view a larger version of this figure.

ComponentSpecification
HardwarePC with 12th Gen Intel(R) Core(TM) i7-12700H (2.30 GHz), 32GB RAM, NVIDIA RTX 3080Ti GPU (16GB)
Operating SystemWindows 11
BackendFlask~2.1 (Python~3.10)
FrontendJinja2 server-rendered templates
AI ModelsOpenAI GPT-5 API (for feedback generation), schema-based post-processing
Feedback SchedulerCustom controller managing multi-stage feedback (3 cycles)
Error SchemaGrammar, Vocabulary, Organization, Reasoning
Version ControlGitHub
LoggingAutomatic logging of submissions, feedback, and revisions for analysis

Table 1: Experimental setup and implementation parameters. Technical specifications of the AI feedback system, including hardware configuration, operating system, backend and frontend frameworks, AI model settings, feedback scheduler, error schema categories, version control, and logging infrastructure.

VariableWCF (n=16)STEP-DWCF-R (n=16)
Age (years), mean (SD)20.9 (1.5)21.1 (1.6)
Female, n (%)9 (56%)8 (50%)
Prior IELTS preparation (yes), n (%)7 (44%)6 (38%)
Years of prior writing instruction3.1 (1.2)3.2 (1.1)
Baseline IELTS Overall (band)5.625 (0.387)5.500 (0.365)
Baseline TR (band)5.56 (0.50)5.50 (0.50)
Baseline CC (band)5.62 (0.49)5.53 (0.49)
Baseline LR (band)5.60 (0.46)5.52 (0.46)
Baseline GRA (band)5.56 (0.48)5.49 (0.45)

Table 2: Baseline characteristics of participants by group (Week 1). Demographic variables and pre-test International English Language Testing System (IELTS) Task 2 writing performance for the WCF and STEP-DWCF-R groups. Values are mean (standard deviation) or n (%).

Fixed effectEstimateSEdftp95% CI
Intercept (WCF, Week~1)5.6250.0971558.1<.001[5.419, 5.831]
Group (STEP-DWCF-R vs. WCF)-0.1250.13329.9-0.939.355[-0.397, 0.147]
Time (Week~8 vs. Week~1)0.31250.128152.44.028[0.039, 0.586]
Group × Time0.71880.1929.753.78.0007[0.330, 1.107]

Table 3: Linear mixed-effects model for IELTS overall band from Week 1/Week 8 scores. Fixed-effect estimates, standard errors (SE), degrees of freedom (df), t-values, p-values, and 95% confidence intervals (CIs) for the Group × Time interaction and model terms.

DimensionWCF Δ (bands)STEP-DWCF-R Δ (bands)ΔΔ (95% CI)adj-p
Task Response (TR)0.250.950.70 (0.28, 1.12).006
Coherence and Cohesion (CC)0.31.150.85 (0.44, 1.26).002
Lexical Resource (LR)0.350.950.60 (0.18, 1.02).012
Grammatical Range and Accuracy (GRA)0.250.970.72 (0.31, 1.13).004

Table 4: Dimension-level outcomes (Task 2): pre–post changes and between-group differences. Pre–post changes (Δ) and difference-in-differences (ΔΔ) with 95% confidence intervals (CIs) and Holm-adjusted p-values for Task Response (TR), Coherence and Cohesion (CC), Lexical Resource (LR), and Grammatical Range and Accuracy (GRA).

ThresholdWCF (%)STEP-DWCF-R (%)
Overall ≥ 6.51381
Overall ≥ 7.0025

Table 5: Post-test band attainment (Week 8). Proportion of learners in the WCF and STEP-DWCF-R groups reaching IELTS Overall band score thresholds of at least 6.5 and at least 7.0.

OutcomeICC95% CI
Overall0.91[0.86, 0.94]
Task Response (TR)0.88[0.82, 0.92]
Coherence and Cohesion (CC)0.9[0.85, 0.94]
Lexical Resource (LR)0.89[0.83, 0.93]
Grammatical Range and Accuracy (GRA)0.87[0.80, 0.91]

Table 6: Inter-rater reliability for IELTS outcomes. Two blinded raters on N = 64 essays (Week 1 and Week 8 combined). Average-measures intraclass correlation coefficients (ICC[2,2]) with 95% confidence intervals (CIs) for Overall and dimension scores.

MeasureWCF GroupSTEP-DWCF-R Group
Revision rounds per task12.26
Feedback uptake rate (adopted or modified, %)68.582.3
Persistent error rate after final round (%)67.445.6
Teacher moderation time (min per task)23.513.8
AI agent delivery time (min per task)3.9
Learner revision time (min per task)44.871.2
Total time on feedback + revision (min/task)68.388.9

Table 7: Means across 96 tasks (6 tasks × 16 learners). Uptake and persistent error rates were calculated at the feedback-point level. Time measures were recorded through teacher logs and system timestamps. AI agent delivery time is not applicable for the WCF group.

Discussion

Loading...
$$\rightleftharpoonup{xx}$$ $$\longleftharp{xx}$$, $$\longrightharp{xx}$$,

This study compared STEP-DWCF-R, a multi-component dynamic written corrective feedback framework with a robotic AI agent, with single-round WCF in an eight-week IELTS writing course. STEP-DWCF-R yielded larger pre–post gains in overall band and across TR, CC, LR, and GRA. Learners under STEP-DWCF-R also completed more revision rounds and showed faster error reduction, with models estimating a decline of about 4.1 errors per round. The greatest improvement was in coherence and cohesion (ΔΔ = 0.85), indicating that structured, rubric-linked feedback fosters higher-level organization as well as accuracy. These findings align with prior DWCF research showing that iterative, comprehensive feedback improves writing accuracy over time14,15,16,17.

The STEP-DWCF-R workflow involves three critical steps. First, the AI system generates itemized feedback across four categories (Grammar, Vocabulary, Organization, Reasoning) using a constrained JSON schema and a standardized system prompt. Second, a teacher moderates the output to remove false positives, add missed critical issues aligned with the IELTS rubric, ensure actionable and encouraging phrasing, and maintain consistency with the four-category schema. This moderation step serves as the primary troubleshooting mechanism, correcting AI hallucinations or over-marking that could otherwise overwhelm learners. Notably, teacher moderation time per task was lower in STEP-DWCF-R than in WCF (13.8 min versus 23.5 min), because the AI generated the initial structured feedback and the teacher only moderated it. Third, the moderated feedback is delivered via the web-based robotic AI agent interface, which logs learner acknowledgments and resubmissions across multiple rounds.

Compared with existing approaches, STEP-DWCF-R addresses distinct limitations. Conventional single-round WCF remains constrained by instructor time and typically provides limited opportunity for sustained revision. Automated writing evaluation tools offer scalability but often deliver static, fragmented comments that learners struggle to translate into meaningful revisions20,21,22. Early DWCF studies demonstrated the efficacy of multi-round feedback cycles, yet their reliance on instructor-authored corrections raised feasibility concerns in large classes14,15,16,17. Recent generative AI feedback studies report parity with teacher feedback on writing outcomes, but without structured moderation or iterative delivery18,19. STEP-DWCF-R combines AI scalability with teacher oversight and iterative robotic AI agent delivery, achieving lower teacher workload while increasing revision frequency.

We acknowledge several limitations. The relatively small sample size (N = 32) limits the generalizability of our findings, and the study should be viewed as a pilot investigation. The two conditions differed in feedback dosage: the WCF group received only one round of feedback per task, whereas the STEP-DWCF-R group underwent 2–3 iterative rounds. Therefore, the superior performance of the STEP-DWCF-R group reflects the combined effects of multiple revision cycles, AI-generated feedback, and teacher moderation, rather than the isolated effect of the robotic AI agent or its delivery modality. The robotic AI agent is a purely virtual web-based interface, so its effects cannot be disentangled from those of the iterative structure itself. Multimodal human feedback (e.g., speech plus gesture) was not included as a control condition because it is not standard practice in conventional EFL writing classrooms and would require a separate experimental paradigm. Future studies should include dosage-matched control groups (e.g., multi-round teacher feedback) to further disentangle these components.

Despite these limitations, the STEP-DWCF-R framework offers practical directions for AI-assisted writing instruction. Its modular architecture permits integration into learning management systems for large-scale EFL courses, and the structured four-category schema could be adapted for academic writing or discipline-specific genres. Future iterations might explore multimodal delivery to leverage the theoretical engagement benefits of embodied agents, though the current pilot establishes the feasibility of text-based iterative AI-teacher collaboration. Nevertheless, the consistent pattern across outcome measures, process indicators, and strong inter-rater reliability supports the feasibility and potential effectiveness of this multi-component approach in similar EFL contexts.

Disclosures

Loading...
$$\rightleftharpoonup{xx}$$ $$\longleftharp{xx}$$, $$\longrightharp{xx}$$,

The authors have no competing interests.

Acknowledgements

Loading...
$$\rightleftharpoonup{xx}$$ $$\longleftharp{xx}$$, $$\longrightharp{xx}$$,

This work was funded by a Universiti Sains Malaysia Bridging Grant, Project No: R501-LR-RND003-0000001342-0000.

Materials

List of materials used in this article
NameCompanyCatalog NumberComments
Flask (Web Framework)Pallets Projectshttps://flask.palletsprojects.comVersion 2.1; used for robotic AI agent web interface
GPT-5 APIOpenAIhttps://platform.openai.comModel gpt-5, temperature=0.7; for structured JSON feedback generation
Jinja2Pallets Projectshttps://jinja.palletsprojects.comServer-rendered templates for web interface
PythonPython Software Foundationhttps://www.python.orgVersion 3.10; backend scripting
Windows 11Microsofthttps://www.microsoft.com/windowsOperating system for AI feedback system

Reprints and Permissions

Request permission to reuse the text or figures of this JoVE article

Request Permission

Tags

EngineeringEnglish writingwritten corrective feedback WCFdynamic written corrective feedback DWCFhuman AI collaborationfeedback uptakemixed effects models

Related Articles