Teaching experiment design
This study used the basic open-source version of DeepSeek, relying on the platform's open API interface for program calls. The complete set of five prompt template texts, four levels of task example entries, transfer annotation category tags, and writing error-template matching mapping rules were compiled in the supplementary appendix. The system had fixed calling parameters: a temperature coefficient of 0.7, and the generated length was dynamically limited based on the writing question type. It strictly followed a fixed sequence of semantic reconstruction > stylistic adjustment > cultural adaptation, executing nested prompt chain calls. The model output was manually selected by the instructor for effective text in classroom teaching.
This study invited three full-time university-level English writing teachers with more than 5 years of experience to participate in manual scoring. All scorers received standardized training before the official scoring, familiarizing themselves with the assessment dimensions, the 5-point scoring system, cultural transfer, syntax, discourse, and other scoring details, and completed pre-scoring calibration. All student essays were scored independently by two teachers in a blind, double-blind scoring mode. If the difference between the two scorers' scores exceeded 1 point (out of 5), a third senior teacher would arbitrate the scoring, and the average of the two valid scores would be taken as the final manual score for that sample. Two groups of raters were involved in the evaluation phase for different research purposes. Three full-time university English writing teachers with over 5 years of teaching experience conducted blind, double-scoring of all student writing samples in the formal experiment, and a third senior teacher served as an arbitrator when scoring discrepancies exceeded 1 point on the 5-point scale. In contrast, five raters participated in the inter-rater reliability test using the Intraclass Correlation Coefficient (ICC), which was conducted separately to verify the consistency of evaluation standards across assessors. The difference in the number of raters stems from distinct functional requirements: three core teachers ensured the practical scoring of experimental data, while the expanded five-rater panel provided more robust evidence for the reliability of the whole evaluation system.
Syntactic complexity, BERT text structure scoring, and semantic similarity are all scored within a 0–5 point range. The weights for each dimension were set according to foreign language writing teaching and research standards: syntax 0.35, text structure 0.35, and logic and culture 0.3. The weight values were based on mature quantitative evaluation systems for writing in the same field. The AI automatic scoring module was calibrated using thresholds based on 15 pre-classified model essays. The human scoring followed a unified five-point grading scale. The machine scoring and human calibration were combined for a pre-experiment verification before the formal evaluation.
The teaching task was carried out around the three-level transfer goals of "syntax-structure-culture", with three rounds of training, each lasting one week. The first round involved sentence-level rewriting to improve syntactic diversity and expression accuracy; the second round entailed paragraph-level structural reconstruction to enhance paragraph organization and logical coherence; the third round involved cultural-level transfer to strengthen cross-cultural pragmatic awareness and expression adaptability. Given repeated measurements collected from identical participants across pre-test, three sequential training rounds, and post-test, repeated-measures analysis of variance (RM-ANOVA) was adopted as the primary statistical approach to examine main effects of group (experimental vs. control), main effects of testing time, as well as the group × time interaction effect, which reflected whether score changes over training differed between the two cohorts. The sphericity assumption was verified via Mauchly’s test; Greenhouse–Geisser correction was applied if sphericity was violated. Independent-samples t-tests were reserved only for post-hoc pairwise between-group comparisons at individual time points, while Pearson correlation and Cohen’s d remained for scoring consistency and effect size quantification, respectively.
All statistical analyses were pre-specified before data collection. Independent-samples t-tests were adopted to compare pre-post writing indicators between two groups, with the primary null hypothesis assuming no inter-group differences in writing transfer performance. Pearson correlation analysis was used to quantify the linear association between AI automatic scoring and teacher manual ratings, and Cohen’s d was calculated as the standardized effect size for between-group comparisons. The significance threshold was set at α = 0.05 for all hypothesis testing, and 95% confidence intervals were computed alongside all test statistics. SPSS 26.0 was utilized to conduct all statistical calculations. No missing data occurred across student test scores and questionnaire datasets, as all participants finished the full three-round training and relevant assessments; hence, no imputation or case deletion for missing values was required.
Construction of the evaluation index system
To comprehensively evaluate students' performance in English writing transfer training, this study constructed a multidimensional evaluation system that covered the three-level transfer goals of "syntax-structure-culture", integrated language ability, model response, teaching feedback, and self-cognition, and established six core indicators. First, the "writing transfer index" was a comprehensive evaluation indicator that covered syntactic diversity, structural clarity, and cultural adaptability, combining the system's automatic scoring and the teacher's manual scoring for quantitative analysis. The "syntactic complexity score" utilized natural language processing technology to extract features such as average sentence length, the number of clause nesting levels, and the frequency of verb phrase use to evaluate the richness and complexity of students' sentences. The "academic style matching" analysis examined the proportion of passive voice use and the proportion of formal vocabulary, based on an NLP (Natural Language Processing) model, to determine whether the text adhered to the style norms of English academic writing. In addition, the study introduced a subjective evaluation dimension, collecting feedback on the acceptance and practicality of the model-generated suggestions through teacher questionnaires to assess the actual impact of AI-assisted teaching. "Changes in teacher review scores" reflected the impact of AI-assisted teaching on the improvement of writing quality by comparing teachers' scores on students' essays before and after the experiment. Finally, "Student self-evaluation of transfer ability" utilized the transfer cognition table to guide students in self-evaluating their mastery of structure, syntax, culture, and other dimensions, thereby enhancing their metacognitive awareness and reflection abilities. The specific descriptions and data sources for each indicator are shown in Table 3.
Syntactic and stylistic indicators were automatically calculated via dependency parsing and BERT-based classification, with normalized scores ranging from 0 to 5; manual evaluation uses a 5-point scale. Calculation formulas for syntactic complexity used the total word count/clause count as the core denominator. Raw data source: pretest & posttest writing compositions of all 60 enrolled students.
Experimental results and analysis
To intuitively present the comprehensive performance differences of students in the three dimensions of syntactic transfer, structural transfer, and cultural transfer, Figure 3 compared the average writing transfer index of students in the experimental group and the control group before and after the experiment: Figure 3 showed the difference in average scores between the experimental group and the control group in the three dimensions of syntactic transfer, structural transfer, and cultural transfer. The experimental group showed a significant improvement in transfer ability in each dimension after the experiment, and the outer circle of the radar chart is significantly larger than the inner circle, indicating that the instruction optimization method based on the DeepSeek model has a positive effect in promoting the development of students' English writing transfer ability. In contrast, the control group's various transfer ability indicators remained relatively unchanged before and after the experiment. The overall improvement was limited, indicating that traditional teaching methods or AI-assisted tools that had not optimized system instructions have a limited role in transfer training. It was challenging to effectively stimulate students' learning potential in multidimensional language transfer.
The average writing transfer indexes of the experimental group in the dimensions of syntactic transfer, structural transfer, and cultural transfer before the experiment were 3.0, 2.9 and 2.8, respectively, which increased to 4.3, 4.1, and 4.5 after the experiment, with increases of 1.3, 1.2, and 1.7, respectively, indicating that the teaching method had a good intervention effect at different levels of transfer. The control group's scores before the experiment were 3.0, 3.0, and 2.9. After the experiment, the values increased to 3.4, 3.3, and 3.2, respectively, which are significantly lower than those of the experimental group. It is particularly noteworthy that in the dimension of cultural transfer, the experimental group showed the most significant improvement, increasing from 2.8 to 4.5, a 1.7-point increase, which is significantly higher than the 0.3-point increase of the control group. This result suggested that the proposed teaching method plays a significant role in helping students identify and adapt to English cultural expression norms. It was not only reflected in adjustments to language form but also in the enhancement of students' cross-cultural pragmatic awareness and the depth of their understanding of the target language's cultural expression habits. Language transfer was primarily reflected in syntactic complexity, language style adaptation, and other aspects. The study utilized automated NLP technology to conduct an in-depth analysis of students' writing texts, extracted key language features, and combined teacher scores to systematically evaluate changes in students' language transfer ability across three rounds of tasks.
Syntactic complexity and linguistic style
In terms of syntactic complexity, the study recorded the average sentence length and number of clause nesting levels of the text generated by students after completing the designated writing task in each round of tasks. These two indicators were widely used to measure the complexity of language expression. They effectively reflected the development level of students in sentence diversity and structural organization ability, as shown in Figure 4.
In the longitudinal evolution of syntactic complexity, the experimental group exhibited a clear upward trend in the three rounds of tasks, reflecting the effective promotion of transfer training on students' ability to express their language structure. From the perspective of average sentence length, the experimental group increased from 12.5 words in Task 1 to 16.1 words in Task 3, representing a 3.6-word increase, which was significantly higher than the control group's increase of only 0.9 words in the same period (from 12.4 to 13.3). Similarly, from the perspective of the structural complexity dimension, measured by the number of clause nesting layers, the experimental group increased from 1.8 layers to 2.7 layers, while the control group increased from 1.7 to 1.9 layers, with a relatively slow increase. To further verify whether the difference in syntactic complexity between the two groups of students is statistically significant, the study employed independent sample t-tests to analyze the average sentence length and the number of clause nesting levels at the three stages of the task. The results showed that there is a significant difference between the experimental group and the control group in terms of average sentence length, t(58) = 4.23, p < 0.001, Cohen’s d = 0.98; there was also a significant difference in the number of clause nesting levels, t(58) = 3.15, p = 0.002, Cohen’s d = 0.78. This comparison showed that systematic transfer training not only improved students’ ability to construct complex sentences but also strengthened their mastery of grammatical organization strategies. In particular, during the third stage of the task, students in the experimental group demonstrated greater flexibility in using multi-layered clauses and complex sentence structures. This reflected a notable shift in their language transfer from a “functional” level to a more “strategic” level. The transfer-oriented training pathway continuously enhanced students’ syntactic transfer skills, effectively overcoming the limitations of single-sentence writing and the fixed expression patterns often seen in traditional teaching. In terms of language style adaptation, the study extracted students' writing texts before and after the experiment. It used the NLP model to determine the proportion of formal vocabulary and the frequency of passive voice use as core indicators for evaluating academic language style matching. These indicators reflected whether students have the language awareness and expression ability to adapt to the target style in English writing. The relevant results are shown in Figure 5.
Figure 5 shows the changes in the academic style matching of the experimental group and the control group before and after the task, mainly focusing on the two core indicators of formal vocabulary proportion and passive voice usage frequency. Overall, after completing transfer training based on the DeepSeek model, the academic style adaptation ability of the experimental group was significantly improved, demonstrating the effectiveness of this teaching method in cultivating students' awareness of language norms and their ability to adapt to academic styles. Regarding the proportion of formal vocabulary, the experimental group increased from 0.42 before the task to 0.56 after the task, a relatively significant increase. In contrast, the control group increased only slightly, from 0.43 to 0.48, suggesting a limited improvement. This result showed that after receiving systematic prompt guidance and model feedback, the students in the experimental group were more inclined to use formal vocabulary that meets the requirements of academic style in the writing process, and their language style gradually approached the standards of English academic writing. In terms of the frequency of passive voice use, the experimental group increased significantly from 0.18 before the task to 0.27 after the task, representing an increase of 0.09; in contrast, the control group increased by only 0.02.
To further verify whether the above changes are statistically significant, the study conducted an independent sample t-test on the data before and after the task. The results showed that the experimental group had a significant improvement in the proportion of formal vocabulary, t(58) = 3.89, p < 0.001, Cohen's d = 0.92; the increase in the frequency of passive voice use was also significant, t(58) = 4.56, p < 0.001, Cohen’s d = 1.05. This showed that the students in the experimental group had made substantial progress in language style adaptation, and this progress was not accidental, but the result of the combined effect of systematic prompt guidance and model feedback.
Statistical verification
In addition, to further verify the reliability of AI-generated content in teaching practice, the study also compared the mean scores of AI-generated content and teacher-generated content under different prompt types. The consistency between the two was evaluated by calculating the Pearson correlation coefficient. The relevant comparison results are shown in Table 4. Table 4 shows the consistency of AI-generated content and teacher scores in five types of transfer-oriented prompts. From the perspective of the mean score, the AI score was slightly lower than the teacher score overall. Still, the gap was within a reasonable range in most task types, indicating that the DeepSeek model has high stability and reference value in evaluating writing transfer tasks. Under the syntactic transfer prompt, the average teacher score was 4.1, and the AI score was 4.0, with a difference of only 0.1 between the two. In the structural and logical transfer prompts, the AI score was only 0.1 points lower than the teacher's score, indicating that the model can better simulate the teacher's evaluation criteria in these tasks, which involve a high degree of language structure and clear rules. In contrast, the scoring deviations under the cultural transfer and register transfer prompts are relatively obvious, with AI scores of 3.6 and 3.7, respectively, which are 0.2 lower than the teacher's scores, reflecting that AI still had certain limitations when dealing with tasks with high subjectivity, such as semantic implications and cultural pragmatics. Inter-rater reliability was assessed via ICC (n = 5 raters). The overall ICC for manual scoring was 0.86, and ICCs for each dimension ranged from 0.82 to 0.89, indicating satisfactory inter-rater agreement.
Further analysis of the Pearson correlation coefficient revealed that the consistency of the scores under various prompts was at a high level, indicating that the AI score was not only close to the teacher's judgment in terms of value but also exhibits a good linear relationship in its scoring trend. Among them, the consistency coefficient of the syntactic transfer prompt was the highest at 0.87, and the structural transfer and logical transfer prompts were also 0.83 and 0.85, respectively, indicating that the AI's evaluation results in terms of syntactic complexity, paragraph organization, and logical coherence are highly consistent with the teacher's judgment. This may be because these tasks have clear objectives and quantifiable language features, which facilitated model recognition and stable judgment. The correlation coefficient of the domain transfer prompt is 0.81. Although it had decreased slightly, it still demonstrated a strong evaluation ability, indicating that AI possessed a certain degree of judgment at the level of stylistic adaptation. However, the consistency coefficient of the cultural transfer prompt was only 0.79, which was lower than other types, but still within an acceptable range.
To examine the applicability of different types of prompts in teaching practice and the degree of acceptance by teachers, a prompt response satisfaction questionnaire was designed. Five English teachers (numbered T1–T5) were invited to rate the generation suggestions of five types of prompts using a five-level scale (very dissatisfied to very satisfied), as shown in Figure 6: the scores of the five teachers on the five types of prompts vary somewhat, but overall recognition was high. Among them, the "syntactic transfer" and "cultural transfer" prompts scored the highest, with an average of 4.0, reflecting strong practicality and adaptability in teaching. In contrast, the "register transfer" prompt scores relatively low, averaging only 2.4, and some teachers, such as T4 and T5, gave the lowest scores, indicating that its guidance and adaptability in teaching applications still need improvement.
At the individual level, there were obvious differences in teachers' scoring tendencies. T1's overall score was positive, indicating a high degree of acceptance of AI-assisted writing, while T5's scores were low in multiple Prompt dimensions, especially in "register transfer" and "logic transfer". This difference might be related to teachers' teaching experience, language preference, or familiarity with AI tools, suggesting that future Prompt design needs to consider personalized adaptation mechanisms to enhance the acceptance of different teacher groups. To further verify the actual effect of AI-assisted teaching on improving students' writing quality, the study compared the changes in overall scores of the experimental group and the control group, as evaluated by teachers, before and after the experiment. The comparison results are shown in Figure 7.
As shown in Figure 7, the overall scores of the experimental group and the control group, as evaluated by teachers before and after the experiment, showed significant differences. Overall, after the instruction optimization teaching intervention based on the DeepSeek model, the experimental group's overall score showed a significant upward trend. In contrast, the score change in the control group was relatively gentle. The average teacher score of the experimental group before the experiment was 59.1 points, and the average score after the experiment increased significantly to 74.4 points, representing a 15.3-point increase. This indicated that AI-assisted teaching had a positive impact on students' writing quality. From the box plot, it was observed that the distribution range of the experimental group's scores had also expanded. In particular, the score box after the experiment is significantly higher than before, and the upper maximum and lower minimum values have increased, indicating that the overall level of students has improved significantly. In contrast, the score changes in the control group were relatively limited. Before the experiment, the average score of the control group was 60.1 points, and after the experiment, it was 64.9 points, representing a 4.8-point increase. This increase was much lower than that of the experimental group, and the scoring box of the control group did not change much before and after the experiment, indicating that traditional teaching methods or unoptimized AI-assisted tools have a limited effect on improving writing quality.
In addition, this study utilized the transfer cognition table to guide students through three rounds of self-assessment on their ability development in structural adjustment, sentence transformation, and cultural expression, among other areas, to reflect the changing trend of students' metacognitive awareness and their mastery of transfer strategies. The results are shown in Figure 8. According to the radar chart data of students' self-assessment of transfer ability shown in Figure 8, the self-assessment of students in the five dimensions of structural transfer, syntactic transfer, cultural transfer, register transfer and logical transfer in the three rounds of tasks showed a continuous increasing trend, indicating that under the instruction optimization teaching intervention based on the DeepSeek model, students' ability to use transfer strategies had been significantly improved. In the first round of tasks, the students' self-assessment scores were generally low. Among the five dimensions, "structural transfer" and "logical transfer" score 3.2 and 3.0, respectively, while "cultural transfer" and "register transfer" score 2.5 and 2.7, indicating that students were still immature in identifying and using transfer strategies. By the second round of tasks, the scores of each item had improved, with "structural transfer" rising to 3.8, "syntactic transfer" to 3.5, and "logical transfer" to 3.7. This showed that with the guidance of teachers and the support of model feedback, students had gradually mastered the key points of basic transfer operations and had initially applied them to actual writing. In the third round of tasks, students' self-evaluation scores increased significantly, with "structural transfer" and "logical transfer" reaching 4.5 and 4.3, respectively, and the other dimensions also stabilized above 4.0 points, showing the sustained effect of multiple rounds of transfer training in improving students' language form control and text construction ability. At this stage, students had formed a basic cognitive closed loop between understanding, execution, and reflection of transfer strategies. In addition to teacher feedback, the study also distributed a satisfaction survey questionnaire about the AI-assisted writing function to all students in the experimental group, and a total of 30 valid questionnaires were collected. The questionnaire assessed the performance of AI in three functional modules: rewriting suggestions, structural optimization, and cultural prompts, and requires students to select options according to five levels of standards. The results are shown in Table 5.
According to the results of the student satisfaction survey on AI-assisted writing functions shown in Table 5, students in the experimental group generally gave positive comments on the performance of AI in the three functional modules of rewriting suggestions, structural optimization, and cultural prompts, showing the good application prospects of AI-assisted writing systems in improving writing training efficiency and teaching support quality. Overall, more than 90% of students expressed "very satisfied" or "satisfied" in the two functional items of "rewriting suggestions" and "structural optimization", among which "rewriting suggestions" had the highest satisfaction, reaching 91%, indicating that students had a high degree of recognition of AI's ability in syntactic reconstruction and language polishing. In comparison, the satisfaction with the "cultural prompts" function was relatively lower. Still, it reached 83%, indicating that although AI has certain complexity and subjectivity in guiding cross-cultural expression, its role in assisting students to identify and adjust for native-language cultural interference was still recognized by most students. In terms of dissatisfaction, the percentage of students who chose "dissatisfied" or "very dissatisfied" across the three functions was less than 5%, indicating that AI-assisted functions had good usability and acceptance in most application scenarios. It is worth noting that the proportion of students who rated "average" in the "cultural prompts" item was relatively high, suggesting that the current AI might not fully meet students' expectations when dealing with cultural adaptation issues, and there was still room for improvement. A comprehensive analysis reveals that the DeepSeek-based instruction optimization method has been widely recognized by students, particularly for its support of language form.
RM-ANOVA
A two-way repeated-measures analysis of variance (RM-ANOVA) was conducted to examine the effects of group (between-subjects factor: experimental group vs. control group) and time (within-subjects factor: pre-test, Task 1, Task 2, Task 3, post-test), as well as their interaction on students’ English writing transfer performance. Mauchly’s test indicated a violation of the sphericity assumption (p < 0.05), so the Greenhouse–Geisser correction was applied to adjust the degrees of freedom for within-subjects effects. All analyses were based on n = 30 participants per group. The results are shown in Table 6:
The results of the repeated-measures ANOVA with Greenhouse–Geisser correction indicated significant main effects of Group, Time, and Group × Time interaction across all measured dimensions (p < 0.001). For the Overall Transfer Index, a significant Group effect was observed (F(1,58) = 76.32), alongside a significant Time effect (F(2.14,124.12) = 92.75) and a significant Group × Time interaction (F(2.14,124.12) = 68.49), suggesting that changes in overall transfer performance over time differed significantly between groups.
Similarly, syntactic transfer showed significant effects of Group (F(1,58) = 52.18), Time (F(2.11,122.38) = 71.26), and Group × Time interaction (F(2.11,122.38) = 45.73), indicating differential development patterns in syntactic transfer ability across groups and measurement points. For Structural Transfer, significant Group (F(1,58)=48.65), Time (F(2.09,121.22) = 67.81), and interaction effects (F(2.09,121.22) = 41.36) were also identified, reflecting consistent improvements with group-dependent trajectories. Notably, Cultural Transfer exhibited the strongest effects, with a significant Group effect (F(1,58) = 81.44), a pronounced Time effect (F(2.07,119.96) = 98.52), and a robust Group × Time interaction (F(2.07,119.96) = 79.27), suggesting the most substantial and differentiated growth pattern in this dimension. Overall, all indices demonstrate statistically significant effects, confirming that the intervention exerted a stable and differentiated impact on transfer development over time.
DATA AVAILABILITY:
All data generated or analyzed during this study are included in this article. The raw evaluation data are provided in Supplementary Table 1.

Figure 1: Steps for sentence structure transformation. (A) Strategies for combining and strengthening verbs to improve sentence structure. (B) Alternative subject transformations, including gerund, infinitive, and nominalized subjects. (C) Use of non-finite phrases to enhance sentence variety and conciseness. (D) Examples of polished final versions demonstrating improved academic style and cross-cultural English expression. Please click here to view a larger version of this figure.

Figure 2: Task evolution tree diagram. It presents the progressive structure of four-tier transfer tasks ranging from sentence, paragraph, discourse, to the cultural level. Please click here to view a larger version of this figure.

Figure 3: Comparison of writing transfer ability before and after the experiment. The radar chart shows pre-test and post-test scores of the experimental group and control group in syntactic, structural, and cultural transfer dimensions. Please click here to view a larger version of this figure.

Figure 4: Syntactic complexity trend line chart. Changes in average sentence length and clause nesting layers across three training tasks are displayed. Please click here to view a larger version of this figure.

Figure 5: Academic terminology matching. This figure shows variations in the proportion of formal vocabulary and the frequency of passive voice before and after the intervention. Please click here to view a larger version of this figure.

Figure 6: Prompt response satisfaction heat map. Ratings of five types of prompt templates from participating teachers are summarized here. Please click here to view a larger version of this figure.

Figure 7: Box plot of changes in teacher ratings. The box plot demonstrates the distribution and overall changes of teacher-assessed writing scores for two groups pre- and post-experiment. Please click here to view a larger version of this figure.

Figure 8: Radar chart of students' self-assessed transfer ability. It shows students’ self-evaluation scores across five transfer dimensions in three training rounds. Please click here to view a larger version of this figure.
| Input | Teacher Operation | Prompt Type | Student Operation | Output | Evaluation Criteria |
| Students’ original writing drafts (sentences/paragraphs/essays) | 1. Label transfer types and task levels2. Adjust prompt parameters if needed | Single transfer prompt / Three-stage nested prompt chain | 1. Learn transfer strategies 2. Revise drafts based on AI feedback 3. Complete self-assessment | AI revised text + Student revised final draft | Syntactic complexity, structural rationality, cultural appropriateness, logical coherence, register standardization (0–5 scale) |
Table 1: Workflow Summary Table. This table summarizes the complete operational workflow of the AI-assisted English writing transfer training, covering input materials, teacher and student operations, prompt types, final outputs and corresponding evaluation criteria.
| Project Category | Contents |
| Understanding of mission objectives | Please briefly describe the migration goals of this round of writing tasks, such as structural adjustment, cultural adaptation, and language conversion. |
| Migration strategy used | Please list the migration strategies you adopted (multiple selections are allowed): |
| Structural Adjustment |
| Sentence transformation |
| Language switching |
| Cultural manifestation |
| Strengthening logical connection |
| Other:_______ |
| Rewrite examples and instructions | Select 1–2 rewritten sentences, show the versions before and after the rewrite, and explain the reasons for the changes and the desired transfer goals. |
| Difficulties and doubts encountered | Describe any challenges you encountered during the migration or questions you had about a specific expression. |
Self-assessment score
(out of 5 points) | Please rate your text rewriting based on the following three aspects: |
| • Strategy application accuracy (/5) |
| • Natural fluency of expression (/5) |
| • Cultural fit (/5) |
| Future writing improvement goals | Briefly describe the transferability you would like to strengthen or optimize in the next round of training |
Table 2: Transfer cognition table. A 1–5 rating scale (1 = no mastery, 5 = full mastery) is adopted for students’ self-assessment of writing transfer strategies.
| Indicator name | Description | Data Source |
| Writing transfer Index (0–5) | A quantitative indicator that comprehensively measures language transfer ability, covering three levels: syntax, structure, and culture. | Automatic scoring by the system + manual scoring by teachers |
| Syntactic complexity score | Including average sentence length, number of clause embeddings, verb phrase richness, etc. | NLP automatic analysis |
| Academic style matching | Does it conform to English academic writing standards? | Natural language processing model judgment |
| Prompt response satisfaction | Teachers’ acceptance and practical evaluation of model-generated suggestions | Teacher Questionnaire |
| Teacher review grading changes | Comparison of teacher scores before and after the experiment to reflect the improvement in writing quality | Teacher Rating Records |
| Students' self-assessed transfer ability | Students self-assess their mastery of different transfer dimensions | Fill out the migration awareness form |
Table 3: Summary of evaluation indicators, descriptions, and data sources. This table clarifies the definitions, measurement methods, and original data sources for each core evaluation indicator in this study.
| Prompt type | Teachers' average rating | AI score average | Pearson correlation coefficient |
| Structural transfer | 4 | 3.9 | 0.83 |
| Syntactic transfer | 4.1 | 4 | 0.87 |
| Cultural transfer | 3.8 | 3.6 | 0.79 |
| Register transfer | 3.9 | 3.7 | 0.81 |
| Logical transfer | 4.2 | 4.1 | 0.85 |
Table 4: Comparison of consistency between AI-generated content and teacher ratings. The results demonstrate the consistency between AI-generated content and teacher ratings across different Prompt types.
| Function items | Very satisfied | Satisfied | Generally | Dissatisfied | Very dissatisfied |
| Rewrite suggestion | 66% | 25% | 7% | 1.50% | 0.50% |
| Structural optimization | 60% | 30% | 7% | 2% | 1% |
| Cultural tips | 55% | 28% | 12% | 3.50% | 1.50% |
Table 5: Survey statistics on students' satisfaction with AI-assisted functions. The statistics show the distribution of student satisfaction for AI rewriting, structural optimization, and cultural prompt functions.
| Measure | Group F(df) | Time F(df)GG | Group × Time F(df)GG | p |
| Overall Transfer Index | 76.32 (1, 58) | 92.75 (2.14, 124.12) | 68.49 (2.14, 124.12) | <0.001 |
| Syntactic Transfer | 52.18 (1, 58) | 71.26 (2.11, 122.38) | 45.73 (2.11, 122.38) | <0.001 |
| Structural Transfer | 48.65 (1, 58) | 67.81 (2.09, 121.22) | 41.36 (2.09, 121.22) | <0.001 |
| Cultural Transfer | 81.44 (1, 58) | 98.52 (2.07, 119.96) | 79.27 (2.07, 119.96) | <0.001 |
Table 6: Summary of repeated-measures ANOVA results. It presents the two-way repeated-measures ANOVA results for overall writing transfer performance and its three sub-dimensions, including main effects of group, time, and the group-by-time interaction. Abbreviations: GG = Greenhouse–Geisser correction.
Supplementary Table 1: Raw data evaluation. Includes the raw data used in all analyses in this study.Please click here to download this file.