$$\rightleftharpoonup{xx}$$
$$\longleftharp{xx}$$,
$$\longrightharp{xx}$$,
The intensive globalization of education and digital technologies has increased the need to evaluate the level of writing in English scientifically and credibly for language teaching, academic development, and career advancement1. Conventional writing evaluations, as practiced by human rating, can measure subjective aspects of writing like the thoroughness of argumentation and cultural suitability2, but are susceptible to long turnaround times, high labor expenses, and bias due to evaluator experience and leanings3,4. These constraints are especially acute in large-scale practice, like international language tests (IELTS, TOEFL) or other courses in English taught in universities where manual scoring cannot be all that is required in terms of instant feedback and coverage5.
AWE systems have become widely used in this context due to their real-time processing, standardization, and scalability6. Such popular tools as Grammarly (which focuses on grammar errors and style refinement) and ETS Criterion (which adheres to formal writing norms) are currently used by millions of students in K-12 education, language schools, higher education, and individual training7. Although these are the benefits, the technological efficiency and education applicability of AWE systems are still disputed8. Technically speaking, the existing systems are highly accurate on objective dimensions, including error detection and lexical diversity, where the correlation with human scoring can be over 0.859. However, in more subjective areas, such as content relevance, logical argumentation and organization of a text, the correlations often become lower than 0.7010. Such a disproportion has the danger of promoting superficial accuracy among the learners at the cost of overall competence in writing11.
The issue of equity also limits the educational usefulness of AWE. The current studies are also inclined to focus on the aggregate indicators of accuracy, neglecting the possibility of deviations that systematically disadvantage some group12. Indicatively, characteristics of interlanguage shared by Chinese or Spanish learners would be mistaken as errors, and this would result in systematic underestimation13,14. In addition, the subjective acceptance of AI feedback by learners is generally little known15. Surveys indicate that almost one-third of non-native learners report an inappropriateness between AI scores and actual performance, with the processes of technical accuracy, group equity, and learner satisfaction still being poorly comprehended16.
These weaknesses reflect the shortcomings of the classical paradigm of accuracy17. A framework that considers only the alignment between AI and human scoring cannot capture issues of equity or the learner's trust in the system. In practice, the educational value of AWE must satisfy three conditions simultaneously: technical precision, fairness across groups, and learner acceptance18. The absence of such a comprehensive validation approach helps explain why AWE systems enjoy widespread adoption yet limited trust in educational practice19,20.
To address this challenge, the present study introduces a multi-level validation framework that integrates technical accuracy, group and individual fairness, and learner perception into a coherent structure. The proposed XAI framework is designed to be practically implemented within existing AWE platforms by providing teachers and students with fairness diagnostics and transparent score explanations, and can be applied in writing courses or test-preparation classes to evaluate its ability to enhance fairness, interpretability, and instructional usefulness in real assessment settings.
In this context, the hypothesis is an AFMM to investigate the mediating role of perceived fairness in determining the relationship between accuracy and satisfaction, as well as the moderating role of language proficiency on fairness sensitivity. Therefore, it contributes in two ways both theoretically by enriching the evaluation models of AWE by describing fairness as one of the key validation dimensions alongside accuracy and perception, and practically, by providing developers with strategies to maximize fairness, educators with group-sensitive system selection criteria, and the educational value of AWE by explaining the manner in which the perceptions of the learners are formed. In addition to education, the framework is also aligned with the broader concept of XAI, demonstrating how fairness and user perception can enhance transparency, trust, and acceptance in other areas, such as healthcare, autonomous systems, and cybersecurity.
Research Questions:
1.To what extent does the AWE system demonstrate technical accuracy and fairness across different native-language and proficiency groups?
2.How can an XAI-based multi-level evaluation framework improve transparency and equity in automated English writing assessment?
LITERATURE REVIEW:
The factors that affect the acceptance of AWE feedback by college students were examined using an extended Technology Acceptance Model (TAM)21. Based on survey data from 448 Chinese students using SEM, it was determined that usefulness, ease of use, and intention had a significant influence on subjective norm, trust, self-efficacy, cognitive feedback, and system characteristics. However, the study was limited to a single nation and a single group of students, which limits the applicability of generalization. To explore how Chinese EFL students respond to Pigai AWE feedback22, a study analyzed repeated submissions (n = 5) from university students. It noted an early emphasis on error correction, low intake of linguistic feedback, and gradual deepening of response. However, the sample size was very limited, as was the AWE system, which restricts applicability and generalizability. The beliefs held by EFL teachers regarding the application of the AI grading tool (CoGrader) were examined to identify the factors that influence their views23. Through a mixed-method study of 10 Saudi university teachers, a survey and an interview revealed that there was a mixed positive opinion, but reluctance to be completely sure of reliability and complete teacher replacement. This impedes the generalization due to the limited sample and the one-country setting.
Considering developments in corpus linguistics and AI technology, a study investigated AES frameworks24. It employed PCA to improve linguistic indicators for evaluating writing quality and discovered that combining micro-characteristics with aggregated characteristics defined writing quality more effectively than aggregated characteristics alone. The non-linear AES approach based on Random Forest Regression surpassed the other approaches. Furthermore, SHAP identified essential language elements for each evaluated attribute, increasing system transparency via explainable AI. The results possibly help to enhance multidimensional methods in writing evaluation and education. The human-machine collaboration system was introduced to address the challenges of annotating Arabic writings, which are often expensive and time-consuming. The method considers essays based on seven features of literature with the help of LLM. Validation processes and prompting tactics were personalized to ensure consistency and accuracy. The cooperation results in a higher supply of labeled resources and does not affect the quality of evaluation, demonstrating it to be a scalable data annotation method suitable for lower-resource languages.
The use of AI in the educational sphere offers an opportunity to decrease grading requirements significantly and enhance writing education25,26. At the same time, researchers have emphasized that the accuracy of AI is not the only aspect relevant to its responsible use. There are principles of fairness and bias reduction, security and privacy, accountability, explainability, transparency, educational effect, integrity, and continuous development. Recent research has empirically evaluated zero-shot scoring based on GPT-4o with a focus on these requirements. The research focused on the perceptions that educators held towards ADWTs regarding the aspect of educational integrity27. The cross-sectional study involving 100 graduate students and professors of 10 subjects suggests that, despite teachers attributing the benefits of ADWTs in achieving the educational objective, it has some limitations, such as limited accessibility, lack of knowledge, and worry about its impact on integrity and creativity. The research suggested that as AI technologies become more integrated into education, ethical concerns and stakeholder participation are necessary for their successful and responsible use. Research investigated the efficacy of AI technologies compared to human assessors in evaluating essays submitted by EFL pupils28. Assessing 30 essays revealed that, while AI offered high-quality comments in terms of content, language, organization, and correctness, it constantly provided lower ratings than human raters. Furthermore, AI provided more thorough feedback, but the scores from various AI tools were not substantially different.
Research Gap:
Currently, most research on AWE scholarship examines either accuracy or user acceptance. Very few examine if scoring differences systematically disadvantage native-language or proficiency groups. While previous studies have examined user acceptance or are limited to a specific AWE system from a specific country and sample size, questions around generalizability arise. Although both SHAP and PCA are XAI strategies and were developed to increase transparency, no studies have examined fairness mechanisms or how learners use AI feedback from the AWE. There are no extensive frameworks in the literature that contemplate defined dimensions of accuracy, fairness analysis, and learner perceptions. There is no example of an explainable model of evaluation that considers intra- and inter-rater accuracy, fairness, and learner perceptions. An explainable framework, TLEF, and a combined model, AFMM, are proposed and validated in this research to assess accuracy, fairness, and learner perceptions at the same time among multilingual and proficiency diverse learners.