Research Article

Explainable AI Framework for Accuracy, Fairness, and Learner Perception in English Writing Assessment

DOI:

10.3791/69841

December 23rd, 2025

In This Article

Summary

Loading...
$$\rightleftharpoonup{xx}$$ $$\longleftharp{xx}$$, $$\longrightharp{xx}$$,

This study develops a three-tier evaluation framework and fairness mediation model to assess AI-assisted English writing systems. Using 764 cross-linguistic samples, results show accuracy disparities, fairness bias against non-native learners (especially Chinese A2 proficiency level), and fairness perception as the key mediator of user satisfaction, offering theoretical and practical implications.

Abstract

Loading...
$$\rightleftharpoonup{xx}$$ $$\longleftharp{xx}$$, $$\longrightharp{xx}$$,

In the context of global educational digital transformation, automated writing evaluation (AWE) has been widely adopted due to its real-time and standardized advantages; however, traditional accuracy-oriented frameworks often neglect equity concerns and learners' perceptions, thereby limiting transparency and educational value. To address this limitation, this research proposes an explainable AI (XAI) framework designed to provide transparent and interpretable feedback, allowing learners to understand and trust automated evaluation, and integrates a multi-level validation model, the Three-Level Evaluation Framework (TLEF), spanning technical accuracy, group and individual equity, and learner perception, together with the AI Fairness Mediation Model (AFMM). Using stratified random sampling, data were collected from 764 multilingual learners (native speakers of English, Chinese, and Spanish) across Common European Framework of Reference for Languages (CEFR) levels A2 to C1 through writing tasks, dual scoring by AI and human experts, and structured questionnaires. Instead of listing individual tests, multiple statistical analysis was employed to examine validity, fairness, and the learner-perception relationship. Statistical analyses combined correlation, root mean square error (RMSE), Equalized Odds testing, and Structural Equation Modeling (SEM). The findings reveal that while the AI-assisted writing evaluation (AWE) system (ETS Criterion) achieves overall validity (r = 0.82), significant disparities remain: Chinese native speakers demonstrate the lowest agreement with human raters (0.72) and the highest RMSE (median 2.15), fairness biases are most pronounced at lower proficiency levels (ΔEO = 0.15 for A2 learners), and perceived fairness fully mediates the link between perceived accuracy and learner satisfaction, with proficiency moderating fairness sensitivity. By reframing fairness and perception as essential dimensions of explainability, the research enhances the theoretical grounding of AWE and provides a practical pathway for increasing transparency, equity, and social acceptance in educational technologies.

Introduction

Loading...
$$\rightleftharpoonup{xx}$$ $$\longleftharp{xx}$$, $$\longrightharp{xx}$$,

The intensive globalization of education and digital technologies has increased the need to evaluate the level of writing in English scientifically and credibly for language teaching, academic development, and career advancement1. Conventional writing evaluations, as practiced by human rating, can measure subjective aspects of writing like the thoroughness of argumentation and cultural suitability2, but are susceptible to long turnaround times, high labor expenses, and bias due to evaluator experience and leanings3,4. These constraints are especially acute in large-scale practice, like international language tests (IELTS, TOEFL) or other courses in English taught in universities where manual scoring cannot be all that is required in terms of instant feedback and coverage5.

AWE systems have become widely used in this context due to their real-time processing, standardization, and scalability6. Such popular tools as Grammarly (which focuses on grammar errors and style refinement) and ETS Criterion (which adheres to formal writing norms) are currently used by millions of students in K-12 education, language schools, higher education, and individual training7. Although these are the benefits, the technological efficiency and education applicability of AWE systems are still disputed8. Technically speaking, the existing systems are highly accurate on objective dimensions, including error detection and lexical diversity, where the correlation with human scoring can be over 0.859. However, in more subjective areas, such as content relevance, logical argumentation and organization of a text, the correlations often become lower than 0.7010. Such a disproportion has the danger of promoting superficial accuracy among the learners at the cost of overall competence in writing11.

The issue of equity also limits the educational usefulness of AWE. The current studies are also inclined to focus on the aggregate indicators of accuracy, neglecting the possibility of deviations that systematically disadvantage some group12. Indicatively, characteristics of interlanguage shared by Chinese or Spanish learners would be mistaken as errors, and this would result in systematic underestimation13,14. In addition, the subjective acceptance of AI feedback by learners is generally little known15. Surveys indicate that almost one-third of non-native learners report an inappropriateness between AI scores and actual performance, with the processes of technical accuracy, group equity, and learner satisfaction still being poorly comprehended16.

These weaknesses reflect the shortcomings of the classical paradigm of accuracy17. A framework that considers only the alignment between AI and human scoring cannot capture issues of equity or the learner's trust in the system. In practice, the educational value of AWE must satisfy three conditions simultaneously: technical precision, fairness across groups, and learner acceptance18. The absence of such a comprehensive validation approach helps explain why AWE systems enjoy widespread adoption yet limited trust in educational practice19,20.

To address this challenge, the present study introduces a multi-level validation framework that integrates technical accuracy, group and individual fairness, and learner perception into a coherent structure. The proposed XAI framework is designed to be practically implemented within existing AWE platforms by providing teachers and students with fairness diagnostics and transparent score explanations, and can be applied in writing courses or test-preparation classes to evaluate its ability to enhance fairness, interpretability, and instructional usefulness in real assessment settings.

In this context, the hypothesis is an AFMM to investigate the mediating role of perceived fairness in determining the relationship between accuracy and satisfaction, as well as the moderating role of language proficiency on fairness sensitivity. Therefore, it contributes in two ways both theoretically by enriching the evaluation models of AWE by describing fairness as one of the key validation dimensions alongside accuracy and perception, and practically, by providing developers with strategies to maximize fairness, educators with group-sensitive system selection criteria, and the educational value of AWE by explaining the manner in which the perceptions of the learners are formed. In addition to education, the framework is also aligned with the broader concept of XAI, demonstrating how fairness and user perception can enhance transparency, trust, and acceptance in other areas, such as healthcare, autonomous systems, and cybersecurity.

Research Questions:

1.To what extent does the AWE system demonstrate technical accuracy and fairness across different native-language and proficiency groups?

2.How can an XAI-based multi-level evaluation framework improve transparency and equity in automated English writing assessment?

LITERATURE REVIEW:

The factors that affect the acceptance of AWE feedback by college students were examined using an extended Technology Acceptance Model (TAM)21. Based on survey data from 448 Chinese students using SEM, it was determined that usefulness, ease of use, and intention had a significant influence on subjective norm, trust, self-efficacy, cognitive feedback, and system characteristics. However, the study was limited to a single nation and a single group of students, which limits the applicability of generalization. To explore how Chinese EFL students respond to Pigai AWE feedback22, a study analyzed repeated submissions (n = 5) from university students. It noted an early emphasis on error correction, low intake of linguistic feedback, and gradual deepening of response. However, the sample size was very limited, as was the AWE system, which restricts applicability and generalizability. The beliefs held by EFL teachers regarding the application of the AI grading tool (CoGrader) were examined to identify the factors that influence their views23. Through a mixed-method study of 10 Saudi university teachers, a survey and an interview revealed that there was a mixed positive opinion, but reluctance to be completely sure of reliability and complete teacher replacement. This impedes the generalization due to the limited sample and the one-country setting.

Considering developments in corpus linguistics and AI technology, a study investigated AES frameworks24. It employed PCA to improve linguistic indicators for evaluating writing quality and discovered that combining micro-characteristics with aggregated characteristics defined writing quality more effectively than aggregated characteristics alone. The non-linear AES approach based on Random Forest Regression surpassed the other approaches. Furthermore, SHAP identified essential language elements for each evaluated attribute, increasing system transparency via explainable AI. The results possibly help to enhance multidimensional methods in writing evaluation and education. The human-machine collaboration system was introduced to address the challenges of annotating Arabic writings, which are often expensive and time-consuming. The method considers essays based on seven features of literature with the help of LLM. Validation processes and prompting tactics were personalized to ensure consistency and accuracy. The cooperation results in a higher supply of labeled resources and does not affect the quality of evaluation, demonstrating it to be a scalable data annotation method suitable for lower-resource languages.

The use of AI in the educational sphere offers an opportunity to decrease grading requirements significantly and enhance writing education25,26. At the same time, researchers have emphasized that the accuracy of AI is not the only aspect relevant to its responsible use. There are principles of fairness and bias reduction, security and privacy, accountability, explainability, transparency, educational effect, integrity, and continuous development. Recent research has empirically evaluated zero-shot scoring based on GPT-4o with a focus on these requirements. The research focused on the perceptions that educators held towards ADWTs regarding the aspect of educational integrity27. The cross-sectional study involving 100 graduate students and professors of 10 subjects suggests that, despite teachers attributing the benefits of ADWTs in achieving the educational objective, it has some limitations, such as limited accessibility, lack of knowledge, and worry about its impact on integrity and creativity. The research suggested that as AI technologies become more integrated into education, ethical concerns and stakeholder participation are necessary for their successful and responsible use. Research investigated the efficacy of AI technologies compared to human assessors in evaluating essays submitted by EFL pupils28. Assessing 30 essays revealed that, while AI offered high-quality comments in terms of content, language, organization, and correctness, it constantly provided lower ratings than human raters. Furthermore, AI provided more thorough feedback, but the scores from various AI tools were not substantially different.

Research Gap:

Currently, most research on AWE scholarship examines either accuracy or user acceptance. Very few examine if scoring differences systematically disadvantage native-language or proficiency groups. While previous studies have examined user acceptance or are limited to a specific AWE system from a specific country and sample size, questions around generalizability arise. Although both SHAP and PCA are XAI strategies and were developed to increase transparency, no studies have examined fairness mechanisms or how learners use AI feedback from the AWE. There are no extensive frameworks in the literature that contemplate defined dimensions of accuracy, fairness analysis, and learner perceptions. There is no example of an explainable model of evaluation that considers intra- and inter-rater accuracy, fairness, and learner perceptions. An explainable framework, TLEF, and a combined model, AFMM, are proposed and validated in this research to assess accuracy, fairness, and learner perceptions at the same time among multilingual and proficiency diverse learners.

Access restricted. Please log in or start a trial to view this content.

Protocol

Loading...
$$\rightleftharpoonup{xx}$$ $$\longleftharp{xx}$$, $$\longrightharp{xx}$$,

The ethical approval and participant recruitment process, including essay administration, dual scoring by ETS Criterion and experts, learner-perception evaluation, and statistical analysis, are summarized in this section. It highlights how accuracy, fairness, and SEM-based perception modeling are integrated into a unified XAI validation pipeline. The XAI-driven AWE evaluation framework is illustrated in Figure 1.

Procedure:

The procedure involved several steps. First, IRB approval was obtained, and informed consent was collected from all participants. Independent, dependent, and control variables were then defined. Standardized writing tasks were administered on Moodle using three neutral essay topics, and writing samples were collected while ensuring adherence to essay requirements, such as word count, time limit, and structure. Dual scoring was conducted using ETS Criterion outputs combined with human expert ratings. Learner-perception questionnaires were distributed immediately after essay submission. Data screening and quality control procedures were implemented to address anomalies, such as cheating or invalid responses. Fairness analysis thresholds (ΔEO, RMSE checks) were also applied. Finally, all anonymized data were stored securely on encrypted, access-controlled servers.

Ethical approval and informed consent

This study received ethics approval from the Institutional Review Board of the authors' institution. All procedures were conducted in accordance with the Declaration of Helsinki and applicable regulations. All participants were adults (≥18 years) and provided written informed consent before participation. Writing samples and questionnaire responses were de-identified at source and stored on encrypted, access-controlled servers; only authorized investigators had access. Human raters were blinded to participants' native language, proficiency level, and demographics. Participation was voluntary, with the right to withdraw at any time, and no deception or sensitive interventions were involved. Formal approval documentation can be provided to the journal upon request.

Variable design

A total of three groups of variables were defined in the study to guide the analysis. Table 1 summarizes the measurement and data types used in measurement methods for each construct and provides the full operational definitions of the independent, dependent, and control variables.

AI scoring accuracy was the first independent variable assessed in terms of RMSE and Pearson correlation coefficient (r) between the outputs of ETS Criterion and the ratings of the experts. Calibration performed by experts yielded an ICC of 0.91, validating reliability.

The second independent variable was the linguistic background of the learners, which was divided into native and non-native speakers, and further subdivision was made into Chinese, Spanish, Arabic, and other groups. Chinese students were one of the target populations because preliminary indications of systematic underestimation were observed.

The third independent variable was writing proficiency, which was rated according to the CEFR levels A2 to C1, as confirmed by official certificates and pre-class proficiency tests, and was also aligned with IELTS equivalencies. Another moderator introduced in the AI Fairness Mediation Model was writing proficiency to test whether sensitivity to fairness differs across proficiency levels.

Perception of fairness and learner satisfaction were the dependent variables. Perception of fairness was assessed by means of an eight-item questionnaire rated on a seven-point Likert scale, which included the individual consistency and group impartiality (Cronbachs 87; CVI 92). The satisfaction of learners was assessed using six Likert questions that indicated willingness to use and perceived improvement in skill (α = 0.85).

The variables were controlled for in terms of age, sex, and writing experience. Age was divided into three groups (18-22 years, 23-28 years, and ≥29 years), and gender was categorized into male and female. Writing experience was categorized into three levels of frequency per year.

Writing task texts

Standardized argumentative essay prompts were formulated to obtain writing data for three neutral topics, Globalization's Impact on Local Cultures, Advantages and Challenges of Online Education, and Ethical Boundaries of Artificial Intelligence. These themes were aimed at balancing cognitive difficulty and accessibility on the one hand, and reducing performance differences due to previous knowledge on the other. The distribution of topics and descriptive statistics for essay length are reported in Table 2.

Each essay was required to be 250 words ±10% and written within 45 minutes on a Moodle-based platform. Auxiliary tools were prohibited, and late submissions were excluded. Essays followed a standardized structure of introduction, two argument paragraphs, and conclusion. In total, 764 valid essays were collected, with an average length of 252.3 words (SD = 8.7).

Scoring comparison data

Accuracy of AWE scoring was assessed using a dual procedure that combined ETS Criterion outputs with human expert ratings. Scores were retrieved from Criterion via its open API. Three linguists with more than ten years of assessment experience independently scored all essays. Before formal scoring, the raters completed three calibration sessions. During calibration, inter-rater reliability reached ICC = 0.87; during formal scoring, ICC rose to 0.91, with dimension-specific ICCs above 0.88. Essays with score discrepancies greater than two points were resolved collectively (18 cases). The scoring workflow and reliability outcomes are summarized in Table 3.

Learner perception questionnaire

Learners' perceptions of AI feedback were captured through a 22-item questionnaire based on the TAM and extended to include fairness. The instrument contained three domains: fairness perception (8 items), satisfaction (6 items), and moderating factors such as comprehensibility and transparency (8 items). Validation by five experts yielded a CVI of 0.92, and pilot testing with 60 learners produced an overall reliability of α = 0.90. The questionnaire structure and psychometric indices are provided in Table 4.

Questionnaires in the main study were administered right after submission of the essays, and there were minimum completion time requirements to diminish thoughtless completion. Out of the 764 surveys issued, 756 were valid following quality checks, and a resultant effective rate of 98.95 was obtained.

Data collection and quality control

The data were recorded for 8 weeks (March-April 2024) in four stages: recruitment and consent; essay writing; dual scoring and questionnaire distribution; and compilation of the database. The proficiency certificates based on pre-class writing performance were reviewed through dual screening, and this process eliminated 16 participants. Four potential cases of cheating were eliminated by real-time monitoring, and three suspect AI performances (deviations of at least 8 points) were subsequently amended following a manual assessment. Eight invalid questionnaires were eliminated based on reverse-item consistency checks.

Data storage and ethics

All the data were anonymized and stored using unique identifiers that consisted of the native language, proficiency level, and serial number. Texts, scores, and questionnaires were encrypted and stored on ISO27001-compliant servers with restricted access. Data will be retained for 3 years before permanent deletion. Ethical approval was obtained from the institutional review board, and written informed consent was collected from all participants.

Access restricted. Please log in or start a trial to view this content.

Results

Loading...
$$\rightleftharpoonup{xx}$$ $$\longleftharp{xx}$$, $$\longrightharp{xx}$$,

The section presents the research results based on five analytical dimensions: experimental design, participant characteristics, scoring accuracy, fairness assessment, and modeling of learning and perception. The outcomes include statistical performance, group differences, fairness disparities, and SEM-based mediation and moderation.

Experimental setup

The key software steps involved setting up ETS Criterion through its API to automatically score th...

Access restricted. Please log in or start a trial to view this content.

Discussion

Loading...
$$\rightleftharpoonup{xx}$$ $$\longleftharp{xx}$$, $$\longrightharp{xx}$$,

The research explored an AWE system under a three-level approach, encompassing technical accuracy, group and individual fairness, and learner perception, and identified that overall validity and systematic group differences are simultaneously present. There were strong correlations between AI and expert ratings (aggregate r = 0.82), but differences were observed by subgroup (native r = 0.89 vs. non-native r = 0.76; Chinese r = 0.72; Table 6). The distributions of RMSEs also indicated higher errors and variability in Chin...

Access restricted. Please log in or start a trial to view this content.

Disclosures

Loading...
$$\rightleftharpoonup{xx}$$ $$\longleftharp{xx}$$, $$\longrightharp{xx}$$,

The author has no conflicts of interest to disclose.

Materials

List of materials used in this article
NameCompanyCatalog NumberComments
Data Storage SystemEncrypted, access-controlled servers for storing anonymized data.Institutional serversSTORAGE-002
ETS Criterion SystemAI-assisted writing evaluation system used for scoring the writing tasks.Educational Testing Service (ETS)ETS-001
Fairness and Accuracy Analysis ToolsTools for RMSE, Equalized Odds, and statistical analysis.Custom scripts/stat packagesTOOL-FA-001
Human Expert RatingsIndependent ratings provided by three linguists with over 10 years of experience.In-house ratersHR-EXP-003
Learner Perception QuestionnaireAn 8-item questionnaire on fairness and satisfaction, rated on a 7-point Likert scale.In-house developedQUES-008
Statistical Software (R 4.3.1)Used for data analysis, including SEM (Structural Equation Modeling).R FoundationR-SW-431
Stratified Random Sampling DataData collected from 764 multilingual learners across CEFR levels A2 to C1.Study participantsDATA-764
Writing Task PromptsThree standardized essay topics on globalization, online education, and AI ethics.Moodle-based platformPROMPT-003

References

Loading...
$$\rightleftharpoonup{xx}$$ $$\longleftharp{xx}$$, $$\longrightharp{xx}$$,
  1. Voogt, J., Roblin, N. P. 21st century skills. Discussienota. 23 (03), 2000(2000).
  2. Weigle, S. C. Assessing writing. , Cambridge University Press. (2002).
  3. Barkaoui, K. Do ESL essay raters' evaluation criteria change with experience? A mixed-methods, cross-sectional study. TESOL Q. 44 (1), 31-57 (2010).
  4. Bitchener, J., Knoch, U. The contribution of written corrective feedback to language development: A ten-month investigation. Appl Linguist. 31 (2), 193-214 (2009).
  5. Chapelle, C. A., Douglas, D. Assessing language through computer technology. , Cambridge University Press. (2006).
  6. Aldosemani, T. I., et al. Automated writing evaluation in EFL contexts. Int J Comput Assist Lang Learn Teach. 13 (1), 1-19 (2023).
  7. Dikli, S. An overview of automated scoring of essays. J Technol Learn Assess. 5 (1), (2006).
  8. Stevenson, M., Phakiti, A. The effects of computer-generated feedback on the quality of writing. Assess Writ. 19, 51-65 (2014).
  9. Attali, Y., Burstein, J. Automated essay scoring with e-rater v.2. J Technol Learn Assess. 4 (3), (2006).
  10. Shermis, M. D., Burstein, J. Handbook of automated essay evaluation Routledge. , (2013).
  11. Hyland, K., Hyland, F. Feedback in second language writing: Contexts and issues. , Cambridge University Press. (2019).
  12. Williamson, D. M., Xi, X., Breyer, F. J. A framework for evaluation and use of automated scoring. Educ Meas Issues Pract. 31 (1), 2-13 (2012).
  13. Selinker, L. Interlanguage. Int Rev Appl Linguist. 10, 209-241 (1972).
  14. Odlin, T. Cross-linguistic influence. Handbook of second language acquisition. , 436-486 (2003).
  15. Ranalli, J., Link, S., Chukharev-Hudilainen, E. Automated writing evaluation for formative assessment. Educ Psychol. 37 (1), 8-25 (2017).
  16. Picón, A., Castro, I., Roldán, J. L. The relationship between satisfaction and loyalty: A mediator analysis. J Bus Res. 67 (5), 746-751 (2014).
  17. Weigle, S. English language learners and automated scoring of essays: Critical considerations. Assess Writ. 18, 85-99 (2013).
  18. Messick, S. Educational measurement  Macmillan. Linn, R. L. , 13-103 (1989).
  19. Floridi, L., Cowls, J. A unified framework of five principles for AI in society. Machine learning and the city. , 535-545 (2022).
  20. Doshi-Velez, F., Kim, B. Towards a rigorous science of interpretable machine learning. arXiv preprint. , (2017).
  21. Zhai, N., Ma, X. Automated writing evaluation feedback. Comput Assist Lang Learn. 35 (9), 2817-2842 (2022).
  22. Yang, H., Gao, C., Shen, H. Learner interaction with AI-programmed AWE feedback. Educ Inf Technol. 29 (4), 3837-3858 (2024).
  23. Alsalem, M. S. EFL teachers' perceptions of an AI grading tool. Cogent Educ. 11 (1), 2430865(2024).
  24. Tang, X., et al. Incorporating linguistic features and explainable AI in automated writing assessment. Appl Sci. 14 (10), 4182(2024).
  25. Elsayed, Y., et al. ZaQQ: A new Arabic dataset for automatic essay scoring. Data. 10 (9), 148(2025).
  26. Johnson, M., Zhang, M. Responsible use of zero-shot AI for essay scoring. Sci Rep. 14 (1), 30064(2024).
  27. Gustilo, L., Ong, E., Lapinid, M. Algorithmically-driven writing and academic integrity. Int J Educ Integr. 20 (1), 3(2024).
  28. Almegren, A., et al. Evaluating the quality of AI feedback. Innov Educ Teach Int. 62 (6), 1-16 (2024).
  29. Attali, Y., Burstein, J. Automated essay scoring with e-rater v.2. J Technol Learn Assess. 4 (3), (2006).
  30. Hardt, M., Price, E., Srebro, N. Equality of opportunity in supervised learning. arXiv preprint arXiv:1610.02413. , (2016).
  31. Davis, F. Perceived usefulness, perceived ease of use, and user acceptance of information technology. MIS Q. 13 (3), 319(1989).
  32. Colquitt, J. A., et al. Justice at the millennium: A meta-analytic review of 25 years of organizational justice research. J Appl Psychol. 86 (3), 425-445 (2001).
  33. Greenberg, J. A taxonomy of organizational justice theories. Acad Manag Rev. 12 (1), 9-22 (1987).
  34. Odlin, T. Language transfer: Cross-linguistic influence in language learning. , Cambridge University Press. (1989).
  35. Vandenberg, R., Lance, C. A review and synthesis of the measurement invariance literature. Organ Res Methods. 3 (1), 4-69 (2000).
  36. Kane, M. T. Validating the interpretations and uses of test scores. J Educ Meas. 50 (1), 1-73 (2013).
  37. Pleiss, G., et al. On fairness and calibration. arXiv preprint arXiv:1709.02012. , (2017).
  38. Mitchell, M., et al. Proceedings of the Conference on Fairness, Accountability, and Transparency. ACM. , 220-229 (2019).
  39. Knoch, U. Rating scales for diagnostic assessment of writing. Assess Writ. 16 (2), 81-96 (2011).
  40. Barkaoui, K. Effects of marking method and rater experience. Assess Educ Policy Pract. 18 (3), 279-293 (2011).
  41. Chouldechova, A. Fair prediction with disparate impact. arXiv preprint arXiv. , (2017).

Access restricted. Please log in or start a trial to view this content.

Reprints and Permissions

Request permission to reuse the text or figures of this JoVE article

Request Permission

Tags

Explainable AIAutomated Writing EvaluationAI FairnessLearner PerceptionWriting AssessmentThree Level EvaluationStructural Equation ModelingEqualized OddsMultilingual LearnersEducational Technology
Video Coming Soon

Related Articles