Ethics statement
This study was conducted in accordance with the ethical standards of the Institutional Review Board (IRB) of Chengdu University, China (Approval No: 202509). All participants provided written informed consent prior to participation. The study adhered to institutional guidelines for research involving human subjects, ensuring voluntary participation, confidentiality, and the right to withdraw at any time without penalty.
To maintain confidentiality, all collected data were anonymized using coded identifiers, and personal information was stored separately from research data in password-protected files accessible only to the principal investigators. The questionnaire design incorporated measures to minimize psychological discomfort, particularly for questions about anxiety, and participants were provided with contact information for university counselling services should any distress arise from survey participation. To address potential power imbalances in the teacher-student relationship, recruitment and data collection were conducted by trained research assistants rather than course instructors, and participants were assured that their responses would not affect their academic standing or relationship with faculty.
This study employed a quantitative, cross-sectional design involving 290 Chinese undergraduate EFL students from a public university. Participants representing diverse academic years and disciplines were recruited through convenience sampling with proportional representation and completed paper-based questionnaires in controlled classroom settings during regular sessions, requiring approximately 25 min. The questionnaire battery included validated scales measuring teaching style, digital literacy, ChatGPT usage, motivation, anxiety, and creativity. To minimize bias, trained research assistants administered the questionnaires using procedural countermeasures such as psychological separation of items. Following data collection, the responses were analyzed using statistical software and Partial Least Squares Structural Equation Modeling (PLS-SEM) software, including data screening, measurement model evaluation, structural model testing, and mediation analysis using a bootstrapping procedure with 5,000 resamples.
Research Design
This study employed a quantitative, cross-sectional research design to examine the relationships between teaching style, ChatGPT usage, digital literacy, motivation, anxiety, and creativity among EFL learners. A cross-sectional approach was deemed appropriate for capturing a snapshot of these variables at a specific point in time, allowing for the analysis of their interrelationships without the temporal demands of longitudinal research. The design facilitated the testing of a hypothesized mediation model in which motivation serves as the mechanism linking teaching style, ChatGPT usage, and digital literacy to reduced anxiety and enhanced creativity.
Data and Sample
The sampling frame consisted of all undergraduate students enrolled in English as a Foreign Language (EFL) courses at a public university in China during the 2024–2025 academic year. Eligibility criteria required participants to be (a) native Chinese speakers, (b) currently enrolled in at least one required EFL course, and (c) without prior formal training in AI-based language learning tools beyond general coursework. A stratified convenience sampling approach was used, with proportional allocation across academic years (freshman, sophomore, junior, senior) and disciplinary areas (humanities, social sciences, natural sciences, engineering). Target strata sizes were set at 25% per academic year and 25% per disciplinary area, based on the university’s overall student distribution. The final sample comprised 290 undergraduate EFL students. Of the 328 students approached during regular class sessions, 11 declined to participate (non-participation rate = 3.4%) and 27 provided incomplete questionnaires (incomplete response rate = 8.2%), yielding a final usable response rate of 88.4% (290/328). The a priori sample‑size calculation was performed using G*Power 3.1 with the F tests family, linear multiple regression: fixed model, R2 deviation from zero as the test. Input parameters were: effect size f2 = 0.15 (conventional medium effect size for behavioural research), α err prob = 0.05, power (1‑β err prob) = 0.95, and number of predictors = 3 (teaching style, ChatGPT use, digital literacy) for the motivation outcome. The calculation yielded a minimum required sample size of 119 for detecting a significant R2. To further account for the full SEM model (including indirect paths and additional endogenous variables), we applied the more conservative rule of 10 cases per estimated parameter. With 29 estimated parameters in our model, the required sample size was 290, which exactly matches the final sample. Thus, the obtained sample is fully powered for both regression‑based and SEM analyses.
Participants and Procedure
Participants were recruited through stratified convenience sampling to ensure proportionality across academic years (freshman to senior) and disciplinary backgrounds (humanities, sciences, and engineering). Strata sizes in the final sample were: freshmen (n = 72, 24.8%), sophomores (n = 74, 25.5%), juniors (n = 71, 24.5%), seniors (n = 73, 25.2%); and humanities (n = 76, 26.2%), social sciences (n = 68, 23.4%), natural sciences (n = 72, 24.8%), engineering (n = 74, 25.5%). The sample comprised 62% female and 38% male students, with an average age of 20.3 years (SD = 1.7). Data collection occurred during regular EFL class sessions. Trained research assistants administered paper‑based questionnaires in controlled classroom settings to ensure standardization. Prior to survey administration, all eligible students present in class were informed of the study’s purpose, voluntary nature, and confidentiality measures (anonymous responses, secure data storage). Students who did not meet eligibility criteria (n = 6, no EFL enrolment) or declined participation (n = 11) were excluded; reasons for declination included time constraints (n = 7) and discomfort with self‑reported technology use (n = 4). The questionnaire battery took approximately 25 minutes to complete and included validated scales for all constructs. To mitigate common method bias, procedural countermeasures were applied (e.g., psychological separation of scale items, anonymity of responses). In addition, common method bias was assessed using Harman’s single-factor test, and the first factor accounted for less than 50% of the total variance, indicating that common method bias was not a major concern in this study.
Study Measures
All scales were adapted to the Chinese EFL context through a rigorous translation and adaptation procedure. Two bilingual researchers independently translated the original English items into Chinese, and a third researcher reconciled discrepancies to produce a preliminary version. This version was back‑translated into English by two different translators, and any inconsistencies were resolved through discussion with the research team. The Chinese versions were then pilot‑tested with 30 undergraduate EFL students (not part of the main sample) to check clarity and comprehension, resulting in minor wording adjustments. For each scale, responses were collected on a 5‑point Likert scale ranging from 1 (strongly disagree) to 5 (strongly agree). Summed or averaged scores were used for each construct, with higher scores indicating higher levels of the respective variable. All adapted scales were used with permission from the original authors or under open‑license terms where applicable. No items were removed from the original scales except where noted below; item retention was based on factor loadings > 0.40 and theoretical relevance. Regarding ChatGPT usage, the study assessed students’ self‑reported interaction with the free version of ChatGPT (GPT‑3.5/4 available via web or mobile app during the 2024–2025 academic year). All participants had equal access to the tool as the university did not restrict ChatGPT; however, actual usage varied individually. The scale measured the frequency of ChatGPT use (ranging from “never” to “daily”), the types of language learning tasks for which it was employed (e.g., vocabulary building, grammar checking, essay drafting, conversation practice, and error correction), and the perceived quality of interaction (e.g., usefulness, clarity of feedback). The scale did not distinguish between specific ChatGPT versions, as version updates occurred during the data collection period; instead, participants reported the version they primarily used, with over 90% reporting GPT‑3.5 or GPT‑4. Detailed item examples and adaptation notes for each construct are provided below.
This study employed rigorously validated scales, thoughtfully adapted to the EFL education context, to assess the key variables: teaching approaches, digital competence, ChatGPT integration, language learning anxiety, learner motivation, and creative thinking. The measurement tools were chosen for their demonstrated reliability and ability to comprehensively evaluate each construct. To examine instructional methods, the research utilized a modified version of Grasha's Teaching Style Inventory34 which evaluates five distinct pedagogical approaches: authoritative, expert, Personal model Teacher, facilitative, and delegatory. To assess digital Literacy, the study incorporated a modified version of the scale developed by Rodríguez-de-Dueños et al.35, consisting of 29 items across six key domains: (1) technical literacy, (2) personal data protection literacy, (3) critical evaluation literacy, (4) device security literacy, (5) information processing literacy, and (6) digital communication literacy. This multidimensional framework provided a thorough evaluation of learners' digital capabilities necessary for contemporary education. The usage of ChatGPT was evaluated using an 8-item scale adapted from Abbas et al.36, focusing on students’ interaction with the AI tool for language learning tasks. This scale captures students perceived and self-reported use of ChatGPT in language learning tasks rather than objective usage frequency. Foreign language classroom anxiety (FLCA) was measured using an adapted scale based on Briesmaster and Briesmaster37. This instrument examined anxiety across three critical dimensions: Communication Apprehension, Test Anxiety, and Fear of Negative Evaluation, with 10, 10, and 7 items, respectively. Students’ motivation was assessed using a scale adapted from Kanoksilapatham et al.38, measuring three key dimensions: Instrumentality Promotion and Prevention, Ethnocentrism and Integrativeness, and Attitude Towards Learning English. The scale included a total of 20 items. EFL students’ creativity was measured using a scale adapted from Govindasamy et al.39, which assessed four dimensions: Originality, Flexibility, Fluency, and Elaboration. Each dimension was measured using three items.
Data Analysis Methods
The collected data were analysed using a two-stage analytical approach combining SPSS (Version 27) and Partial Least Squares Structural Equation Modeling (PLS-SEM) through SmartPL software. Initial data screening and preliminary analyses were conducted using SPSS to examine data quality, including checks for missing values, outliers, and normality assumptions. Less than 2% of data points were missing completely at random (MCAR), as confirmed by Little's MCAR test (χ2 = 18.24, p = 0.212), and were handled using the expectation-maximization algorithm to preserve statistical power. The dataset demonstrated acceptable univariate normality (skewness < |2|, kurtosis < |7|) and multivariate normality (Mardia's coefficient = 18.37, p < 0.001), justifying the use of both parametric tests and PLS-SEM, which is robust to non-normality. Although univariate normality was acceptable, Mardia’s coefficient indicated a violation of multivariate normality (p < 0.001). However, this does not pose a concern as PLS-SEM is robust to non-normal data distributions.
Reliability analysis revealed strong internal consistency for all scales (Cronbach's α > 0.82, composite reliability > 0.85), while exploratory factor analysis confirmed the anticipated factor structures with all items loading cleanly on their respective constructs (loadings > 0.65, cross-loadings < 0.40).
For hypothesis testing, PLS-SEM was selected over covariance-based SEM due to its superior handling of complex models with small-to-medium sample sizes and its ability to maximize explained variance in endogenous constructs. The analysis followed the two-step procedure, first evaluating the measurement model and then assessing the structural model. The reflective measurement model demonstrated adequate convergent validity (average variance extracted > 0.50 for all constructs) and discriminant validity as confirmed by both Fornell-Larcker criterion and heterotrait-monotrait ratio (HTMT < 0.85). In the structural model, path coefficients were estimated using a bootstrapping procedure with 5,000 resamples to generate stable estimates and confidence intervals. The model's predictive power was evaluated through R2 values for endogenous variables (motivation = 0.53, anxiety = 0.41, creativity = 0.38) and predictive relevance (Q2 > 0), with effect sizes (f2) calculated to determine substantive impact. Mediation analyses employed the bias-corrected bootstrap method to test indirect effects, with significant mediation established when 95% confidence intervals excluded zero.
All constructs were modelled as reflective (mode A) based on theoretical considerations, as each latent variable was assumed to cause its observed indicators. No items were dropped from any scale, as all items demonstrated outer loadings above the 0.50 threshold (ranging from 0.62 to 0.89) and all average variance extracted (AVE) values exceeded 0.50, confirming convergent validity. For mediation testing, the structural model included direct paths from the three antecedents (teaching style, ChatGPT use, digital literacy) to motivation, from motivation to anxiety, and from anxiety to creativity. Direct paths from antecedents to anxiety and from antecedents to creativity were also estimated to test for partial vs. full mediation. Full mediation was determined based on the criteria that (a) the indirect effects (antecedents → motivation → anxiety) were significant (95% bias-corrected bootstrap confidence interval excluding zero), and (b) the direct effects from antecedents to anxiety became non-significant after including the mediator, while the direct effects from antecedents to creativity were never significant. All analyses used the PLS-SEM algorithm in SmartPLS 4 with path weighting scheme, maximum iterations set to 300 and stop criterion of 1.0e-7. Bootstrapping was performed with 5,000 subsamples, bias-corrected and accelerated (BCa) confidence intervals at 95%, and no sign changes option. Model fit was assessed using the standardized root mean square residual (SRMR = 0.061), which falls below the recommended 0.08 threshold.