Although the study did not require formal ethical approval from an institutional review board because of its non-experimental, anonymous survey design involving no physical interventions, manipulation, or deception of participants, all procedures were conducted in accordance with the ethical principles of the Declaration of Helsinki and the American Psychological Association’s guidelines for research involving human participants. Informed consent was obtained from all participants prior to data collection. The institution did not require or issue a formal ethics approval or exemption for this type of study. All respondents were explicitly informed of the research purpose, their voluntary participation, their right to withdraw at any time without consequence, and the strict confidentiality measures applied to their data. No personally identifiable information was collected or reported, and all data were used exclusively for aggregated academic analysis. Particular attention was given to non-intrusive, respectful wording to ensure that the survey did not interfere with the participants’ academic activities or psychological well-being. For the informed consent form, see Supplementary File 1. Moreover, the materials used in this study are listed in the Table of Materials.
Research design
This section outlines the research design, data collection procedures, and sample details used in the study to investigate the relationships among teaching style, ChatGPT usage, digital literacy, motivation, anxiety, and creativity among EFL students. This study employed a quantitative, cross-sectional research design to examine the mediating role of motivation in the relationships among teaching style, ChatGPT usage, digital literacy, and anxiety among EFL students, and how these variables collectively relate to creativity. A structured survey methodology was adopted to collect primary data, ensuring standardization and reliability in measuring the constructs. The study utilized validated scales to assess all variables, including teaching style, ChatGPT usage, digital literacy, motivation, FLCA, and creativity.
Data and Sample
Data for this study were collected from EFL students enrolled at the University of Jeddah, a public-sector university in Saudi Arabia recognized for its emphasis on language education and the integration of technology in teaching. The participants were selected using a purposive sampling technique, which was deemed appropriate for this study given the specific inclusion criteria requiring participants to have active experience with AI tools such as ChatGPT in their learning environment. This non-probability sampling approach was necessary because the study’s research questions focused on a specific population EFL students who had been exposed to both traditional and technology-assisted learning methods and random sampling would not have guaranteed the inclusion of participants meeting these criteria. While purposive sampling limits the generalizability of the findings to the broader population, it was justified for this exploratory investigation to ensure that participants possessed the relevant experiences and perspectives necessary to address the research objectives. To mitigate the limitations of purposive sampling, efforts were made to include participants with diverse academic backgrounds and varying levels of technological proficiency, thereby enhancing the representativeness of the sample within the target population. The final sample consisted of 280 undergraduate EFL students, a size deemed adequate for the statistical techniques employed.
The determination of sample size adequacy was based on established guidelines for Structural Equation Modeling (SEM). Researchers recommend a minimum ratio of 10 cases per estimated parameter for SEM, with more conservative guidelines suggesting 20 cases per parameter for Partial Least Squares SEM (PLS-SEM). The measurement model comprised 29 items for digital literacy, 25 items for teaching style, and 32 items for FLCA, alongside additional items for ChatGPT usage, motivation, and creativity. PLS-SEM guidelines also recommend a minimum sample size equal to 10 times the largest number of formative indicators or the largest number of structural paths directed at a particular construct; therefore, the sample of 280 participants exceeded these minimum requirements. Furthermore, a post hoc power analysis using G*Power53 indicated that a sample of 280 participants provided statistical power exceeding 0.95 for detecting medium effect sizes (f2 = 0.15) with α = 0.05, exceeding the commonly recommended threshold of 0.80. This sample size provides reliable parameter estimates and sufficient statistical power to detect meaningful relationships among the study variables while accommodating the complexity of the proposed measurement and structural models.
Participants and Procedure
The participants were EFL students aged between 18 and 25 years, with a majority being first- and second-year undergraduates enrolled in English language courses. The survey was conducted between September and October 2025, during which students were invited to participate voluntarily. The recruitment process involved outreach through official university channels, including emails and announcements in classes. The participants were EFL students aged 18–25 years who were enrolled as first- or second-year undergraduates in English language courses. Eligibility was determined using predefined inclusion and exclusion criteria. Participants were eligible if they were (a) officially enrolled in undergraduate English language courses at the University of Jeddah, (b) aged 18–25 years, and (c) registered as first- or second-year students. Students were excluded if they were (a) graduate students or upper-division undergraduates outside the foundational course sequence, (b) outside the specified age range, or (c) involved in the questionnaire pretesting phase. To ensure informed consent, participants were briefed on the purpose of the study, the voluntary nature of their participation, and the confidentiality of their responses. The data collection instrument was a structured, self-administered questionnaire distributed physically and electronically using a secure online platform. Validated scales were adapted for measuring teaching style, ChatGPT usage, digital literacy, students’ motivation, FLCA, and creativity. Participants were asked to rate their responses on a 5-point Likert scale, ranging from 1 (Strongly Disagree) to 5 (Strongly Agree). Pre-testing of the questionnaire was conducted with a small group of 30 students to ensure clarity and reliability. Based on the pre-testing results, minor wording adjustments were made to enhance item comprehensibility, and no significant issues with scale reliability or validity were identified. To maintain data integrity and encourage honest responses, anonymity was guaranteed, and participants were assured that their data would only be used for research purposes. Out of the initial 300 responses collected, 280 were deemed valid after screening for incomplete or inconsistent data, yielding a retention rate of 93%. Data screening procedures included examining response patterns for straight-lining, excessive missing data (more than 10%), and multivariate outliers using Mahalanobis distance. This methodological approach provides a robust foundation for analyzing the relationships among the study variables and offers valuable insights into the role of teaching style, AI tools, and digital literacy in shaping EFL students’ motivation, anxiety, and creativity.
Following data collection, 20 responses were excluded from the initial pool of 300 based on predefined criteria: 12 responses were removed due to excessive missing data exceeding 10% of questionnaire items, 5 were eliminated for demonstrating straight-lining response patterns (uniform responses across all items indicating non-engagement), and 3 were identified as multivariate outliers using Mahalanobis distance (p < 0.001) and subsequently removed to prevent distortion of the SEM results. Missing values for retained cases were handled using mean imputation where missing data were minimal (less than 5% per case), as this approach is acceptable for small proportions of missingness in large-scale survey research. The final retention rate of 93% (280 valid responses out of 300 collected questionnaires) represents the proportion of usable responses retained from the total collected questionnaires; however, this should not be interpreted as a true response rate, as the exact number of invited participants could not be calculated because the combined physical and electronic distribution methods through official university channels precluded precise tracking of total invitations. This limitation is acknowledged, and the 93% figure is therefore reported as the proportion of valid responses retained from the collected questionnaires rather than as a response rate relative to the total invited population.
Tools and Measures
This study utilized well-established and validated scales, carefully adapted to the context of EFL education, to measure the core constructs: teaching style, digital literacy, ChatGPT usage, FLCA, motivation, and creativity. Each scale was selected based on its relevance and comprehensiveness in capturing the respective dimensions of the variables under investigation. Digital literacy was measured using an adapted scale based on Rodríguez-de-Dios et al.54, comprising six dimensions with 29 items. These dimensions included Technological Skill, Personal Security Skill, Critical Skill, Device Security Skill, Informational Skill, and Communication Skill, allowing for a comprehensive assessment of students’ proficiency in digital skills essential for modern educational contexts. The usage of ChatGPT was evaluated using an 8-item scale adapted from Abbas et al.55, focusing on students’ interaction with the AI tool for language learning tasks. Teaching style was assessed using an adapted version of Grasha’s Teaching Style Inventory56. This instrument measures five dimensions of teaching styles: Authority, Expert, Personal Model, Facilitator, and Delegator. These dimensions reflect various instructional approaches that influence EFL students’ learning experiences. FLCA was measured using an adapted scale based on Briesmaster and Briesmaster57. This instrument examined anxiety across three dimensions: Communication Apprehension (10 items), Test Anxiety (10 items), and Fear of Negative Evaluation (7 items). Students’ motivation was assessed using a scale adapted from Kanoksilapatham et al.58, measuring three dimensions: Instrumentality Promotion and Prevention, Ethnocentrism and Integrativeness, and Attitude Towards Learning English. The scale included a total of 20 items. EFL students’ creativity was measured using a scale adapted from Govindasamy et al.59, which assessed four dimensions: Originality, Flexibility, Fluency, and Elaboration. Each dimension was measured using three items. To ensure the cultural and contextual suitability of the instruments for the EFL setting, a systematic adaptation procedure was applied. Because the original scales were developed in English, two bilingual EFL experts reviewed and simplified the language to improve clarity for intermediate-level undergraduate students while preserving the original meaning of each construct. The adapted versions were then back-translated into the participants’ first language by an independent translator to verify linguistic and conceptual equivalence, and any discrepancies were resolved through discussion among the research team. Subsequently, three experienced EFL instructors evaluated the adapted instruments for content relevance, cultural appropriateness, and item clarity, resulting in minor wording revisions and the replacement of culturally unfamiliar examples where appropriate. Finally, the revised instruments were pilot-tested with 30 students (see Procedure section), and participant feedback was used to make additional minor wording refinements to improve comprehensibility. This multi-step process ensured that the adapted instruments maintained their original psychometric properties while being appropriate for the target EFL population. For the complete adapted questionnaire, see Supplementary File 2.
Data Analysis Methods
The data collected for this study were analyzed using the Statistical Package for the Social Sciences (SPSS 26) and SmartPLS 4, employing PLS-SEM. This dual-method approach ensured a comprehensive analysis of the data, addressing both descriptive and inferential aspects while testing the hypothesized relationships and mediation effects in the research model. Initially, SPSS was used to perform preliminary data analysis, including data cleaning, descriptive statistics, and reliability testing. Data screening involved checking for missing values, outliers, and normality to ensure the quality and integrity of the dataset. Descriptive statistics, including the mean, standard deviation, and frequency distributions, were computed to summarize the sample characteristics. The main analysis was conducted using SmartPLS 4. PLS-SEM was selected because of its suitability for analyzing complex models with multiple constructs and its robustness for small to medium sample sizes. The analysis was conducted in two stages: evaluation of the measurement model and evaluation of the structural model. In the measurement model assessment, indicator reliability, internal consistency reliability (using composite reliability [CR]), convergent validity (using average variance extracted [AVE]), and discriminant validity (using the heterotrait–monotrait ratio [HTMT]) were evaluated.
Following validation of the measurement model, the structural model was assessed to evaluate the hypothesized relationships among teaching style, ChatGPT usage, digital literacy, motivation, anxiety, and creativity. Path coefficients were analyzed to determine the strength and significance of the direct and indirect effects, with significance assessed using bootstrapping with 5,000 resamples. To assess the significance of direct, indirect, and mediation effects, bias-corrected (BC) bootstrap confidence intervals were constructed using 5,000 bootstrap resamples with a 95% confidence level. The mediation effects of motivation on the relationships between teaching style, ChatGPT usage, digital literacy, and anxiety were also evaluated using bootstrapped confidence intervals. Model fit and predictive power were assessed using the coefficient of determination (R2), predictive relevance (Q2), and standardized root mean square residual (SRMR).