$$\rightleftharpoonup{xx}$$
$$\longleftharp{xx}$$,
$$\longrightharp{xx}$$,
The table of materials summarizes the instruments, software, and scales used in the study, including adapted and validated scales for generative AI use, trust in AI, cognitive load, self-perceived academic performance, sampling and survey procedures, and data analysis tools (SPSS and SmartPLS). Moreover, ethical approval was not applicable for this study as it did not involve any experimental manipulation, medical procedures, or interventions that could pose physical or psychological risks to participants. The research used a non-invasive, self-administered online questionnaire to collect data on students' perceptions, experiences, and behaviors regarding the use of generative AI in academic settings. Since the study involved no direct interaction beyond completing a survey, no clinical trials, biomedical procedures, or experimental treatments were administered, and participants were not subjected to any experimental conditions that could cause harm or distress. Prior to data collection, all participants were provided with a detailed informed consent form explaining the purpose of the study, the voluntary nature of their participation, and the measures taken to ensure anonymity and confidentiality. Informed consent was obtained implicitly through the completion and submission of the questionnaire, as explicitly stated in the informed consent form. Participants were assured that their responses would be used solely for academic research purposes and that no personally identifiable information would be collected or reported. Furthermore, participants were informed of their right to withdraw from the study at any time before submitting the questionnaire, without consequences. All data were stored securely on password-protected devices accessible only to the research team, and aggregated findings were reported to prevent the identification of individual respondents. These measures ensured that the study was conducted in accordance with the ethical principles of the Declaration of Helsinki and the institutional guidelines for research involving human participants, despite the formal requirement for IRB review not being applicable.
Research design
This study employs a quantitative, cross-sectional research design to investigate the complex relationships between generative AI usage, trust in AI, cognitive load, and academic performance among higher education students. This design was operationalized through a survey strategy, chosen for its efficiency and effectiveness in collecting standardized data from a large sample, thereby enabling robust statistical testing of the hypothesized relationships. The survey method facilitates the systematic measurement of each construct using validated scales, ensuring that the data collected is both reliable and comparable across all respondents. Furthermore, the adoption of a cross-sectional approach, wherein data is gathered at a single point in time, is particularly appropriate for this investigation as it allows for the examination of the current prevalence of generative AI usage and trust among students, while simultaneously assessing the interconnections between these variables and their collective influence on cognitive load and academic performance within the study's target population.
Participants and sampling procedure
The target population for this study comprised students enrolled in higher education institutions who had active experience using generative AI tools, specifically ChatGPT and similar large language models, for academic purposes. To ensure a representative cross-section of this population and enhance the generalizability of findings, a stratified random sampling technique was employed. The sampling frame was obtained from the university registrar's office, comprising a complete list of currently enrolled students across all faculties. Stratification was based on two key demographic variables: academic discipline and year of study. Academic discipline was categorized into four main strata: humanities and social sciences, natural sciences, engineering and technology, and business and economics. Year of study was stratified into four levels: first-year, second-year, third-year, and fourth-year undergraduate students, as well as postgraduate students (Master's and Doctoral). This dual-stratification approach ensured proportional representation from each subgroup within the student body, preventing over-representation or under-representation of any academic field or academic level. From each stratum, participants were randomly selected using a random-number generator, with the number drawn proportional to that stratum's representation in the overall student population. This systematic approach to participant selection ensured that the final sample accurately reflected the diversity of the higher education student population, including academic backgrounds and educational progression.
Regarding participant demographics, the sample of 390 respondents comprised a diverse mix of students across various age groups, typically ranging from 18 to 35 years or older, with a balanced representation of gender identities proportionate to the institutional demographics. The sample included students from all four undergraduate years as well as postgraduate students, ensuring coverage of different stages of academic development and varying levels of exposure to AI tools. Academic disciplines were represented proportionally, with students from humanities, sciences, engineering, and business backgrounds included in the sample. Additionally, basic demographic information, such as prior technology experience, frequency of generative AI use, and native language status, was collected to provide context for the primary variables under investigation. This comprehensive demographic profiling served two purposes: first, it allowed for descriptive characterization of the sample, and second, it enabled researchers to control for potential confounding variables in subsequent analyses. The detailed documentation of the sampling procedure and demographic characteristics ensures that other researchers can replicate this study precisely, thereby fulfilling the scientific requirement of reproducibility and enabling meaningful comparisons across different institutional and cultural contexts. Furthermore, this rigorous sampling approach strengthened the external validity of the findings, allowing for more confident generalization of results to the broader higher education student population. The data for this study were collected using a structured, self-administered online questionnaire. The questionnaire is divided into five sections, designed to measure the study's core constructs using established and validated scales. All items, unless otherwise stated, will be measured on a five-point Likert scale, ranging from 1 ("Strongly Disagree") to 5 ("Strongly Agree").
Instrumentation and measures
Generative AI usage, specifically ChatGPT, was quantified using an 8-item scale adapted from a study45. This scale moves beyond simple frequency of use to capture the multifaceted nature of student interaction with AI. The items focus on three key dimensions: (a) the frequency of use for various academic tasks, (b) the specific purposes for which it is employed (e.g., brainstorming, content summarization, proofreading), and (c) the students' perceived efficacy of the tool in enhancing their language-related academic work. Trust in AI was conceptualized as a multidimensional construct, drawing from established frameworks of interpersonal trust46, and technology trust. The scale comprises two distinct dimensions. Human-like trust: This 6-item subscale measures the extent to which students attribute human-like qualities to the AI, such as benevolence, integrity, and predictability. For example, items assess whether students believe the AI acts in their best interest or provides unbiased information. Functionality trust: This 5-item subscale assesses trust in the AI's functional performance. It focuses on the reliability, competence, and dependability of the technology itself, capturing the belief that the AI has the functionality required to perform academic tasks effectively.
Cognitive load was measured using a scale adapted from another study47, which is grounded in CLT. To capture the triarchic structure of CLT, the scale is divided into three subscales: Intrinsic Load (3 items): this subscale assesses the perceived complexity and difficulty inherent in the learning materials and tasks themselves. A sample item might be, "The topics covered in my coursework are very complex." Extraneous load (3 items): This subscale measures the cognitive burden imposed by the instructional design and the way information is presented. An example is, "The instructions for assignments are often unclear and difficult to follow. Self-perceived learning (4 items): Following prior educational research, this subscale was used as an approximate indicator of germane cognitive processing rather than a direct measure of germane load itself48. The items capture students’ perceptions of the mental effort invested in understanding, comprehension, and schema construction processes associated with meaningful learning. However, because self-perceived learning may conceptually overlap with broader perceptions of academic capability and achievement, it should be interpreted cautiously as a partial operational representation of productive cognitive engagement rather than a definitive measure of germane load. It includes items related to understanding, comprehension, and the feeling of having mastered the material. Following established precedents in educational research46, academic performance was operationalized using a self-perception scale. A 4-item instrument was employed. This approach is appropriate for capturing students' subjective assessment of their learning efficacy and academic achievement. The scale focuses on perceived academic efficacy and includes statements such as, "I am confident in my ability to complete my coursework successfully," and "I feel I have learned to manage my academic tasks efficiently and effectively."
Data analysis strategy
Partial least squares structural equation modeling (PLS-SEM) was employed to examine the relationships among generative AI usage, trust in AI, cognitive load, and self-perceived academic performance. The selection of PLS-SEM was guided by both methodological and analytical considerations. First, the study adopts a prediction-oriented and exploratory modeling perspective aimed at examining associative pathways among relatively emerging constructs within AI-enhanced learning contexts rather than testing a fully established causal theory. Second, PLS-SEM is considered appropriate for complex models involving multiple latent constructs and mediation pathways, particularly when the primary objective is variance explanation and predictive analysis. In addition, PLS-SEM is robust under conditions of non-normal data distributions and is suitable for social science survey research using moderate sample sizes. Given the exploratory nature of generative AI research in higher education and the study’s emphasis on examining predictive associations among perceptual constructs, PLS-SEM was considered methodologically appropriate for the present analysis.
The collected data will be analyzed using a two-step approach, employing partial least squares structural equation modeling (PLS-SEM) with SmartPLS 4 software. PLS-SEM is a variance-based technique suitable for complex predictive models and does not require the data to be normally distributed. Step 1: Assessment of the measurement model. The first step involves evaluating the reliability and validity of the constructs. Internal consistency reliability will be assessed using Cronbach's alpha and composite reliability, with a threshold of 0.70. Convergent validity will be established by examining the outer loadings of the indicators (should be > 0.70) and the average variance extracted (AVE) for each construct (should be > 0.50). Discriminant validity, which ensures that the constructs are empirically distinct, will be assessed using the Fornell-Larcker criterion and the Heterotrait-Monotrait (HTMT) ratio of correlations. Step 2: Assessment of the structural model. Once the measurement model is confirmed as reliable and valid, the structural model will be evaluated to test the study's hypotheses. This involves examining the path coefficients (β) for significance and relevance using a bootstrapping procedure with 5,000 resamples. The key criteria for assessment will include the coefficient of determination (R2): To measure the model's predictive power (i.e., the variance explained in the endogenous variables, particularly academic performance). Predictive relevance (Q2): Using the blindfolding procedure to assess the model's predictive accuracy. Effect sizes (f2): To evaluate the substantive impact of each independent variable on the dependent variables. Mediation analysis: To test the mediating role of cognitive load in the relationship between the AI-related antecedents (generative AI usage, trust in AI) and academic performance. The significance of the indirect effects will be assessed using bootstrapped confidence intervals.