The protocol presented here summarizes the main steps required to obtain reliable eye-tracking data in reading research. Although many of these steps may be well known to experienced researchers, they are often scattered across separate methodological papers, vendor documentation, and informal lab practices. By organizing them into a single workflow from lab preparation to preprocessing, we aim to provide a practical guide for researchers who are new to eye-tracking studies of reading. Compared to general-purpose eye-tracking guidelines or resources mostly focusing on eye-movement detection algorithms, this protocol provides an end-to-end framework as a response to reading research demanding data recording with high spatial resolution and sampling frequency.
A central theme across stages of the workflow is the importance of proactive data-quality management, starting with environmental control and equipment setup. Small issues such as reflections on the monitor, chair instability, or inconsistent viewing distances can produce systematic biases in fixation location or precision. Addressing these factors early minimizes the need for heavy post-hoc correction and supports more consistent validation accuracy (mean error ≤ 0.5–1°).
The protocol also emphasizes the importance of the alignment between research questions and hardware characteristics. Many studies in reading rely on precise spatial and temporal measurement of fixations, saccade landing positions, and regressions1,2. Choosing an eye tracker with an appropriate sampling rate, precision, and accuracy is therefore essential, especially for measures such as landing position or parafoveal preview that are sensitive to small spatial errors. Depending on the specific study constraints, these methods can be modified. For example, when working with neurodivergent populations who cannot tolerate a chinrest, researchers might shift to remote, head-free trackers and compensate for the loss of spatial precision by utilizing larger fonts, wider line spacing, and broader Interest Areas. When head-free systems are used, researchers must account for increased movement and potential data loss. Although remote trackers support more naturalistic setups, head stabilization still remains preferable for most reading tasks today.
Experiment design is another key factor that affects both data quality and interpretability. Standardized on-screen instructions, trial randomization, and appropriate block structure help maintain stable participant behavior across the session. Including drift checks at regular intervals is especially important in reading experiments, as even small gradual offsets can accumulate across lines of text. The protocol also supports the use of language-dependent lexical norms and transparent reporting of independent and dependent variables, which facilitates replication and comparison across studies.
During data collection, continuous monitoring of accuracy, posture, and participant comfort helps prevent large-scale data loss. Active troubleshooting is often necessary for common problems. For instance, calibration may fail at screen corners, then adjusting the monitor or chair height can resolve peripheral tracking loss. Issues from reflective glasses or contact lenses can sometimes be resolved by adjusting ambient lighting or changing monitor angles. Maintaining a detailed session log supports later troubleshooting and strengthens transparency by documenting interruptions, recalibrations, and unexpected events. This is increasingly important because many published studies now require explicit reporting of quality metrics, including precision and percentage data loss.
The preprocessing section emphasizes consistent and well-documented choices in event detection, drift correction, and cleaning routines. Because fixation classification algorithms and thresholds can strongly influence the resulting measures, researchers should specify all parameter settings and justify them when they differ from common defaults. Visual inspection remains a critical step, even when automated tools are applied. The protocol also encourages reporting standard quality indicators such as validation accuracy, precision, data loss, and drift behavior, so that studies can be compared more reliably.
Moreover, the protocol recommends simple sanity checks of reading behavior, such as effects of word length, frequency, or predictability. Reproducing these canonical patterns provides an additional safeguard against unnoticed alignment issues, problematic drift, or preprocessing errors. These checks do not replace full statistical analysis but offer quick feedback on whether the data behaves as expected for skilled reading.
In addition to the details of the experimental protocol already mentioned, it is important to describe the linguistic and psychometric characteristics of the participants. The main rationale for this is the challenge experimental researchers face regarding the comparability of different studies. This is particularly important in the age of open science, where publicly available datasets may be used by researchers who did not originally collect the data and may therefore be unaware of potential idiosyncrasies arising from variations in participants’ (socio)linguistic backgrounds.
Regarding linguistic background, the mother tongue of the participants should be noted, and information about the dialects they use should be collected to account for any dialectal differences that may impact the results. If participants speak more than one language – a common situation in today’s world – second language proficiency should be recorded. In studies of bilingualism or second language processing, second language proficiency should be assessed using validated second language placement tests. Additionally, relevant aspects of socio-economic status may be recorded (e.g., participants’ reading habits).
Tests should be conducted to assess psychometric characteristics of the participants, such as working memory span and cognitive flexibility (especially in bilingual studies). From a methodological perspective, one should consider which tests to select, their cognitive and time demands, and when to administer these measurements within the experimental protocol – whether before, after the eye-tracking session, or in a separate session.
It is essential to remember that in reading studies, language behavior is measured and explained using both the linguistic characteristics of the stimulus and the psychometric characteristics of the readers (or any other features relevant to the linguistic behavior of the participants)11,12,17,20. To enable comparison across studies, all factors that may be relevant for data interpretation must be included. The choice of which characteristics to include will depend on the purpose of the study. The data must be comparable. In summary, data should be collected and described in sufficient detail to allow any potential user the greatest possible flexibility in its use. These considerations may not be part of the protocol itself, but they are relevant for every individual study. In today’s global scientific environment, such details increase the future potential use of the collected data.
This protocol has limitations, most notably the trade-off between data quality and ecological validity, as reading on a static monitor with a restrained head does not fully capture naturalistic reading behavior, and strict visual inclusion criteria can limit sample representativeness. Moreover, eye trackers differ substantially in coordinate scaling, timestamp precision, and event classification. A limitation of the current protocol is that it does not clarify how the variability affects reproducibility across laboratories or hardware brands. Creating guidance for cross-device validation, and for realistic future applications, such as adapting laboratory paradigms into immersive Virtual Reality or Augmented Reality reading environments, and developing digital biomarkers for cognitive or attentional disorders, would fill an important gap.
Overall, the protocol consolidates existing methodological guidance into a clear and transparent workflow for collecting high-quality eye-tracking data in reading. By following the steps and reporting the associated quality metrics, researchers can improve the reproducibility and credibility of reading studies and help establish stronger methodological consistency across the field.