All methods involving human subjects were conducted in compliance with institutional guidelines and were approved by the institutional review board (IRB) or the human research ethics committee of Hanshan Normal University. Informed consent was obtained from all participants prior to the experiment.
Conceptual framework and research design
This study employed a process–product research design with an auxiliary discourse layer to examine how audience-generated danmu comments influence subtitle translation. The analytical framework integrated textual characteristics, indicators of the translation process, and translation quality outcomes. Specifically, it investigated whether audience attention, as reflected by danmu density, directed translators' cognitive resources toward particular textual features and whether this attentional allocation subsequently affected translation quality.
To operationalize this framework, the study first established segment-level textual features based on four dimensions: cultural load, evaluative density, narrative compression, and danmu salience. Cultural load referred to the presence of culturally specific references that required adaptation or explanation. Evaluative density measured the concentration of emotionally or attitudinally charged language within a segment. Narrative compression captured the extent to which semantic content was condensed under subtitle constraints. Danmu salience represented the degree of audience attention directed toward each subtitle segment, quantified by the volume of synchronized danmu comments. Unlike conventional measures of translation difficulty, danmu salience reflected socially distributed attention rather than intrinsic textual complexity. A segment receiving extensive audience engagement was not necessarily more difficult to translate; instead, it represented a point of heightened communicative significance within the viewing experience.
The analytical framework comprised three interconnected datasets. The participant-level dataset included individual demographic characteristics and translation expertise. The process–product dataset linked each subtitle segment to behavioral indicators extracted from keystroke logging, together with expert translation quality ratings. The danmu dataset contained synchronized audience comments aligned with subtitle timing, from which danmu salience values were computed. These datasets were linked via participant and segment identifiers, enabling simultaneous analysis of each translated subtitle segment in relation to translator characteristics, cognitive processing behavior, audience attention, and translation outcomes. The final integrated database comprised 216 complete translation sessions and 3,168 participant–text–segment observations, providing the empirical foundation for the subsequent process–product analyses.
Participant screening and background profiling
The sample comprised 72 English majors recruited from a public university in southern China, evenly distributed across the second, third, and fourth years of undergraduate study (24 participants per year group). All participants were native Chinese speakers receiving formal instruction in English writing and translation. To ensure comparability within the trainee-translator population, individuals who had resided in an English-speaking country for more than 6 months or held professional translator certification were excluded. Participants were not stratified into separate experimental groups according to translation experience, cultural familiarity, or second-language proficiency. Instead, year level, recent English proficiency test scores, prior translation coursework, self-rated familiarity with China-related cultural topics, and average weekly engagement in bilingual practice were recorded and included as participant-level covariates in the statistical models. This clarification also resolves the participant-number inconsistency raised during peer review: both the manuscript and the replication dataset include 72 participants, 216 task sessions, and 3,168 segment-level observations.
Source text selection and task construction
The translation materials consisted of three contemporary China-focused narrative texts developed specifically for this study to reflect authentic public discourse and avoid reproducing copyrighted materials. Each text contained 205–225 Chinese characters but differed systematically in its internal characteristics. Text A presents a Spring Festival homecoming narrative with moderate cultural load and relatively straightforward syntax. Text B recounts a Dragon Boat Festival memory and contains a higher density of culture-specific expressions and implicit background knowledge. Text C describes a rural livestreaming and e-commerce narrative characterized by dense evaluative language and compressed narrative positioning.
The texts were systematically segmented using clause boundaries, followed by meaning-unit boundaries, yielding 44 analyzable units across the three translation tasks (Table 1). Before formal coding, a coding manual defined cultural load, evaluative density, and narrative compression using 1–5 ordinal scales. A pilot study involving 12 students who did not participate in the main experiment confirmed the intended difficulty calibration, the clarity of the task instructions, and the stability of the segment boundaries for keystroke alignment.
Auxiliary corpus construction and salience mapping
To identify publicly salient narrative cues, an auxiliary danmu corpus was compiled from 24 Bilibili videos uploaded between January 2023 and December 2025. Videos were included only if they matched one of the three semantic domains represented in the source texts, had accumulated more than 50,000 views, contained dense danmu interaction, and permitted the export and thematic mapping of time-synchronized comments. After duplicate removal and filtering of emojis, onomatopoeic fillers, and irrelevant fragments, the final corpus comprised 18,642 comments.
Content words were standardized and clustered using source-text-related keyword anchors. Each source-text segment was then assigned a continuous danmu salience score based on a weighted composite of normalized keyword-cluster frequency, comment concentration within the corresponding thematic cluster, and mean evaluative intensity. This procedure enables comparisons between translator processing effort and audience salience at the semantic-cluster level rather than through direct lexical matching. Because danmu reflects audience responses to online videos rather than the experimental translation task itself, danmu salience is interpreted as an external indicator of audience engagement rather than direct evidence of the linguistic difficulty of a source segment.
Experimental procedure and data capture
Translation sessions were conducted individually in a controlled computer laboratory under standardized hardware, software, and internet-access conditions (Figure 2). Following a five-minute orientation emphasizing the production of readable English for an international audience, participants completed a three-minute English typing warm-up to minimize the effects of keyboard acclimatization. The three translation tasks were administered sequentially using Translog-II, with the task order counterbalanced in a Latin-square design to distribute potential fatigue and order effects evenly across participants. Participants were allocated a maximum of 20 minutes per text, separated by five-minute rest intervals.
Although the built-in bilingual dictionary was permitted, machine translation systems and external search engines were prohibited. All keystrokes, pauses, and cursor movements were logged automatically and supplemented by synchronous screen recordings to verify non-linear revision episodes. After completing each task, participants completed a self-reported difficulty assessment used solely to assist in interpreting borderline coding cases.
Process measurement and variable coding
Keystroke records were aligned with the predefined segment structure. An alignment episode began with the first target-language keystroke corresponding to a segment and ended when the participant definitively progressed to the next segment. Overlapping revisions were reassigned using time-stamped cursor-tracking records.
Decision-making behavior was quantified using five indicators: first-pause duration before initial segment translation, mean within-segment pause duration, the frequency of pauses exceeding two seconds, dictionary consultation counts, and cursor-based interruption counts indicating departures from linear drafting (Table 2). Revision behavior was classified along two dimensions: timing and function. Timing distinguished immediate local revisions from delayed revisions performed after drafting had progressed beyond the current segment. Functional coding classified revisions as surface revisions (form-level corrections without semantic change), meaning revisions (lexical or syntactic modifications), discourse-level revisions (adjustments to clause relationships or coherence), and cultural adaptation revisions (target-audience-oriented reformulations). Two trained coders independently analyzed each revision episode before adjudication and assigned a single primary functional category to each event. Inter-rater reliability was robust across all coded measures. Cohen’s ĸ indicated substantial agreement for the categorical coding of revision functions (ĸ = 0.84, p < 0.001). For the ordinal text-level variables, intraclass correlation coefficients (ICC, two-way mixed-effects model, absolute agreement) demonstrated high reliability for cultural load (ICC = 0.86), evaluative density (ICC = 0.78), and narrative compression (ICC = 0.82). Any initial coding discrepancies were subsequently resolved through joint adjudication.
Translation quality assessment
Translation quality was evaluated independently of the process data by two raters with doctoral-level training and extensive experience in translation instruction. Blinded to participant identities and process measures, the raters applied a 100-point rubric, evenly distributed across five dimensions: semantic accuracy, linguistic fluency, narrative coherence, cultural rendering, and register appropriateness (Table 3). Score discrepancies greater than eight points triggered joint review and consensus rescoring, with the adjudicated average used as the final translation quality score. The pre-adjudication rating records were retained to facilitate transparent reporting of inter-rater reliability in the replication materials.
Statistical modeling
The analyses proceeded in three sequential stages. First, descriptive statistics were computed to summarize all variables across year groups and text types. Second, segment-level processing effort was analyzed using mixed-effects regression models to account for the nested structure of segments within texts and repeated observations from individual participants. Dependent variables, including first-pause duration and revision frequency, were modeled as functions of cultural load, evaluative density, narrative compression, and danmu salience while controlling for participant proficiency. Predictor correlations, variance inflation diagnostics, 95% confidence intervals, effect-size estimates, and model-fit indices were included to improve statistical transparency.
Finally, translation quality was modeled to evaluate the predictive contributions of revision timing and revision function relative to overall revision frequency. The overlap between danmu-salient segments and high-effort translation segments was assessed using the upper quintile of each distribution, a standardized composite effort index, and a chance-overlap analysis. No exploratory factor analysis or structural equation modeling was performed. Accordingly, the accompanying figures are presented solely as conceptual, descriptive, or regression-based summaries.