Overall profile of the translation sessions
All 72 participants completed the three translation tasks under standardized experimental conditions, yielding 216 complete translation sessions. Following segment alignment, the final analytical dataset comprised 3,168 participant-text-segment observations. The mean task completion time across the sample was 998.4 s per text, with clear differences across year groups and source texts (Table 4; Figure 3). Accordingly, both the manuscript and the accompanying data consistently support a sample size of n = 72 participants; any figures summarizing subsets should explicitly identify those subsets.
Fourth-year students completed the translation tasks more quickly and with fewer interruptions than second-year students. This difference extended beyond processing speed. The fourth-year cohort also exhibited shorter mean first-pause durations, fewer long pauses, and a lower total number of revisions, while achieving the highest translation quality scores. In contrast, second-year students produced the most revisions, most of which were surface-level modifications and immediate local corrections. Advanced students demonstrated a higher proportion of delayed revisions and cultural adaptation revisions, indicating that superior translation performance was associated with selective revision strategies rather than greater revision volume.
Differences among the three source texts were similarly pronounced. Text A produced the shortest completion times, the fewest long pauses, and the lowest revision density. Texts B and C both imposed greater processing demands, although through different mechanisms. Text B elicited the highest number of culture-related revision events, whereas Text C produced the highest frequency of discourse-level revisions and the longest mean within-segment pause duration. Thus, the source of processing difficulty differed across texts. In Text B, difficulty clustered around culture-bound references and implicit background assumptions, whereas in Text C, it was primarily associated with narrative compression and evaluative positioning.
A one-way analysis of variance confirmed significant differences among year groups in total translation quality, F (2, 69) = 31.84, p < 0.001, with a large effect size (η2 ≈ 0.48 based on participant-level mean scores). Post hoc comparisons showed that fourth-year students achieved significantly higher scores than third-year students, who in turn outperformed second-year students. The same pattern was observed for first-pause duration, long-pause frequency, and the delayed revision ratio. Collectively, these baseline findings establish two important points for the subsequent analyses. First, the three source texts differed in processing difficulty rather than merely in topical content. Second, higher translation quality was associated with a reorganization of translation processes rather than with a simple reduction in processing effort.
Effects of textual salience on decision-making and revision behavior
To determine whether segment-level textual properties predicted processing effort, mixed-effects models were estimated for first-pause duration, long-pause frequency, total revision density, and delayed revision density. Cultural load, evaluative density, narrative compression, and danmu salience were entered as fixed effects, with participant and segment included as random intercepts (Table 5; Figure 4). Predictor screening confirmed the absence of problematic multicollinearity, with all Variance Inflation Factor (VIF) values well below the conservative threshold of 2.5 (Maximum VIF = 1.45). To evaluate model fit and the proportion of variance explained, both marginal and conditional R2 values were calculated for all mixed-effects models (marginal R2 ranging from 0.08 to 0.19; conditional R2 ranging from 0.46 to 0.64). The complete variance components for random effects (participant and segment intercepts) and corresponding model-fit indices are detailed in Table 5. Coefficient estimates are interpreted together with their 95% confidence intervals rather than p-values alone.
Cultural load exhibited the strongest and most consistent effect across the models. In the first-pause model, each one-unit increase in cultural-load score was associated with a 0.91 s increase in first-pause duration (b = 0.91, SE = 0.12, 95% CI [0.67, 1.15], p < 0.001). Cultural load also significantly increased long-pause frequency (b = 0.19, SE = 0.04, 95% CI [0.11, 0.27], p < .001) and total revision density (b = 0.27, SE = 0.06, 95% CI [0.15, 0.39], p < .001). Thus, culture-loaded segments not only delayed initial processing but also increased the likelihood of subsequent reformulation.
Narrative compression demonstrated an equally robust but distinct pattern of effects. It was a strong predictor of discourse-level revision (b = 0.31, SE = 0.07, 95% CI [0.17, 0.45], p < .001) and significantly increased mean within-segment pause duration (b = 0.58, SE = 0.14, 95% CI [0.31, 0.85], p < .001). In practical terms, segments that condensed substantial background information into short textual spans were less likely to prompt immediate lexical substitution and more likely to require broader structural reconsideration.
Evaluative density also contributed significantly, although its primary effect was on delayed revision rather than initial decision latency. In the delayed-revision model, evaluative density yielded a positive coefficient (b = 0.23, SE = 0.05, 95% CI [0.13, 0.33], p < .001). Its effect on first-pause duration was comparatively modest (b = 0.34, SE = 0.13, 95% CI [0.09, 0.59], p = 0.010). These findings suggest that evaluative cues were not always recognized as bottlenecks in processing during initial translation. Instead, they frequently became problematic only after an initial draft had been produced, prompting translators to revisit issues of stance, tone, or narrative appropriateness.
Danmu salience remained a significant predictor after controlling for text-internal variables. It significantly predicted both first-pause duration (b = 0.36, SE = 0.14, 95% CI [0.09, 0.63], p = 0.011) and total revision density (b = 0.17, SE = 0.06, 95% CI [0.05, 0.29], p = 0.006). Although its effect size was smaller than those of cultural load and narrative compression, its consistent contribution suggests that segments that attract greater public attention in the danmu corpus also tend to require greater processing effort during translation. This finding is interpreted as convergent evidence of semantic salience rather than as evidence that danmu frequency directly causes translation difficulty.
Revision behavior and translation quality
Subsequent analyses examined whether translation quality was better explained by the overall quantity of revisions or by their temporal and functional characteristics. The quality model (Table 5) produced a clear pattern of results. The total number of revisions did not significantly predict translation quality (b = 0.04, 95% CI [-0.06, 0.14], p = 0.412), whereas both revision timing and revision function were significant predictors. Specifically, the delayed revision ratio (b = 1.12, 95% CI [0.39, 1.85], p = 0.003), discourse-level revisions (b = 1.76, 95% CI [0.96, 2.56], p < 0.001), and cultural adaptation revisions (b = 2.04, 95% CI [1.14, 2.94], p < 0.001) were all positively associated with translation quality. In contrast, surface revisions exhibited a small but statistically significant negative association (b = -0.63, 95% CI [-1.18, -0.08], p = 0.028).
These findings indicate that revision quantity alone cannot distinguish productive reformulation from unstable drafting (Figure 5). Qualitative inspection of the translation scripts supported this interpretation. In lower-scoring translations, revisions tended to cluster around local lexical substitutions and grammatical corrections without resolving deeper mismatches between Chinese narrative cues and English rhetorical framing. In higher-scoring translations, delayed revisions frequently involved reorganizing the information structure, introducing selective explicitation, or adjusting evaluative language to improve accessibility for an international readership. Representative translated danmu comments—including “this expression is very Chinese,” “the background needs to be explained,” and “foreign audiences may not understand this reference”—illustrate the types of culturally salient audience concerns associated with the corresponding segment clusters. These examples are presented solely to facilitate interpretation and should not be construed as evidence of participants’ intentions.
Overlap between danmu salience and translation bottlenecks
The final analysis examined whether publicly salient segments in the danmu corpus overlapped with those generating the greatest processing effort during translation. A composite effort index was calculated for each segment by standardizing and summing first-pause duration, long-pause frequency, and revision density. The top 20% of segments based on this composite effort index were then compared with the top 20% ranked by danmu salience.
The overlap was substantial: six of the nine highest-effort segments also appeared among the nine most danmu-salient segments, corresponding to a 66.7% overlap rate. Because the top-quintile comparison involved only nine segments, the Wilson confidence interval was relatively wide (95% CI approximately 35.4%–87.9%). Nevertheless, an exact chance-overlap analysis indicated that an overlap of six or more segments would be unlikely under random allocation (p ≈ .001). At the segment level, danmu salience was positively correlated with the composite effort index (Spearman's ρ = 0.41, p = 0.006). Most overlapping segments occurred in Texts B and C and centered on culturally embedded references, digital entrepreneurship, and evaluative positioning. These findings suggest that public discourse serves as an external indicator of where narrative meaning becomes particularly salient, though the relationship remains correlational rather than causal.
Taken together, these findings demonstrate that processing effort in translating China-focused narratives is segment-specific and primarily shaped by cultural load, narrative compression, and evaluative density. Translation quality depends not on overall revision volume but on delayed, discourse-sensitive, and culturally oriented revision strategies. Finally, the observed alignment between high-effort translation segments and danmu-salient narrative points supports an associative relationship between audience salience and translator processing effort, without implying that danmu salience itself causes translation difficulty.
DATA AVAILABILITY:
All raw data, keystroke logging files, processed segment-level observations, session-level observations, danmu-video metadata, and the first 5,000 displayed danmu comments are provided in the replication dataset workbook. The full danmu-comment file contains 18,642 comments and is deposited with the dataset package. The dataset can be accessed through the Zenodo repository at https://doi.org/10.5281/zenodo.19289665.

Figure 1: Conceptual framework of the study. (A) Segment-level textual predictors, including cultural load, evaluative density, narrative compression, and danmu salience. (B) Process indicators used to operationalize online decision-making behavior, including first-pause duration, mean pause duration, long-pause frequency, dictionary consultations, and cursor interruptions. (C) Two-dimensional classification of revision behavior according to revision timing and revision function. (D) Five dimensions used to evaluate translation quality: semantic accuracy, linguistic fluency, narrative coherence, cultural rendering, and register appropriateness. (E) Overall analytical framework showing the relationships among source-text salience, decision-making behavior, revision behavior, and translation quality, with danmu salience serving as an external indicator of semantic audience prominence. No exploratory factor analysis (EFA) or structural equation modeling (SEM) was performed or is presented in this study. Please click here to view a larger version of this figure.

Figure 2: Workflow of data collection and data alignment. (A) Participant recruitment, eligibility screening, informed consent, and collection of background information. (B) Administration of the three translation tasks in Translog-II under standardized experimental conditions. (C) Collection of process data, including keystrokes, pauses, cursor movements, revision episodes, and synchronized screen recordings. (D) Blind expert assessment of translation products, followed by adjudication of scoring discrepancies. (E) Construction of the auxiliary danmu corpus and its alignment with participant-, text-, segment-, process-, and product-level datasets to generate the final integrated dataset used for inferential analyses. Please click here to view a larger version of this figure.

Figure 3: Distribution of process and product measures across year groups and source texts. (A) Mean task completion time by source text and year group. (B) Distribution of first-pause duration across the three source texts. (C) Long-pause frequency by source text and year group. (D) Proportional distribution of revision categories across year groups, including surface revisions, meaning revisions, discourse-level revisions, and cultural adaptation revisions. (E) Total translation quality scores by year group for each source text. Collectively, the panels show that higher year levels are associated with fewer processing interruptions, more selective revision strategies, and higher translation quality, whereas Texts B and C impose greater processing demands than Text A. Please click here to view a larger version of this figure.

Figure 4: Effects of segment-level textual salience on translation process behavior. (A) Fitted marginal effect of cultural load on first-pause duration. (B) Fitted marginal effect of evaluative density on delayed revision density. (C) Fitted marginal effect of narrative compression on discourse-level revision density. (D) Heatmap of the standardized processing-effort index across all source-text segments, grouped by source text. (E) Standardized fixed-effect coefficients estimated from the mixed-effects regression models. The figure demonstrates that different textual characteristics activate distinct processing pathways, with cultural load, evaluative density, and narrative compression contributing to hesitation and revision through partially different mechanisms. Please click here to view a larger version of this figure.

Figure 5: Relationships between translation process, translation quality, and danmu salience. (A) Association between the delayed revision ratio and total translation quality score. (B) Relationship between revisions to cultural adaptation and the cultural-rendering score. (C) Association between total revision count and overall translation quality, demonstrating the limited predictive value of revision quantity alone. (D) Overlap between the top 20% of danmu-salient segments and the top 20% of high-effort translation segments. (E) Segment-level association between danmu salience and the standardized processing-effort index. Collectively, the panels show that successful translation performance is associated with strategically timed, functionally targeted revision rather than revision volume alone, and that publicly salient narrative segments also tend to coincide with translation bottlenecks. Please click here to view a larger version of this figure.
| Text ID | Narrative topic | Chinese characters | Number of segments | Cultural-load score (1–5) | Evaluative-density score (1–5) | Narrative-compression score (1–5) | Predicted difficulty | Validation note |
| Text A | Spring Festival return-home narrative | 212 | 14 | 3 | 2 | 3 | Moderate | Pilot-tested with 12 non-participant students |
| Text B | Dragon Boat Festival memory narrative | 219 | 15 | 5 | 3 | 4 | High | Pilot-tested with 12 non-participant students |
| Text C | Rural livestreaming and hometown development narrative | 223 | 15 | 4 | 5 | 4 | High | Pilot-tested with 12 non-participant students |
Table 1: Source-text profile and segment structure.
| Variable | Level | Operational definition | Measurement unit |
| First-pause duration | Segment | Time from segment onset to the first relevant target-language keystroke | Seconds |
| Mean within-segment pause duration | Segment | Mean duration of all pauses within the segment | Seconds |
| Long-pause frequency | Segment | Number of pauses longer than 2 seconds | Count |
| Dictionary consultations | Segment | Number of dictionary-window activations during segment processing | Count |
| Cursor interruptions | Segment | Number of cursor-based departures from linear drafting before stable completion | Count |
| Immediate revision | Segment | Revision made before drafting moves beyond the current local text area | Count / density |
| Delayed revision | Segment | Revision made after drafting has progressed to a later text area | Count / density |
| Surface revision | Revision function | Form-level correction without a semantic shift | Count / density |
| Meaning revision | Revision function | Lexical or syntactic alteration affecting meaning expression | Count / density |
| Discourse-level revision | Revision function | Adjustment to clause relations, coherence, or information order | Count / density |
| Cultural adaptation revision | Revision function | Target-audience-oriented reframing or explicitation of culture-bound meaning | Count / density |
| Composite effort index | Segment | Standardized composite of first-pause duration, long-pause frequency, and revision density | z-score |
| Danmu salience | Segment / semantic cluster | Weighted semantic-cluster proxy based on normalized keyword-cluster frequency, comment concentration, and evaluative intensity | Continuous score |
Table 2: Operational definitions of process variables.
| Dimension | Score range | Scoring focus |
| Semantic accuracy | 1-20 | Fidelity to source meaning, precision of transferred content, and absence of major distortion |
| Linguistic fluency | 1-20 | Grammaticality, naturalness, idiomaticity, and sentence-level readability |
| Narrative coherence | 1-20 | Information flow, clause linkage, progression, and textual stability |
| Cultural rendering | 1-20 | Handling of culture-bound references, explicitation, framing, and interpretive accessibility |
| Register appropriateness | 1-20 | Fit to intended readership, stylistic suitability, and control of tone |
Table 3: Translation quality rating rubric.
| Section | Variable | Year 2 (n = 24), mean ± SD | Year 3 (n = 24), mean ± SD | Year 4 (n = 24), mean ± SD |
| Participant-level session profile | Task completion time (sec) | 1070.84 ± 18.22 | 996.59 ± 20.08 | 927.76 ± 19.91 |
| Participant-level session profile | Mean first-pause duration (sec) | 7.09 ± 0.50 | 6.65 ± 0.40 | 6.38 ± 0.41 |
| Participant-level session profile | Long-pause frequency per task | 21.04 ± 0.32 | 18.47 ± 0.30 | 16.31 ± 0.19 |
| Participant-level session profile | Total revisions per task | 28.65 ± 2.87 | 26.24 ± 2.63 | 24.59 ± 2.42 |
| Revision configuration | Delayed revision ratio | 0.45 ± 0.01 | 0.54 ± 0.01 | 0.63 ± 0.01 |
| Revision configuration | Discourse-level revision ratio | 0.37 ± 0.01 | 0.40 ± 0.01 | 0.45 ± 0.01 |
| Revision configuration | Cultural adaptation revision ratio | 0.38 ± 0.01 | 0.42 ± 0.00 | 0.48 ± 0.01 |
| Revision configuration | Surface revision ratio | 0.04 ± 0.00 | 0.02 ± 0.00 | 0.00 ± 0.00 |
| Product quality | Total quality score | 77.82 ± 4.65 | 84.03 ± 4.44 | 88.27 ± 4.56 |
| Product quality | Cultural-rendering score | 77.73 ± 4.49 | 83.95 ± 4.85 | 88.18 ± 4.86 |
Table 4: Descriptive statistics for process and product variables.
| Outcome variable | Predictor | b | SE | 95% CI | t/z | p |
| First-pause duration | Cultural load | 0.91 | 0.12 | [0.67, 1.15] | 7.58 | < 0.001 |
| First-pause duration | Evaluative density | 0.34 | 0.13 | [0.09, 0.59] | 2.62 | 0.01 |
| First-pause duration | Narrative compression | 0.58 | 0.14 | [0.31, 0.85] | 4.14 | < 0.001 |
| First-pause duration | Danmu salience | 0.36 | 0.14 | [0.09, 0.63] | 2.55 | 0.011 |
| Long-pause frequency | Cultural load | 0.19 | 0.04 | [0.11, 0.27] | 4.75 | < 0.001 |
| Long-pause frequency | Evaluative density | 0.08 | 0.04 | [0.00, 0.16] | 2 | 0.046 |
| Long-pause frequency | Narrative compression | 0.14 | 0.05 | [0.04, 0.24] | 2.8 | 0.005 |
| Long-pause frequency | Danmu salience | 0.09 | 0.03 | [0.03, 0.15] | 3 | 0.003 |
| Total revision density | Cultural load | 0.27 | 0.06 | [0.15, 0.39] | 4.5 | < 0.001 |
| Total revision density | Evaluative density | 0.11 | 0.05 | [0.01, 0.21] | 2.2 | 0.028 |
| Total revision density | Narrative compression | 0.21 | 0.06 | [0.09, 0.33] | 3.5 | < 0.001 |
| Total revision density | Danmu salience | 0.17 | 0.06 | [0.05, 0.29] | 2.75 | 0.006 |
| Delayed revision density | Cultural load | 0.15 | 0.05 | [0.05, 0.25] | 3 | 0.003 |
| Delayed revision density | Evaluative density | 0.23 | 0.05 | [0.13, 0.33] | 4.6 | < 0.001 |
| Delayed revision density | Narrative compression | 0.26 | 0.06 | [0.14, 0.38] | 4.33 | < 0.001 |
| Delayed revision density | Danmu salience | 0.1 | 0.04 | [0.02, 0.18] | 2.5 | 0.013 |
| Total quality score | Total revisions | 0.04 | 0.05 | [-0.06, 0.14] | 0.82 | 0.412 |
| Total quality score | Delayed revision ratio | 1.12 | 0.37 | [0.39, 1.85] | 3.03 | 0.003 |
| Total quality score | Discourse-level revisions | 1.76 | 0.41 | [0.96, 2.56] | 4.29 | < 0.001 |
| Total quality score | Cultural adaptation revisions | 2.04 | 0.46 | [1.14, 2.94] | 4.43 | < 0.001 |
| Total quality score | Surface revisions | -0.63 | 0.28 | [-1.18, -0.08] | -2.25 | 0.028 |
| Total quality score | Year group | 3.48 | 0.69 | [2.13, 4.83] | 5.04 | < 0.001 |
Table 5: Mixed-effects models predicting segment-level process behavior and translation quality.