A subscription to JoVE is required to view this content. Sign in or start your free trial.

Research Article

Decision-Making and Revision Behavior in Chinese-to-English Translation: Evidence from English Majors Translating China-Focused Narratives

182 views

⸱

DOI:

10.3791/71607

⸱

September 3rd, 2026

 , 

Corresponding Authors: Chunming Wu <wucm@hstc.edu.cn>

In This Article

Summary

This study examines decision-making and revision behaviors of English majors translating China-focused narratives, integrating keystroke-logging records and a danmu corpus to reveal how textual salience influences processing effort and translation quality.

Abstract

As China-focused narratives increasingly enter international circulation, Chinese-to-English translation has become a site of meaning negotiation rather than a straightforward linguistic transfer. For student translators, the primary challenge extends beyond lexical selection; it involves deciding how to render culturally saturated references, evaluative stances, and narrative positioning for readers outside the source context. This study examines the real-time decision-making and revision behaviors of 72 English majors translating China-focused narratives. Using a process-product design, the research integrates keystroke-logging data collected with Translog-II, translation product evaluations, and an auxiliary danmu (time-synced commentary) corpus treated as an audience-salience layer rather than as a direct measure of segment-level linguistic difficulty. Translation processes were analyzed across pause duration, revision frequency, revision location, and stage distribution, while product quality was assessed through expert rating and targeted coding of culture-related segments. A corpus of 18,642 danmu comments from 24 Bilibili videos was used to identify publicly salient narrative themes and culturally sensitive expressions. Results show that culture-loaded and evaluatively dense segments triggered longer pauses, increased recursive revisions, and denser mid-process reformulations. Higher-quality translations were associated with selective, strategically timed revisions at the discourse and cultural-adaptation levels rather than with a higher total revision count. The overlap between danmu-salient narrative triggers and high-effort translation segments is interpreted as convergent evidence of audience salience and translator processing pressure, not as evidence of causation. The study advances research on the translation process by linking revision behavior to culturally embedded narrative translation and offering implications for translator training in cross-cultural communication.

Introduction

In recent years, the translation of China-related discourse has moved from a relatively specialized concern to a central issue in international communication1. What is at stake in such translation is not simply the transfer of propositional content, but the re-articulation of public narratives that carry political positioning, cultural framing, and audience design. Work on the translation of Chinese political discourse has made this point increasingly clear: when texts are translated for readers outside China, choices about wording, modality, and textual emphasis also become choices about how China is narrated and interpreted abroad2. Once the problem is framed in that way, a product-only approach becomes insufficient. A finished translation can show what was chosen, but it cannot show when uncertainty emerged, where reformulation started, or how competing solutions were tested and abandoned3.

These limitations help explain why empirical translation process research has become an important strand within translation studies, using process-tracing tools to capture pauses, deletions, insertions, and the temporal shape of decision-making rather than inferring them indirectly from the final text4. This does not displace product, sociological, cultural, or corpus-based approaches; rather, it adds evidence about the moment-by-moment formation of translation decisions. Recent work has moved from describing isolated pauses to modeling translation as a temporally structured activity. The hesitation-orientation-flow taxonomy, for instance, treats translation behavior as a patterned alternation between uncertainty, search, and fluent production, explaining why particular stretches of a task become unstable and trigger reformulation5. This shift is especially relevant in student translation, where language proficiency, subject knowledge, and strategic control are still developing concurrently6.

Evidence from combined eye-tracking and keystroke logging shows that novice translators do not distribute attention evenly, and that target-text production often becomes a major site of cognitive load during Chinese-to-English translation7. Consequently, what appears as a local wording problem in the product may actually reflect a broader reorganization of attention and effort during the real-time process8. Recent evidence further suggests that student translators may produce texts with similar surface adequacy while differing substantially in the amount of effort, uncertainty, and self-monitoring required. A satisfactory final rendering may still rest on a fragile or inefficient decision path that can only be captured through process-sensitive measures9. Furthermore, comparative work indicates that translation expertise involves stable metacognitive control over when to pause, search, continue drafting, or return for revision10.

Revision is the exact point at which this hidden cognitive process becomes most visible. It is where translators return to earlier wording, test alternatives, and decide whether a problem is lexical, syntactic, discursive, or cultural. Consequently, revision should not be treated as a late polishing step but as a core indicator of translator development. Better outcomes are linked to how alternatives are consolidated during production, demonstrating that effective revision is selective, strategically timed, and responsive to greater textual demands rather than a simple accumulation of edits11,12,13.

The complexity of decision-making and revision becomes sharper when the source text is a China-focused narrative rather than a neutral informational passage. Such texts frequently rely on culturally saturated assumptions, historically situated expressions, and evaluative cues whose meaning depends on shared background rather than explicit exposition14,15,16. Text difficulty alters the very granularity of processing; narrative texts carrying cultural or evaluative density compel translators to work in smaller units, produce non-linear operations, and engage in repeated local restructuring8. For student translators, the central problem is deciding how much of that implicit background to unpack or reorganize for readers outside the original discourse community, without flattening cultural specificity.

Three specific gaps motivate the present study. First, translation process metrics such as pause duration, long-pause frequency, cursor interruption, and dictionary consultation are often examined separately from revision typologies, making it difficult to see how decision pressure becomes visible as later reformulation. Second, culturally specific narrative categories have received less process-based attention than general academic or expository texts, even though China-focused narratives often require translators to handle cultural load, evaluative density, and compressed background implication at the same time14,15,16. In this study, cultural load refers to the density of culture-bound references and shared presuppositions; evaluative density refers to explicit or implicit stance-taking, affective intensity, and value-laden framing; and narrative compression refers to the degree to which background information, causal relations, or social positioning is condensed into short textual spans. Third, translation tasks are usually designed solely on the basis of text-internal analysis, with little consideration of whether segments identified as difficult align with points that attract public attention in digital circulation. Danmu (time-synced video comments) can function as paratext and participatory discourse, highlighting what audiences notice, contest, and reinterpret17,18; however, danmu salience must be treated as a semantic-level indicator of audience engagement rather than as a direct substitute for segment-level linguistic difficulty.

Addressing these gaps, the present study examines how English majors make decisions and revise as they translate China-focused narratives from Chinese into English. By bringing translation-process data and digital audience discourse into the same analytical frame, the study aims to explain not only where student translators struggle but also why those struggles cluster around particular kinds of narrative meaning.

Specifically, this study addresses four research questions: (1) What decision-making patterns characterize English majors’ Chinese-to-English translation of China-focused narratives? (2) How do revision behaviors vary across segments with different levels of cultural load and narrative density? (3) Which dimensions of revision behavior are associated with higher translation quality? and (4) Do culturally salient points identified in danmu discourse overlap with the segments that generate the greatest processing effort during translation?

Correspondingly, the study tests four linked hypotheses in a single analytic sequence: H1 proposes that segments with higher cultural load will trigger longer pauses, greater decision uncertainty, and more frequent revision than segments with lower cultural load; H2 distinguishes two pathways, predicting that evaluative density will increase delayed revision while narrative compression will increase discourse-level reformulation; H3 predicts that translation quality will be explained more strongly by the selectivity and timing of revision than by total revision count alone; and H4 predicts that danmu-salient segments will overlap with high-effort translation segments as an associative convergence between audience salience and translator processing pressure.

Access restricted. Please log in or start a trial to view this content.

Protocol

All methods involving human subjects were conducted in compliance with institutional guidelines and were approved by the institutional review board (IRB) or the human research ethics committee of Hanshan Normal University. Informed consent was obtained from all participants prior to the experiment.

Conceptual framework and research design

This study employed a process–product research design with an auxiliary discourse layer to examine how audience-generated danmu comments influence subtitle translation. The analytical framework integrated textual characteristics, indicators of the translation process, and translation quality outcomes. Specifically, it investigated whether audience attention, as reflected by danmu density, directed translators' cognitive resources toward particular textual features and whether this attentional allocation subsequently affected translation quality.

To operationalize this framework, the study first established segment-level textual features based on four dimensions: cultural load, evaluative density, narrative compression, and danmu salience. Cultural load referred to the presence of culturally specific references that required adaptation or explanation. Evaluative density measured the concentration of emotionally or attitudinally charged language within a segment. Narrative compression captured the extent to which semantic content was condensed under subtitle constraints. Danmu salience represented the degree of audience attention directed toward each subtitle segment, quantified by the volume of synchronized danmu comments. Unlike conventional measures of translation difficulty, danmu salience reflected socially distributed attention rather than intrinsic textual complexity. A segment receiving extensive audience engagement was not necessarily more difficult to translate; instead, it represented a point of heightened communicative significance within the viewing experience.

The analytical framework comprised three interconnected datasets. The participant-level dataset included individual demographic characteristics and translation expertise. The process–product dataset linked each subtitle segment to behavioral indicators extracted from keystroke logging, together with expert translation quality ratings. The danmu dataset contained synchronized audience comments aligned with subtitle timing, from which danmu salience values were computed. These datasets were linked via participant and segment identifiers, enabling simultaneous analysis of each translated subtitle segment in relation to translator characteristics, cognitive processing behavior, audience attention, and translation outcomes. The final integrated database comprised 216 complete translation sessions and 3,168 participant–text–segment observations, providing the empirical foundation for the subsequent process–product analyses.

Participant screening and background profiling

The sample comprised 72 English majors recruited from a public university in southern China, evenly distributed across the second, third, and fourth years of undergraduate study (24 participants per year group). All participants were native Chinese speakers receiving formal instruction in English writing and translation. To ensure comparability within the trainee-translator population, individuals who had resided in an English-speaking country for more than 6 months or held professional translator certification were excluded. Participants were not stratified into separate experimental groups according to translation experience, cultural familiarity, or second-language proficiency. Instead, year level, recent English proficiency test scores, prior translation coursework, self-rated familiarity with China-related cultural topics, and average weekly engagement in bilingual practice were recorded and included as participant-level covariates in the statistical models. This clarification also resolves the participant-number inconsistency raised during peer review: both the manuscript and the replication dataset include 72 participants, 216 task sessions, and 3,168 segment-level observations.

Source text selection and task construction

The translation materials consisted of three contemporary China-focused narrative texts developed specifically for this study to reflect authentic public discourse and avoid reproducing copyrighted materials. Each text contained 205–225 Chinese characters but differed systematically in its internal characteristics. Text A presents a Spring Festival homecoming narrative with moderate cultural load and relatively straightforward syntax. Text B recounts a Dragon Boat Festival memory and contains a higher density of culture-specific expressions and implicit background knowledge. Text C describes a rural livestreaming and e-commerce narrative characterized by dense evaluative language and compressed narrative positioning.

The texts were systematically segmented using clause boundaries, followed by meaning-unit boundaries, yielding 44 analyzable units across the three translation tasks (Table 1). Before formal coding, a coding manual defined cultural load, evaluative density, and narrative compression using 1–5 ordinal scales. A pilot study involving 12 students who did not participate in the main experiment confirmed the intended difficulty calibration, the clarity of the task instructions, and the stability of the segment boundaries for keystroke alignment.

Auxiliary corpus construction and salience mapping

To identify publicly salient narrative cues, an auxiliary danmu corpus was compiled from 24 Bilibili videos uploaded between January 2023 and December 2025. Videos were included only if they matched one of the three semantic domains represented in the source texts, had accumulated more than 50,000 views, contained dense danmu interaction, and permitted the export and thematic mapping of time-synchronized comments. After duplicate removal and filtering of emojis, onomatopoeic fillers, and irrelevant fragments, the final corpus comprised 18,642 comments.

Content words were standardized and clustered using source-text-related keyword anchors. Each source-text segment was then assigned a continuous danmu salience score based on a weighted composite of normalized keyword-cluster frequency, comment concentration within the corresponding thematic cluster, and mean evaluative intensity. This procedure enables comparisons between translator processing effort and audience salience at the semantic-cluster level rather than through direct lexical matching. Because danmu reflects audience responses to online videos rather than the experimental translation task itself, danmu salience is interpreted as an external indicator of audience engagement rather than direct evidence of the linguistic difficulty of a source segment.

Experimental procedure and data capture

Translation sessions were conducted individually in a controlled computer laboratory under standardized hardware, software, and internet-access conditions (Figure 2). Following a five-minute orientation emphasizing the production of readable English for an international audience, participants completed a three-minute English typing warm-up to minimize the effects of keyboard acclimatization. The three translation tasks were administered sequentially using Translog-II, with the task order counterbalanced in a Latin-square design to distribute potential fatigue and order effects evenly across participants. Participants were allocated a maximum of 20 minutes per text, separated by five-minute rest intervals.

Although the built-in bilingual dictionary was permitted, machine translation systems and external search engines were prohibited. All keystrokes, pauses, and cursor movements were logged automatically and supplemented by synchronous screen recordings to verify non-linear revision episodes. After completing each task, participants completed a self-reported difficulty assessment used solely to assist in interpreting borderline coding cases.

Process measurement and variable coding

Keystroke records were aligned with the predefined segment structure. An alignment episode began with the first target-language keystroke corresponding to a segment and ended when the participant definitively progressed to the next segment. Overlapping revisions were reassigned using time-stamped cursor-tracking records.

Decision-making behavior was quantified using five indicators: first-pause duration before initial segment translation, mean within-segment pause duration, the frequency of pauses exceeding two seconds, dictionary consultation counts, and cursor-based interruption counts indicating departures from linear drafting (Table 2). Revision behavior was classified along two dimensions: timing and function. Timing distinguished immediate local revisions from delayed revisions performed after drafting had progressed beyond the current segment. Functional coding classified revisions as surface revisions (form-level corrections without semantic change), meaning revisions (lexical or syntactic modifications), discourse-level revisions (adjustments to clause relationships or coherence), and cultural adaptation revisions (target-audience-oriented reformulations). Two trained coders independently analyzed each revision episode before adjudication and assigned a single primary functional category to each event. Inter-rater reliability was robust across all coded measures. Cohen’s ĸ indicated substantial agreement for the categorical coding of revision functions (ĸ = 0.84, p < 0.001). For the ordinal text-level variables, intraclass correlation coefficients (ICC, two-way mixed-effects model, absolute agreement) demonstrated high reliability for cultural load (ICC = 0.86), evaluative density (ICC = 0.78), and narrative compression (ICC = 0.82). Any initial coding discrepancies were subsequently resolved through joint adjudication.

Translation quality assessment

Translation quality was evaluated independently of the process data by two raters with doctoral-level training and extensive experience in translation instruction. Blinded to participant identities and process measures, the raters applied a 100-point rubric, evenly distributed across five dimensions: semantic accuracy, linguistic fluency, narrative coherence, cultural rendering, and register appropriateness (Table 3). Score discrepancies greater than eight points triggered joint review and consensus rescoring, with the adjudicated average used as the final translation quality score. The pre-adjudication rating records were retained to facilitate transparent reporting of inter-rater reliability in the replication materials.

Statistical modeling

The analyses proceeded in three sequential stages. First, descriptive statistics were computed to summarize all variables across year groups and text types. Second, segment-level processing effort was analyzed using mixed-effects regression models to account for the nested structure of segments within texts and repeated observations from individual participants. Dependent variables, including first-pause duration and revision frequency, were modeled as functions of cultural load, evaluative density, narrative compression, and danmu salience while controlling for participant proficiency. Predictor correlations, variance inflation diagnostics, 95% confidence intervals, effect-size estimates, and model-fit indices were included to improve statistical transparency.

Finally, translation quality was modeled to evaluate the predictive contributions of revision timing and revision function relative to overall revision frequency. The overlap between danmu-salient segments and high-effort translation segments was assessed using the upper quintile of each distribution, a standardized composite effort index, and a chance-overlap analysis. No exploratory factor analysis or structural equation modeling was performed. Accordingly, the accompanying figures are presented solely as conceptual, descriptive, or regression-based summaries.

Access restricted. Please log in or start a trial to view this content.

Results

Overall profile of the translation sessions

All 72 participants completed the three translation tasks under standardized experimental conditions, yielding 216 complete translation sessions. Following segment alignment, the final analytical dataset comprised 3,168 participant-text-segment observations. The mean task completion time across the sample was 998.4 s per text, with clear differences across year groups and source texts (Table 4; F...

Access restricted. Please log in or start a trial to view this content.

Discussion

The present findings demonstrate that difficulty in the Chinese-to-English translation of China-focused narratives is not evenly distributed across a text. Instead, it clusters in segments where cultural references, evaluative stances, and compressed narrative meanings converge. This supports a process-based, rather than a vocabulary-deficit, view of translation difficulty. Empirical translation process research emphasizes that translation is a temporally organized activity in which hesitation, orientation, and productio...

Access restricted. Please log in or start a trial to view this content.

Disclosures

The authors have nothing to disclose.

Author Contributions
Weijia Chen: Conceptualization, methodology, formal analysis, investigation, writing-original draft.
Chunming Wu: Study design, guidance on formal analysis and data-handling procedures, and review and feedback on revisions in response to reviewer comments.

Acknowledgements

This research was financially supported by two grants from Hanshan Normal University. Specifically, it constitutes a phased research outcome of the 2023 First-Class Undergraduate Course (Offline First-Class Course) program for "Theory and Practice of Chinese–English Translation II" (Approval No. YHS [2023] 172). Furthermore, the study was supported by the 2025 Teaching Quality and Teaching Reform Project, designated as a Model Course for Curriculum-Based Ideological and Political Education for "Chinese–English Translation" (Approval No. YHSJ〔2025〕126). The authors gratefully acknowledge the institutional support provided for the completion of this research.

Access restricted. Please log in or start a trial to view this content.

Materials

List of materials used in this article
NameCompanyCatalog NumberComments
Bilibili danmu corpusBilibili24 videos; 18,642 comments;https://www.bilibili.comConstruction of the auxiliary danmu corpus and calculation of segment-level danmu salience
Microsoft ExcelMicrosoft CorporationExcel 2021;https://www.microsoft.com/microsoft-365/excelData cleaning, coding organization, and preliminary data management
RR Foundation for Statistical ComputingR 4.3.1;https://www.r-project.orgMixed-effects modeling, correlation analysis, and figure generation
Screen recording softwareLaboratory workstationLocal workstation utility; no catalog numberVerification of non-linear revision episodes and backup of process data
SPSSIBM CorporationSPSS Statistics 27.0;https://www.ibm.com/products/spss-statisticsDescriptive statistics, ANOVA, and basic statistical analyses
Translog-IICopenhagen Business School / CRITTTranslog-II;https://sites.google.com/site/centretranslationinnovation/translog-iiKeystroke logging and recording of pauses, cursor movements, and revision episodes during translation tasks

References

  1. Valdeón RA, Li S. Political discourse translation in contemporary Chinese and Western contexts. Translator. 2024;30(4):1-14.
  2. Du L, Afzaal M. Analyzing modality-mediated ideology in translated Chinese political discourses: An ideological square model approach. Humanit Soc Sci Commun. 2024;11:956.
  3. Carl M. Empirical translation process research: Past and possible future perspectives. Transl Cogn Behav. 2024;6(2):252-274.
  4. Qassem M, Al Thowaini BM. Cognitive processes and translation quality: Evidence from keystroke-logging software. J Psycholinguist Res. 2023;52(5):1589-1604.
  5. Carl M, et al. Hesitation, orientation, and flow: A taxonomy for deep temporal translation architectures. Ampersand. 2024;12:100164.
  6. Swar O, Mohsen M. Students' cognitive processes in L1 and L2 translation: Evidence from a keystroke logging program. Interact Learn Environ. 2023;31(10):6247-6263.
  7. Wang Y, Li S, Rasmussen YZ. Translators' allocation of cognitive resources in two translation directions: A study using eye-tracking and keystroke logging. Appl Sci. 2025;15(8):4401.
  8. Wang F, Xu Q. Processing of translation units by student and semi-professional translators in translating texts with different levels of translation difficulty. PLoS One. 2025;20(4):e0320809.
  9. Chen X, Yan JX. Examining student translators' writing and translation products: Quality, errors, and self-perceived mental workload. Interpret Transl Train. 2024;18(3):465-485.
  10. Dong D, Chen ML. Metacognitive strategies in translation: A comparative study of student and professional translators. Humanit Soc Sci Commun. 2025;12:890.
  11. Chen X, Yan JX. Investigating the processing effort in translation students' L2 writing and L2 translating: Evidence from keylogging. Humanit Soc Sci Commun. 2025;12:1810.
  12. Cao L, Doherty S, Lee JF. The process and product of translation revision: Empirical data from student translators using eye tracking and screen recording. Interpret Transl Train. 2023;17(4):548-565.
  13. Van Egdom GW, Schrijver I, Verplaetse H, Segers W. The impact of collaborative processes on target text quality in translator training. Interpret Transl Train. 2024;18(3):486-506.
  14. Riondel A. How to teach revision: Tips from an interview study. Interpret Transl Train. 2024;18(3):507-522.
  15. Lu S, Chen X. An experimental study on viewing perception and gratification of danmu subtitled online video streaming. J Spec Transl. 2024;42:217-238.
  16. Lei Y, Peng M, Han Y. Mistaken presuppositions and the translation of cultural texts: A study of two English translations of The Classic of Tea. Perspectives. 2025. doi:10.1080/0907676X.2025.2599854.
  17. Kim HR. Danmu as paratext: An exploratory study of paratextual functions in Chinese fan-subtitled videos. J Transl Stud. 2025;26(3):197-231.
  18. Wu X, Fitzgerald R. The danmu discourse of user engagement with cross-posted broadcast interviews in Chinese social media. Journalism. 2025;26(3):733-751.
  19. Yamada M, Carl M, Schaeffer MJ. Introduction to the special Ampersand issue on empirical translation process research. Ampersand. 2024;12:100179.
  20. Peng X, Wang X, Li X. When student translators meet with machine translation: The impacts of machine translation quality and perceived self-efficacy on post-editing performance. SAGE Open. 2024;14(4):21582440241291624.
  21. Yang Y. Making sense of the raw meat: A social semiotic interpretation of user translation on the danmu interface. Discourse Context Media. 2021;44:100550.
  22. Zeng D. Danmu subtitling as a self-regulative practice: A descriptive discourse analysis of Bilibili danmu subtitles from ethical perspectives. Front Psychol. 2025;16:1577260.
  23. Li Y, Cai X. The translation of ethnic literature: An analysis of the Chinese and English translations of the Yi folk narrative poem Ashima and their dissemination. Perspectives. 2025. doi:10.1080/0907676X.2025.2502935.

Access restricted. Please log in or start a trial to view this content.

Reprints and Permissions

Explore More Articles

Chinese English TranslationDecision-Making BehaviorCulture-Loaded SegmentsNarrative TranslationAudience SalienceTranslation ProcessCross-Cultural CommunicationKeystroke LoggingTranslation Quality