Participant baseline characteristics and data availability
The questionnaire contained 62 participants (31 novice and 31 trained), and the task sheet contained 124 records, indicating two completed task records per participant (Table 1). The novice group had less than 12 months of structured translation training (observed maximum, 8.6 months), whereas the trained group had at least 12 months (observed minimum, 13.3 months). Age, sex, English-learning duration, proficiency, and task-order distributions were similar between groups. Figure 1 reports the final analyzed sample and correctly distinguishes task-specific data: keystroke logs and screen recordings apply to written translation, audio and verbatim transcripts apply to sight interpreting, and quality ratings apply to both.
Behavioral profiles across task conditions
Behavioral profiles differed between the two task conditions and between training groups (Table 2; Figure 2A–E). Written translation had longer completion times and longer mean pauses, whereas sight interpreting had more pauses per 100 information units under the reported threshold. Within each task condition, the trained group completed tasks faster and had higher mean output scores. These contrasts should not be interpreted as pure modality effects because text length, time limit, output mode, and revision opportunity also differed. Pauses are not uniformly disruptive: a pause may reflect unresolved difficulty, successful planning, monitoring, or input-method activity, so pause measures are interpreted together with revision, delivery, and quality measures.
Strategic revision, self-repair, and predictive modeling
The coding framework separated episode type from functional outcome (Table 3; Figure 3A–D). Trained participants produced more written-revision episodes than novices (12.4 ± 3.7 vs. 8.9 ± 3.0 per task), whereas total oral self-repair did not differ significantly between groups (4.9 ± 1.6 vs. 4.4 ± 1.5 per task). The trained group nevertheless had a larger proportion of effective episodes in both tasks. The reported mixed-effects summary in Table 4 presents associations rather than causal paths; Figure 3E therefore displays the reported fixed-effect estimates and confidence intervals instead of a causal diagram.

Figure 1: Participant flow, task-order allocation, and task-specific data availability. (A) Final analyzed sample of 62 participants and 124 task records. (B) Counterbalanced task-order allocation within novice and trained groups. (C) Availability matrix in which keystroke logs and screen recordings are applicable only to written translation, audio recordings, and verbatim transcripts only to sight interpreting, and quality ratings to both tasks. (D) Group-by-order balance (chi-square test, p = 0.799). n = 62 participants (31 per group). Please click here to view a larger version of this figure.

Figure 2: Behavioral profiles across written-translation and sight-interpreting task conditions. (A) Task-specific recording setup. (B) Initial/onset latency and total task time. (C) Joint distribution of pause frequency and mean pause duration. (D) Task-specific deletion, revision, and oral-disfluency indicators. (E) Written production rate and oral speech rate. In panels B, C, and E, points and error bars show group means ± SD; in panel D, points show trained-minus-novice mean differences and error bars show 95% confidence intervals. n = 31 participants per group. In panel E, the y-axis is displayed on a logarithmic scale. Please click here to view a larger version of this figure.

Figure 3: Revision/self-repair coding, functional outcome, and output-quality associations. (A) Decision framework for independently assigning episode type and functional outcome. (B) Mean counts of lexical, syntactic, semantic, and stylistic episodes in each task condition. (C) Mean proportions of effective, neutral, and detrimental episodes. (D) Descriptive association between effective-episode rate and total output-quality score. (E) forest plot displaying fixed effects from a linear mixed-effects model with 95% confidence intervals (detailed in Table 4); the plot is associative and does not represent a causal path model. Please click here to view a larger version of this figure.
| Variables | Novice group (n = 31) | Trained group (n = 31) | p value |
| Age, years | 21.9 ± 1.7 | 22.4 ± 1.9 | 0.279 |
| Female, n (%) | 22 (71.0) | 21 (67.7) | 0.779 |
| English learning duration, years | 12.4 ± 2.1 | 12.9 ± 2.2 | 0.36 |
| English proficiency score | 548.6 ± 41.3 | 562.4 ± 45.8 | 0.216 |
| Translation training duration, months | 5.6 ± 3.0 | 18.7 ± 5.4 | <0.001 |
| Interpreting training duration, months | 2.3 ± 2.1 | 13.9 ± 5.7 | <0.001 |
| Prior written translation practice ≥5 tasks, n (%) | 9 (29.0) | 24 (77.4) | <0.001 |
| Prior sight interpreting practice ≥5 tasks, n (%) | 5 (16.1) | 20 (64.5) | <0.001 |
| Written translation first, n (%) | 16 (51.6) | 15 (48.4) | 0.799 |
| Sight interpreting first, n (%) | 15 (48.4) | 16 (51.6) | 0.799 |
Table 1: Participant characteristics and baseline language-training profiles by group. Variables are presented as mean ± standard deviation for continuous data and as counts (percentages) for categorical data. The reported p-values reflect between-group comparisons to verify baseline demographic balance and confirm the intended differences in training duration.
| Variables | Written translation Novice (n = 31) | Written translation Trained (n = 31) | Sight interpreting Novice (n = 31) | Sight interpreting Trained (n = 31) | Task effect p | Group effect p |
| Initial/onset latency, s | 31.8 ± 12.4 | 24.6 ± 9.7 | 14.2 ± 5.6 | 10.8 ± 4.2 | <0.001 | 0.003 |
| Total task time, s | 1038.5 ± 168.2 | 925.4 ± 151.6 | 305.6 ± 59.3 | 274.8 ± 50.7 | <0.001 | 0.004 |
| Pause frequency, per 100 information units | 18.6 ± 6.1 | 14.2 ± 5.0 | 24.8 ± 7.4 | 19.6 ± 6.3 | <0.001 | 0.001 |
| Mean pause duration, s | 4.1 ± 1.2 | 3.3 ± 1.0 | 2.7 ± 0.8 | 2.3 ± 0.7 | 0.012 | 0.002 |
| Production/speech rate, Chinese characters/min | 28.6 ± 6.7 | 34.1 ± 7.5 | 93.4 ± 18.9 | 108.7 ± 20.4 | | 0.001 |
| Deletion rate, % | 8.9 ± 3.4 | 6.8 ± 2.7 | | | | 0.009 |
| Revision density, per 100 Chinese characters | 7.6 ± 2.8 | 8.9 ± 3.1 | | | | 0.087 |
| Filled pause frequency, per min | | | 5.7 ± 1.9 | 4.2 ± 1.6 | | 0.002 |
| Repetition frequency, per min | | | 4.9 ± 1.8 | 3.6 ± 1.4 | | 0.004 |
| False-start frequency, per min | | | 2.8 ± 1.1 | 2.0 ± 0.9 | | 0.006 |
| Self-repair frequency, per min | | | 4.8 ± 1.7 | 3.9 ± 1.4 | | 0.031 |
| Total quality score, 0–25 | 16.8 ± 2.3 | 20.1 ± 2.1 | 15.4 ± 2.6 | 18.9 ± 2.4 | 0.002 | <0.001 |
| Accuracy, 0–5 | 3.4 ± 0.7 | 4.1 ± 0.6 | 3.1 ± 0.8 | 3.8 ± 0.7 | 0.004 | <0.001 |
| Completeness, 0–5 | 3.3 ± 0.8 | 4.0 ± 0.6 | 3.0 ± 0.8 | 3.7 ± 0.7 | 0.006 | <0.001 |
| Lexical appropriateness, 0–5 | 3.2 ± 0.7 | 4.0 ± 0.6 | 3.1 ± 0.7 | 3.8 ± 0.6 | 0.041 | <0.001 |
| Coherence, 0–5 | 3.5 ± 0.6 | 4.0 ± 0.6 | 3.1 ± 0.7 | 3.8 ± 0.6 | 0.003 | <0.001 |
| Target-language naturalness, 0–5 | 3.4 ± 0.7 | 4.0 ± 0.6 | 3.1 ± 0.7 | 3.8 ± 0.7 | 0.007 | <0.001 |
Table 2: Behavioral process indicators and output quality across task conditions and groups. Values are expressed as mean ± standard deviation. The reported p-values indicate the main effects of task modality (written translation vs. sight interpreting) and training background (novice vs. trained) on each process and quality variable.
| Variables | Written translation Novice | Written translation Trained | p value | Sight interpreting Novice | Sight interpreting Trained | p value |
| Lexical revision/self-repair, n/task | 3.4 ± 1.5 | 4.2 ± 1.7 | 0.053 | 1.8 ± 0.8 | 1.7 ± 0.7 | 0.612 |
| Syntactic revision/self-repair, n/task | 2.1 ± 1.1 | 3.0 ± 1.2 | 0.004 | 1.1 ± 0.7 | 1.2 ± 0.8 | 0.645 |
| Semantic revision/self-repair, n/task | 1.8 ± 1.0 | 2.5 ± 1.1 | 0.011 | 0.9 ± 0.6 | 1.1 ± 0.7 | 0.232 |
| Stylistic revision/self-repair, n/task | 1.6 ± 0.9 | 2.7 ± 1.2 | <0.001 | 0.6 ± 0.4 | 0.9 ± 0.6 | 0.026 |
| Total revision/self-repair episodes, n/task | 8.9 ± 3.0 | 12.4 ± 3.7 | <0.001 | 4.4 ± 1.5 | 4.9 ± 1.6 | 0.209 |
| Effective episodes, % | 52.6 ± 13.5 | 68.4 ± 12.1 | <0.001 | 45.1 ± 14.2 | 62.3 ± 13.0 | <0.001 |
| Neutral episodes, % | 31.8 ± 10.4 | 22.7 ± 8.8 | 0.001 | 35.6 ± 11.7 | 25.4 ± 9.6 | 0.001 |
| Detrimental episodes, % | 15.6 ± 7.1 | 8.9 ± 5.6 | <0.001 | 19.3 ± 8.6 | 12.3 ± 6.8 | 0.001 |
Table 3: Type and effectiveness of written revisions and oral self-repairs across task conditions. Episode types are reported as mean counts per task (± standard deviation), and functional outcomes are reported as mean percentages (± standard deviation). The p-values indicate the significance of between-group differences within each task modality.
| Fixed effects | β | SE | 95% CI | p value |
| Intercept | 16.74 | 0.42 | 15.91 to 17.57 | <0.001 |
| Task modality: sight interpreting vs written translation | -1.08 | 0.29 | -1.65 to -0.51 | <0.001 |
| Training group: trained vs novice | 2.84 | 0.47 | 1.91 to 3.77 | <0.001 |
| Task modality × training group | 0.26 | 0.47 | -0.67 to 1.19 | 0.581 |
| Pause frequency, per 100 information units | -0.12 | 0.05 | -0.22 to -0.02 | 0.018 |
| Mean pause duration, s | -0.34 | 0.16 | -0.66 to -0.02 | 0.041 |
| Revision/self-repair density | 0.07 | 0.07 | -0.07 to 0.21 | 0.334 |
| Effective revision/self-repair rate, per 10% increase | 0.56 | 0.13 | 0.30 to 0.82 | <0.001 |
| Deletion/disfluency burden | -0.18 | 0.08 | -0.34 to -0.02 | 0.027 |
| Production/speech rate, standardized | 0.42 | 0.19 | 0.04 to 0.80 | 0.032 |
Table 4: Reported fixed effects from the linear mixed-effects model predicting total output quality. The summary details the unstandardized coefficients (β), standard errors (SE), 95% confidence intervals (CI), and exact p-values for all specified task and behavioral predictors.