针对技术演讲同步翻译,精准选择术语上下文可显著提升翻译质量。
When to Use Extra Context: Evidence-Grounded Terminology Adaptation for Simultaneous Speech Translation

- 构建术语记忆库,根据实时流状态筛选关键术语
- 在英译中/德场景下命名实体召回率提升超70%
- 无需微调模型,适用于多种翻译架构
技术演讲的同步语音翻译中,额外上下文有价值,但全文档注入过于粗略。诊断实验发现,上下文收益主要来自论文特定术语的恢复,而非整体语义增强。为此提出EGTA框架:建立文档术语记忆,基于当前流状态选择紧凑候选术语,并仅用选定术语调整ASR/语音端与解码端决策空间。EGTA可在级联、端到端及仅生成型同步翻译设置中应用,无需完整模型微调。在包含MCIF-dev和ACL60/60-dev的ACL技术演讲同步翻译评估集上测试,于MCIF-dev上,EGTA-RG在英→中/英→德任务中分别带来+1.05/+0.59的BLEU提升,XCOMET-XL提升+0.019/+0.006,命名实体召回率相对提升+79%/+73%,缩写召回率提升+0.099/+0.171。在不同延迟设置下,EGTA持续提升XCOMET-XL、命名实体与缩写召回率。外部验证在ACL60/60-dev上也显示无额外微调下的术语召回稳定增长。打乱记忆对照与激活审计表明,性能提升源于论文特异性证据对齐,而非通用上下文提示。
原文摘要 · Abstract (English)
Extra context is valuable for simultaneous speech translation of technical talks, but injecting the entire document context into every streaming segment is often too coarse. Through diagnostic experiments, we find that context gains mainly come from paper-specific terminology recovery rather than uniform semantic enhancement. We therefore propose EGTA, an Evidence-Grounded Terminology Adaptation framework that builds a document terminology memory, selects compact candidate terms conditioned on the current streaming state, and adapts ASR/speech-side and decoder-side decision spaces using only the selected terms. EGTA can be instantiated in cascaded, end-to-end, and generation-only SimulST settings without full-model fine-tuning. We evaluate EGTA on an ACL technical-talk SimulST evaluation suite consisting of MCIF-dev and ACL60/60-dev. On MCIF-dev, EGTA-RG improves BLEU by +1.05/+0.59, XCOMET-XL by +0.019/+0.006, named-entity recall by +79\%/+73\% relative, and acronym recall by +0.099/+0.171 on En$\rightarrow$Zh and En$\rightarrow$De. Across MCIF-dev latency settings, EGTA consistently improves XCOMET-XL, named-entity recall, and acronym recall. External validation on ACL60/60-dev further shows consistent terminology-recall gains without additional fine-tuning. Shuffled-memory controls and activation audits provide evidence that the improvements are tied to paper-specific evidence alignment rather than generic context prompting.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。