arXiv:2601.11002cs.CL2026-01ACL

给大模型翻译加新动作,让机器同传更像人。

Redefining Machine Simultaneous Interpretation: From Incremental Translation to Human-Like Strategies

  • 用四个新操作(切句、删减、简写、代词化)重构实时翻译流程
  • 在多个语言对上提升语义质量并降低延迟,效果优于传统方法
  • 适合研究实时翻译或人机交互的开发者参考

同步机器翻译(SiMT)需在严格实时约束下产出高质量译文,传统仅含读/写动作的策略难以满足。本文扩展动作空间,引入四种自适应操作:断句(Sentence_Cut)、删减(Drop)、部分摘要(Partial_Summarization)和代词化(Pronominalization),实现实时内容重构、省略与简化,同时保持语义一致。在大语言模型框架中集成这些动作,并通过感知动作提示构建训练样本。为评估译文质量与字级单调性,进一步开发了时延感知的文本转语音(TTS)流水线,使输出语音具备真实时间特征。在ACL60/60英语-中文、英语-德语及英语-日语基准上实验表明,本框架在语义指标上持续优于参考译文与萨拉米基线,且延迟更低。尤其当结合使用删减与断句时,流畅性与延迟的平衡显著提升。结果表明,丰富基于大模型的同步翻译动作空间,是缩小机器与人类口译差距的可行方向。

原文摘要 · Abstract (English)

Simultaneous Machine Translation (SiMT) requires high-quality translations under strict real-time constraints, which traditional policies with only READ/WRITE actions cannot fully address. We extend the action space of SiMT with four adaptive actions: Sentence_Cut, Drop, Partial_Summarization and Pronominalization, which enable real-time restructuring, omission, and simplification while preserving semantic fidelity. We adapt these actions in a large language model (LLM) framework and construct training references through action-aware prompting. To evaluate both quality and word-level monotonicity, we further develop a latency-aware TTS pipeline that maps textual outputs to speech with realistic timing. Experiments on the ACL60/60 English-Chinese, English-German and English-Japanese benchmarks show that our framework consistently improves semantic metrics and achieves lower delay compared to reference translations and salami-based baselines. Notably, combining Drop and Sentence_Cut leads to consistent improvements in the balance between fluency and latency. These results demonstrate that enriching the action space of LLM-based SiMT provides a promising direction for bridging the gap between human and machine interpretation.

同步翻译大模型口译动作空间

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。