arXiv:2504.05122cs.CL2025-04ACL被引 2

用在线框架提升语音翻译的上下文理解能力

DoCIA: An Online Document-Level Context Incorporation Agent for Speech Translation

  • 将文档级上下文分阶段融入语音翻译流程
  • 在4个大模型上显著提升句内与语篇指标
  • 防止过度修正导致幻觉,适合实际部署

文档级上下文对解决文本到文本的文档级机器翻译中的语篇挑战至关重要。尽管自动语音识别(ASR)引入了噪声,增加了语篇复杂性,但语音翻译(ST)中对文档级上下文的整合仍研究不足。本文提出DoCIA,一个在线框架,通过引入文档级上下文提升ST性能。DoCIA将ST流程分解为四个阶段,在ASR优化、机器翻译及翻译优化阶段,通过基于大语言模型(LLM)的辅助模块融入文档级上下文。此外,DoCIA以多层级方式利用文档信息,同时控制计算开销。还引入一种简单有效的判定机制,防止过度优化导致的幻觉,保障最终结果可靠性。实验表明,DoCIA在四个大模型上均显著优于传统ST基线,在句子和语篇层面的指标均有提升,验证了其在提升ST性能上的有效性。

原文摘要 · Abstract (English)

Document-level context is crucial for handling discourse challenges in text-to-text document-level machine translation (MT). Despite the increased discourse challenges introduced by noise from automatic speech recognition (ASR), the integration of document-level context in speech translation (ST) remains insufficiently explored. In this paper, we develop DoCIA, an online framework that enhances ST performance by incorporating document-level context. DoCIA decomposes the ST pipeline into four stages. Document-level context is integrated into the ASR refinement, MT, and MT refinement stages through auxiliary LLM (large language model)-based modules. Furthermore, DoCIA leverages document-level information in a multi-level manner while minimizing computational overhead. Additionally, a simple yet effective determination mechanism is introduced to prevent hallucinations from excessive refinement, ensuring the reliability of the final results. Experimental results show that DoCIA significantly outperforms traditional ST baselines in both sentence and discourse metrics across four LLMs, demonstrating its effectiveness in improving ST performance.

语音翻译上下文融合大模型应用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。