arXiv:2507.18263cs.CLcs.AI2025-07ACL被引 6

通过定位术语并聚焦翻译知识,提升语音模型术语翻译准确率

Locate-and-Focus: Enhancing Terminology Translation in Speech Language Models

  • 先定位语音中含术语片段,再融合多模态知识增强关注
  • 术语翻译成功率显著提升,整体翻译性能保持稳定
  • 适合需要精准术语翻译的场景,如医疗、法律语音转写

直接语音翻译(ST)近年来受到广泛关注,但话语中术语的准确翻译仍是重大挑战。现有方法主要依赖引入各类翻译知识,却常受无关噪声干扰,且未能充分使用翻译知识。为此,本文提出一种新的「定位-聚焦」方法:首先有效定位语句中包含术语的语音片段,构建翻译知识,减少对ST模型的无关信息干扰;随后在音频与文本双模态下,将翻译知识与语句及翻译假设关联,使ST模型在翻译过程中更聚焦于相关知识。跨多个数据集的实验结果表明,该方法能有效定位术语并显著提升术语翻译成功率,同时保持稳健的整体翻译性能。

原文摘要 · Abstract (English)

Direct speech translation (ST) has garnered increasing attention nowadays, yet the accurate translation of terminology within utterances remains a great challenge. In this regard, current studies mainly concentrate on leveraging various translation knowledge into ST models. However, these methods often struggle with interference from irrelevant noise and can not fully utilize the translation knowledge. To address these issues, in this paper, we propose a novel Locate-and-Focus method for terminology translation. It first effectively locates the speech clips containing terminologies within the utterance to construct translation knowledge, minimizing irrelevant information for the ST model. Subsequently, it associates the translation knowledge with the utterance and hypothesis from both audio and textual modalities, allowing the ST model to better focus on translation knowledge during translation. Experimental results across various datasets demonstrate that our method effectively locates terminologies within utterances and enhances the success rate of terminology translation, while maintaining robust general translation performance.

语音翻译术语翻译多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。