arXiv:2501.08523cs.CLcs.AI2025-01被引 8

提出Doc-Guided Sent2Sent++,解决文档级翻译中句子遗漏与连贯性难题。

Doc-Guided Sent2Sent++: A Sent2Sent++ Agent with Doc-Guided memory for Document-level Machine Translation

  • 采用逐句强制解码策略,确保每句必译且相邻句更流畅。
  • 在多语言多领域测试中,显著提升s-COMET、d-COMET等指标。
  • 适合需高一致性与流畅性的跨语言文档翻译任务。

人工智能在自然语言处理领域的进展主要得益于大语言模型(LLMs)的能力。这些模型支撑了应对长上下文依赖的智能体,尤其在文档级机器翻译(DocMT)中表现突出。DocMT的关键评价指标为质量、一致性和流畅性。现有方法如Doc2Doc和Doc2Sent或忽略句子,或牺牲流畅性。本文提出Doc-Guided Sent2Sent++,采用增量式逐句强制解码策略,确保每句被翻译并提升相邻句的连贯性。该智能体利用仅关注摘要及其翻译的文档引导记忆机制,有效保持一致性。在多语言、多领域广泛测试中,Sent2Sent++在质量、一致性和流畅性方面均优于其他方法。结果显示,其在s-COMET、d-COMET、LTCR-$1_f$及文档级困惑度(d-ppl)等指标上均有显著提升。本文贡献包括对当前DocMT研究的深入分析、引入Sent2Sent++解码方法、提出文档引导记忆机制,并验证其跨语言与跨领域的有效性。

原文摘要 · Abstract (English)

The field of artificial intelligence has witnessed significant advancements in natural language processing, largely attributed to the capabilities of Large Language Models (LLMs). These models form the backbone of Agents designed to address long-context dependencies, particularly in Document-level Machine Translation (DocMT). DocMT presents unique challenges, with quality, consistency, and fluency being the key metrics for evaluation. Existing approaches, such as Doc2Doc and Doc2Sent, either omit sentences or compromise fluency. This paper introduces Doc-Guided Sent2Sent++, an Agent that employs an incremental sentence-level forced decoding strategy \textbf{to ensure every sentence is translated while enhancing the fluency of adjacent sentences.} Our Agent leverages a Doc-Guided Memory, focusing solely on the summary and its translation, which we find to be an efficient approach to maintaining consistency. Through extensive testing across multiple languages and domains, we demonstrate that Sent2Sent++ outperforms other methods in terms of quality, consistency, and fluency. The results indicate that, our approach has achieved significant improvements in metrics such as s-COMET, d-COMET, LTCR-$1_f$, and document-level perplexity (d-ppl). The contributions of this paper include a detailed analysis of current DocMT research, the introduction of the Sent2Sent++ decoding method, the Doc-Guided Memory mechanism, and validation of its effectiveness across languages and domains.

文档翻译大模型智能体一致性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。