arXiv:2504.12140cs.CL2025-04被引 10

通过高质量文档数据微调,提升大模型长文档翻译能力。

Multilingual Contextualization of Large Language Models for Document-Level Machine Translation

  • 用自建文档数据集DocBlocks进行针对性微调,增强长程依赖建模。
  • 支持多范式翻译,文档级与分块翻译结合,提升翻译质量与速度。
  • 适合需要高精度长文本翻译的研究者和工业应用开发者。

大语言模型在句级机器翻译中表现优异,但扩展到文档级翻译仍面临挑战,尤其在建模跨句子和段落的长程依赖与语篇现象方面。本文提出一种方法,通过在高质量文档级数据上进行目标微调来提升基于大模型的长文档翻译性能,这些数据由我们构建并引入为DocBlocks。该方法支持多种翻译范式,包括直接文档对文档翻译和带上下文的分块级翻译,通过集成有无上下文的指令,使模型更好地捕捉跨句依赖,同时保持强句级翻译性能。实验表明,结合多种翻译范式相比提示工程和代理方法,能显著提升文档级翻译质量和推理速度。

原文摘要 · Abstract (English)

Large language models (LLMs) have demonstrated strong performance in sentence-level machine translation, but scaling to document-level translation remains challenging, particularly in modeling long-range dependencies and discourse phenomena across sentences and paragraphs. In this work, we propose a method to improve LLM-based long-document translation through targeted fine-tuning on high-quality document-level data, which we curate and introduce as DocBlocks. Our approach supports multiple translation paradigms, including direct document-to-document and chunk-level translation, by integrating instructions both with and without surrounding context. This enables models to better capture cross-sentence dependencies while maintaining strong sentence-level translation performance. Experimental results show that incorporating multiple translation paradigms improves document-level translation quality and inference speed compared to prompting and agent-based methods.

文档翻译大模型长程依赖多范式

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。