arXiv:2508.13079cs.CL2025-08被引 6

构建了全球最大多语言文档级翻译数据集,助力低资源语言翻译性能提升。

DocHPLT: A Massively Multilingual Document-Level Translation Dataset

  • 基于网页抓取优化管道,完整保留文档原始内容与结构。
  • 覆盖50种语言,含124万对文档,共42.6亿句,支持非英语配对。
  • 微调后大模型显著超越通用基线,尤其提升低资源语言表现。

现有文档级机器翻译资源仅限少数高资源语言。为推动全球社区在文档级翻译及长文本建模方面的研究,我们构建了迄今为止最大的公开文档级翻译数据集DocHPLT,包含50种语言与英语的1240万组对齐文档,共42.6亿句子。通过引入桥梁对齐,可额外获得2500组不涉及英语的文档对。不同于以往从句级数据重构文档的方法,我们改进现有网络抽取流程,确保源文档完整性,保留全部内容(包括未对齐部分)。初步实验确定最优训练上下文策略后,我们证明在DocHPLT上微调的大模型显著优于现成指令微调基线,尤其在低资源语言上提升明显。数据集已开源,采用宽松许可协议,为多语言文档级翻译研究提供关键基础设施。

原文摘要 · Abstract (English)

Existing document-level machine translation resources are only available for a handful of languages, mostly high-resourced ones. To facilitate the training and evaluation of document-level translation and, more broadly, long-context modeling for global communities, we create DocHPLT, the largest publicly available document-level translation dataset to date. It contains 124 million aligned document pairs across 50 languages paired with English, comprising 4.26 billion sentences. By adding pivoted alignments, practitioners can obtain 2500 additional pairs not involving English. Unlike previous reconstruction-based approaches that piece together documents from sentence-level data, we modify an existing web extraction pipeline to preserve complete document integrity from the source, retaining all content, including unaligned portions. After our preliminary experiments identify the optimal training context strategy for document-level translation, we demonstrate that LLMs fine-tuned on DocHPLT substantially outperform off-the-shelf instruction-tuned baselines, with particularly dramatic improvements for under-resourced languages. We open-source the dataset under a permissive license, providing essential infrastructure for advancing multilingual document-level translation.

文档翻译多语言大模型训练数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。