构建可复现的联合国多语种平行语料库,规模超7亿英文词元。
UPRPRC: Unified Pipeline for Reproducing Parallel Resources -- Corpus from the United Nations
- 从网页抓取到段落对齐,全流程可复现。
- 产出超7亿英文词元的高质量平行语料,规模翻倍。
- 适合机器翻译研究者使用,内容全为人工翻译。
多语言数据集的质量与可获取性对机器翻译发展至关重要。然而,以往基于联合国文件构建的语料库存在流程不透明、难以复现和规模有限等问题。为此,我们提出一个端到端完整解决方案,涵盖从网络爬取到文本对齐的全过程。整个流程完全可复现,提供单机最小化示例,并支持分布式计算以实现扩展。核心提出一种新的图辅助段落对齐(GAPA)算法,实现高效灵活的段落级对齐。最终构建的语料库包含超过7.13亿个英文词元,规模超过以往工作的两倍。据我们所知,这是目前公开可用的最大纯人工翻译、非AI生成的平行语料库。代码与语料库均以MIT许可证开源。
原文摘要 · Abstract (English)
The quality and accessibility of multilingual datasets are crucial for advancing machine translation. However, previous corpora built from United Nations documents have suffered from issues such as opaque process, difficulty of reproduction, and limited scale. To address these challenges, we introduce a complete end-to-end solution, from data acquisition via web scraping to text alignment. The entire process is fully reproducible, with a minimalist single-machine example and optional distributed computing steps for scalability. At its core, we propose a new Graph-Aided Paragraph Alignment (GAPA) algorithm for efficient and flexible paragraph-level alignment. The resulting corpus contains over 713 million English tokens, more than doubling the scale of prior work. To the best of our knowledge, this represents the largest publicly available parallel corpus composed entirely of human-translated, non-AI-generated content. Our code and corpus are accessible under the MIT License.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。