arXiv:2606.20212cs.CL2026-06

构建多语言文档平行数据集,支持格式保持的机器翻译评估

CzechDocs: A Multiway Parallel Dataset of Formatted Documents for Minority Languages in Czechia

  • 收集捷克语及少数语言的网页、文档和PDF格式平行文本
  • 验证多种格式保持翻译方法,提供可复现的评估基准
  • 适合研究文档级翻译与格式保留的团队使用

我们提出CzechDocs,一个涵盖捷克语及捷克境内主要少数语言(乌克兰语、英语)的多语言平行文档数据集,包含HTML、DOCX和PDF三种格式。该数据集旨在支持机器翻译系统在翻译过程中保持原始文档格式的评估。我们对主流格式保持翻译方法在数据集的一个验证子集上进行了对比分析,并公开发布该验证集及配套评估工具包以促进后续研究。一个保留用于未来文档级翻译共享任务的测试集将被预留。

原文摘要 · Abstract (English)

We present CzechDocs, a multiway parallel dataset of formatted documents (HTML, DOCX, and PDF) covering Czech and minority languages used in Czechia-primarily Ukrainian and English, with smaller portions of Vietnamese, Russian and other languages. The dataset is designed to support the evaluation of machine translation systems that aim to preserve document formatting during translation. We provide a comparison of the most common approaches to format-preserving machine translation on a validation subset of the dataset. This validation split, together with the evaluation toolkit, is publicly released for further research. A held-out test split will be reserved for a future shared task focused on document-level translation with formatting preservation.

机器翻译多语言文档生成格式保持

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。