arXiv:2605.15794cs.CL2026-05

构建了15语言对的多模态文档翻译数据集,保持原文排版结构。

ForMaT: Dataset for Visually-Grounded Multilingual PDF Translation

论文配图:ForMaT: Dataset for Visually-Grounded Multilingual PDF Translation
图 1 · 摘自论文原文
  • 基于45个几何特征用K-Medoids采样,确保排版多样性。
  • 现有翻译模型在空间定位和布局同步上表现差,丢失图文关联。
  • 适合研究布局感知翻译、多模态文档重建的团队使用。

我们提出ForMaT(格式保持的多语言翻译),一个包含15种语言对共3,956份PDF的平行语料库,用于多模态机器翻译任务,能保留原始布局元数据。为保证数据集的结构多样性,我们在45个几何特征上采用K-Medoids采样方法,捕捉嵌套表格、公式等复杂元素,仅选取视觉差异显著的PDF文档。评估发现,当前机器翻译系统在空间定位与几何同步方面表现不佳,常丢失文本与其视觉上下文的关联。ForMaT为开发融合视觉与文本信息的布局感知翻译模型提供了基准,助力高保真文档重建。

原文摘要 · Abstract (English)

We present ForMaT (Format-Preserving Multilingual Translation), a parallel corpus of 3,956 PDFs across 15 language pairs that preserves original layout metadata proposed for multimodal machine translation. To ensure structural diversity in the dataset, we employ K-Medoids sampling over 45 geometric features, capturing complex elements like nested tables and formulas to focus only on visually diverse PDF documents. Our evaluation reveals that current MT systems struggle with spatial grounding and geometric synchronization, often losing the link between text and its visual context. ForMaT provides a benchmark for developing layout-aware translation models that integrate visual and textual context for high-fidelity document reconstruction.

多模态翻译文档生成布局感知

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。