arXiv:2603.09392cs.CVcs.AI2026-03中稿 · ICDAR 2025被引 2

2025年文档图像翻译竞赛,挑战复杂版面的端到端跨语言翻译。

ICDAR 2025 Competition on End-to-End Document Image Machine Translation Towards Complex Layouts

  • 分无OCR与有OCR两赛道,支持大中小模型统一系统参赛
  • 共收27份有效提交,大模型在复杂布局翻译中表现更优
  • 适合多模态理解、文档智能与机器翻译研究者参考

文档图像机器翻译(DIMT)旨在联合建模文本内容与页面布局,实现对文档图像中文字的跨语言翻译,融合光学字符识别(OCR)与自然语言处理(NLP)。ICDAR 2025 DIMT竞赛推动端到端文档图像翻译研究,该领域在多模态文档理解中快速发展。竞赛设无OCR与有OCR两个赛道,每赛道含小于10亿参数和大于10亿参数的子任务。参赛者需提交单一统一的DIMT系统,可选择使用提供的OCR转录文本。比赛于2024年12月10日至2025年4月20日举行,共吸引69支队伍,提交27份有效成果。其中,第一赛道34支队伍提交13份有效结果,第二赛道35支队伍提交14份有效结果。本文介绍竞赛动机、数据集构建、任务定义、评估协议,并总结结果。分析表明,大模型方法在复杂版面文档图像翻译中展现出新范式潜力,为未来研究留下广阔空间。

原文摘要 · Abstract (English)

Document Image Machine Translation (DIMT) seeks to translate text embedded in document images from one language to another by jointly modeling both textual content and page layout, bridging optical character recognition (OCR) and natural language processing (NLP). The DIMT 2025 Challenge advances research on end-to-end document image translation, a rapidly evolving area within multimodal document understanding. The competition features two tracks, OCR-free and OCR-based, each with two subtasks for small (less than 1B parameters) and large (greater than 1B parameters) models. Participants submit a single unified DIMT system, with the option to incorporate provided OCR transcripts. Running from December 10, 2024 to April 20, 2025, the competition attracted 69 teams and 27 valid submissions in total. Track 1 had 34 teams and 13 valid submissions, while Track 2 had 35 teams and 14 valid submissions. In this report, we present the challenge motivation, dataset construction, task definitions, evaluation protocol, and a summary of results. Our analysis shows that large-model approaches establish a promising new paradigm for translating complex-layout document images and highlight substantial opportunities for future research.

文档翻译多模态端到端ICDAR

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。