arXiv:2507.07572cs.CLcs.AI2025-07ACL被引 9

用多模态大模型对齐图文特征,提升文档图像翻译泛化能力

Single-to-mix Modality Alignment with Multimodal Large Language Model for Document Image Machine Translation

  • 通过图像编码器与多模态大模型对齐,学习视觉-文本关联
  • 跨域场景下翻译质量显著提升,尤其在复杂文档上表现优异
  • 推理时无需调用大模型,保持高效同时继承其知识

文档图像机器翻译(DIMT)旨在翻译文档图像中的文字,但受限于训练数据少以及视觉与文本信息间复杂的交互关系,面临泛化挑战。为此,我们提出M4Doc,一种基于多模态大模型(MLLM)的单模态到混合模态对齐框架。该框架将仅含图像的编码器与预训练于大规模文档图像数据集的MLLM的多模态表示进行对齐,使轻量级的DIMT模型在训练中可学习关键的视觉-文本关联。推理阶段,M4Doc无需调用MLLM,既保持计算效率,又受益于其多模态知识。大量实验表明,该方法在翻译质量上实现显著提升,特别是在跨领域泛化和复杂文档图像场景下。

原文摘要 · Abstract (English)

Document Image Machine Translation (DIMT) aims to translate text within document images, facing generalization challenges due to limited training data and the complex interplay between visual and textual information. To address these challenges, we introduce M4Doc, a novel single-to-mix modality alignment framework leveraging Multimodal Large Language Models (MLLMs). M4Doc aligns an image-only encoder with the multimodal representations of an MLLM, pre-trained on large-scale document image datasets. This alignment enables a lightweight DIMT model to learn crucial visual-textual correlations during training. During inference, M4Doc bypasses the MLLM, maintaining computational efficiency while benefiting from its multimodal knowledge. Comprehensive experiments demonstrate substantial improvements in translation quality, especially in cross-domain generalization and challenging document image scenarios.

文档翻译多模态大模型对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。