arXiv:2602.21956cs.CV2026-02被引 2

提出双视角视觉感知框架,提升高分辨率图文翻译准确性

Global-Local Dual Perception for MLLMs in High-Resolution Text-Rich Image Translation

  • 融合全局低分辨率图与多尺度局部文本块,指导模型理解上下文
  • 在51万对图像上训练,翻译完整率和准确率显著优于现有方法
  • 适合需要精准图文翻译的跨语言应用,如多语言界面迁移

文本图像机器翻译(TIMT)旨在将源语言图像中的文本内容翻译为目标语言,需协同视觉感知与语言理解。现有方法在高分辨率、文字密集的图像上表现不佳,常因布局杂乱、字体多样及非文本干扰导致漏译、语义偏移和上下文不一致。为此,我们提出GLoTran,一种基于多模态大模型的全局-局部双视觉感知框架。该框架结合低分辨率全局图像与多尺度区域级文本图像切片,在指令引导对齐策略下,使多模态大模型在保持场景上下文一致性的同时,精确捕捉细粒度文本信息。此外,为支持该范式,我们构建了包含51万对高分辨率全局-局部图像-文本对的GLoD大规模数据集,覆盖多样化真实场景。大量实验表明,GLoTran显著优于当前最优多模态大模型,在翻译完整性和准确性方面均有大幅提升,为高分辨率、文字密集条件下的细粒度图文翻译提供了新范式。

原文摘要 · Abstract (English)

Text Image Machine Translation (TIMT) aims to translate text embedded in images in the source-language into target-language, requiring synergistic integration of visual perception and linguistic understanding. Existing TIMT methods, whether cascaded pipelines or end-to-end multimodal large language models (MLLMs),struggle with high-resolution text-rich images due to cluttered layouts, diverse fonts, and non-textual distractions, resulting in text omission, semantic drift, and contextual inconsistency. To address these challenges, we propose GLoTran, a global-local dual visual perception framework for MLLM-based TIMT. GLoTran integrates a low-resolution global image with multi-scale region-level text image slices under an instruction-guided alignment strategy, conditioning MLLMs to maintain scene-level contextual consistency while faithfully capturing fine-grained textual details. Moreover, to realize this dual-perception paradigm, we construct GLoD, a large-scale text-rich TIMT dataset comprising 510K high-resolution global-local image-text pairs covering diverse real-world scenarios. Extensive experiments demonstrate that GLoTran substantially improves translation completeness and accuracy over state-of-the-art MLLMs, offering a new paradigm for fine-grained TIMT under high-resolution and text-rich conditions.

图文翻译多模态模型高分辨率双视角感知

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。