构建多语言多场景图文翻译评测集,提升跨语言视觉理解能力
MMTIT-Bench: A Multilingual and Multi-Scenario Benchmark with Cognition-Perception-Reasoning Guided Text-Image Machine Translation
- 设计认知-感知-推理统一框架,增强图文翻译逻辑性
- 覆盖14种非中英文、1400张图像,涵盖文档/实景/网页等场景
- 适用于多语言视觉理解、跨模态翻译研究者
端到端图文机器翻译(TIMT)直接将图像中的文本内容跨语言转换,在多语言场景理解中至关重要。尽管视觉语言大模型(VLLMs)取得进展,但其在多样化视觉场景和低资源语言下的鲁棒性仍因评估资源有限而未充分探索。我们提出MMTIT-Bench,一个经人工验证的多语言多场景基准,包含1400张图像,覆盖14种非中文和非英语语言,涵盖文档、实景与网页图像等多种场景,支持对端到端TIMT的严格评估。此外,我们研究了以推理为导向的数据设计如何提升翻译效果。现有方法或顺序分解解析与翻译,或仅关注语言推理,忽视了视觉认知在VLLMs中的核心作用。为此,我们提出CPR-Trans:一种将场景认知、文本感知与翻译推理融合的统一推理范式。通过基于VLLM的数据生成流程,CPR-Trans提供结构化、可解释的监督信号,使感知与推理对齐。在3B和7B模型上的实验显示,准确率与可解释性均有稳定提升。论文接受后将公开MMTIT-Bench,推动多语言多场景TIMT研究。
原文摘要 · Abstract (English)
End-to-end text-image machine translation (TIMT), which directly translates textual content in images across languages, is crucial for real-world multilingual scene understanding. Despite advances in vision-language large models (VLLMs), robustness across diverse visual scenes and low-resource languages remains underexplored due to limited evaluation resources. We present MMTIT-Bench, a human-verified multilingual and multi-scenario benchmark with 1,400 images spanning fourteen non-English and non-Chinese languages and diverse settings such as documents, scenes, and web images, enabling rigorous assessment of end-to-end TIMT. Beyond benchmarking, we study how reasoning-oriented data design improves translation. Although recent VLLMs have begun to incorporate long Chain-of-Thought (CoT) reasoning, effective thinking paradigms for TIMT are still immature: existing designs either cascade parsing and translation in a sequential manner or focus on language-only reasoning, overlooking the visual cognition central to VLLMs. We propose Cognition-Perception-Reasoning for Translation (CPR-Trans), a data paradigm that integrates scene cognition, text perception, and translation reasoning within a unified reasoning process. Using a VLLM-driven data generation pipeline, CPR-Trans provides structured, interpretable supervision that aligns perception with reasoning. Experiments on 3B and 7B models show consistent gains in accuracy and interpretability. We will release MMTIT-Bench to promote the multilingual and multi-scenario TIMT research upon acceptance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。