arXiv:2509.12278cs.CVcs.AI2025-09EMNLP被引 1

构建多场景定位感知图文翻译基准,支持精准布局保留的翻译。

PATIMT-Bench: A Multi-Scenario Benchmark for Position-Aware Text Image Machine Translation in Large Vision-Language Models

  • 提出定位感知图文翻译新任务,兼顾区域翻译与全图定位。
  • 构建含1200个专家标注实例的基准数据集,覆盖10类真实场景。
  • 适配不同场景的OCR优化流程,提升模型在复杂图像上的表现。

图文机器翻译(TIMT)旨在将图像中的文字内容翻译为另一种语言。现有研究多关注图像中所有文本的整体翻译,忽略边界框信息,且场景覆盖有限。本文将传统TIMT拓展为定位感知图文翻译(PATIMT),旨在实现细粒度、布局保留的翻译,具有重要实际价值但尚未充分探索。该任务包含两个关键子任务:区域特定翻译和带定位的全图翻译。为支持模型在该任务上的评估,我们构建了PATIMT基准(PATIMTBench),涵盖10种真实世界场景。我们设计自适应图像OCR精炼管道,根据场景动态选择合适的OCR工具并优化文本丰富图像的结果。为确保评估可靠性,构建包含1,200个高质量人工标注与专家审核的测试集。在该数据上微调后,紧凑型大型视觉-语言模型(LVLMs)在两项子任务上均达到当前最优性能。实验还验证了训练数据的可扩展性与泛化能力。

原文摘要 · Abstract (English)

Text Image Machine Translation (TIMT) aims to translate texts embedded within an image into another language. Current TIMT studies primarily focus on providing translations for all the text within an image, while neglecting to provide bounding boxes and covering limited scenarios. In this work, we extend traditional TIMT into position-aware TIMT (PATIMT), aiming to support fine-grained and layoutpreserving translation, which holds great practical value but remains largely unexplored. This task comprises two key sub-tasks: regionspecific translation and full-image translation with grounding. To support existing models on PATIMT and conduct fair evaluation, we construct the PATIMT benchmark (PATIMTBench), which consists of 10 diverse real-world scenarios. Specifically, we introduce an Adaptive Image OCR Refinement Pipeline, which adaptively selects appropriate OCR tools based on scenario and refines the results of text-rich images. To ensure evaluation reliability, we further construct a test set, which contains 1,200 high-quality instances manually annotated and reviewed by human experts. After fine-tuning on our data, compact Large Vision-Language Models (LVLMs) achieve state-of-the-art performance on both sub-tasks. Experimental results also highlight the scalability and generalizability of our training data

图文翻译视觉语言模型基准测试定位感知

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。