arXiv:2603.10495cs.CV2026-03被引 2

构建多场景跨模态评估基准,推动图像内文本翻译真实化评测

IMTBench: A Multi-Scenario Cross-Modal Collaborative Evaluation Benchmark for In-Image Machine Translation

  • 设计覆盖四种场景的2500样本基准,支持多维度评估
  • 引入跨模态对齐评分,量化翻译与渲染一致性
  • 揭示现有模型在自然场景和低资源语言中表现显著不足

端到端图像内机器翻译(IIMT)旨在将图像中的文本转换为目标语言,同时保持原始视觉上下文、布局和渲染风格。然而,现有IIMT基准多为合成数据,难以反映真实世界复杂性;且评估协议仅关注单模态指标,忽略翻译结果与图像渲染文本之间的跨模态一致性。为此,我们提出图像内机器翻译基准(IMTBench),包含2500个图像翻译样本,覆盖四个实际应用场景和九种语言。IMTBench支持多方面评估:翻译质量、背景保留、整体图像质量以及跨模态对齐分数,用于衡量模型输出与图像中渲染文本的一致性。我们对主流商业级级联系统及封闭/开源统一多模态模型进行了基准测试,发现不同场景和语言间存在显著性能差距,尤其在自然场景和低资源语言上表现较差,凸显该任务仍有巨大提升空间。期望IMTBench能建立标准化评测体系,推动该新兴任务发展。

原文摘要 · Abstract (English)

End-to-end In-Image Machine Translation (IIMT) aims to convert text embedded within an image into a target language while preserving the original visual context, layout, and rendering style. However, existing IIMT benchmarks are largely synthetic and thus fail to reflect real-world complexity, while current evaluation protocols focus on single-modality metrics and overlook cross-modal faithfulness between rendered text and model outputs. To address these shortcomings, we present In-image Machine Translation Benchmark (IMTBench), a new benchmark of 2,500 image translation samples covering four practical scenarios and nine languages. IMTBench supports multi-aspect evaluation, including translation quality, background preservation, overall image quality, and a cross-modal alignment score that measures consistency between the translated text produced by the model and the text rendered in the translated image. We benchmark strong commercial cascade systems, and both closed- and open-source unified multi-modal models, and observe large performance gaps across scenarios and languages, especially on natural scenes and resource-limited languages, highlighting substantial headroom for end-to-end image text translation. We hope IMTBench establishes a standardized benchmark to accelerate progress in this emerging task.

图像翻译跨模态评估多语言基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。