arXiv:2509.05146cs.CLcs.CV2025-09EMNLP被引 7

构建真实场景多语言图文翻译数据集与模型,解决实际应用中的文本识别难题。

PRIM: Towards Practical In-Image Multilingual Machine Translation

  • 分离处理图像文字与背景信息,提升多语言翻译鲁棒性
  • 在复杂背景、多种字体下实现更高质量的翻译与视觉效果
  • 开源数据集和代码,推动真实世界图文翻译研究

图文翻译(IIMT)旨在将图像中的文字从一种语言翻译为另一种语言。当前端到端IIMT研究主要基于合成数据,存在背景简单、字体单一、文本位置固定、仅支持双语翻译等问题,难以反映真实场景,导致研究与实际应用间存在显著差距。为此,我们提出实用型图文多语言机器翻译(IIMMT),并构建了公开的PRIM数据集,包含真实拍摄的一行文本图像,具有复杂背景、多样字体、多变文本位置,并支持多语言翻译方向。针对这一挑战,我们提出端到端模型VisTrans,分别处理图像中的文本与背景信息,在保证多语言翻译能力的同时提升视觉质量。实验结果表明,VisTrans在翻译质量和视觉表现上均优于现有方法。代码与数据集已公开:https://github.com/BITHLP/PRIM。

原文摘要 · Abstract (English)

In-Image Machine Translation (IIMT) aims to translate images containing texts from one language to another. Current research of end-to-end IIMT mainly conducts on synthetic data, with simple background, single font, fixed text position, and bilingual translation, which can not fully reflect real world, causing a significant gap between the research and practical conditions. To facilitate research of IIMT in real-world scenarios, we explore Practical In-Image Multilingual Machine Translation (IIMMT). In order to convince the lack of publicly available data, we annotate the PRIM dataset, which contains real-world captured one-line text images with complex background, various fonts, diverse text positions, and supports multilingual translation directions. We propose an end-to-end model VisTrans to handle the challenge of practical conditions in PRIM, which processes visual text and background information in the image separately, ensuring the capability of multilingual translation while improving the visual quality. Experimental results indicate the VisTrans achieves a better translation quality and visual effect compared to other models. The code and dataset are available at: https://github.com/BITHLP/PRIM.

图文翻译多语言真实场景

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。