arXiv:2412.18139cs.CL2024-12被引 1

提升图像内文本翻译的一致性,兼顾语义与风格统一

Ensuring Consistency for In-Image Translation

  • 分两阶段:大语言模型翻译+扩散模型补图
  • 40万伪数据训练,保持文字风格与背景一致
  • 适合需要高质量图像翻译的场景应用

图像内机器翻译任务涉及将嵌入图像中的文本翻译为新语言并以图像形式呈现。尽管该任务在电影海报翻译和日常场景图像翻译中具有广泛应用,但现有方法常忽略一致性问题。本文提出需维护两种一致性:翻译一致性(结合图像信息进行翻译)与图像生成一致性(保持文本图像与原图风格一致,确保背景完整)。为此,我们提出名为HCIIT的两阶段框架:第一阶段使用多模态多语言大模型进行文本图像翻译,并采用思维链学习增强模型对图像信息的利用能力;第二阶段通过训练于风格一致性的扩散模型实现图像补全,确保文字风格统一且保留背景细节。我们构建了一个包含40万组风格一致伪文本图像对的数据集用于模型训练。在自建测试集与真实图像测试集上的结果验证了本框架在保证一致性与生成高质量翻译图像方面的有效性。

原文摘要 · Abstract (English)

The in-image machine translation task involves translating text embedded within images, with the translated results presented in image format. While this task has numerous applications in various scenarios such as film poster translation and everyday scene image translation, existing methods frequently neglect the aspect of consistency throughout this process. We propose the need to uphold two types of consistency in this task: translation consistency and image generation consistency. The former entails incorporating image information during translation, while the latter involves maintaining consistency between the style of the text-image and the original image, ensuring background integrity. To address these consistency requirements, we introduce a novel two-stage framework named HCIIT (High-Consistency In-Image Translation) which involves text-image translation using a multimodal multilingual large language model in the first stage and image backfilling with a diffusion model in the second stage. Chain of thought learning is utilized in the first stage to enhance the model's ability to leverage image information during translation. Subsequently, a diffusion model trained for style-consistent text-image generation ensures uniformity in text style within images and preserves background details. A dataset comprising 400,000 style-consistent pseudo text-image pairs is curated for model training. Results obtained on both curated test sets and authentic image test sets validate the effectiveness of our framework in ensuring consistency and producing high-quality translated images.

图像翻译一致性扩散模型多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。