arXiv:2409.17747cs.CV2024-09ICCV

用高资源语言风格生成低资源语言文字图像,提升识别效果

Text Image Generation for Low-Resource Languages with Dual Translation Learning

论文配图:Text Image Generation for Low-Resource Languages with Dual Translation Learning
图 1 · 摘自论文原文
  • 通过双状态扩散模型,将文本转为真实或合成风格图像
  • 生成图像使低资源语言识别准确率显著提升
  • 适合语言资源匮乏场景下的文字识别研究

低资源语言的场景文本识别常因真实场景训练数据不足而受限。本文提出一种新方法,通过模仿高资源语言的真实文本图像风格,生成低资源语言的文本图像。该方法利用一个基于二元状态(‘合成’与‘真实’)条件的扩散模型,执行双重转换任务:将纯文本图像转化为合成或真实风格图像。该机制不仅能有效区分两个域,还促使模型明确识别目标语言字符。为进一步提升生成图像的准确性和多样性,引入两种引导技术:保真度-多样性平衡引导和保真度增强引导。实验表明,本框架生成的文本图像能显著提升低资源语言场景文本识别模型的性能。

原文摘要 · Abstract (English)

Scene text recognition in low-resource languages frequently faces challenges due to the limited availability of training datasets derived from real-world scenes. This study proposes a novel approach that generates text images in low-resource languages by emulating the style of real text images from high-resource languages. Our approach utilizes a diffusion model that is conditioned on binary states: ``synthetic'' and ``real.'' The training of this model involves dual translation tasks, where it transforms plain text images into either synthetic or real text images, based on the binary states. This approach not only effectively differentiates between the two domains but also facilitates the model's explicit recognition of characters in the target language. Furthermore, to enhance the accuracy and variety of generated text images, we introduce two guidance techniques: Fidelity-Diversity Balancing Guidance and Fidelity Enhancement Guidance. Our experimental results demonstrate that the text images generated by our proposed framework can significantly improve the performance of scene text recognition models for low-resource languages.

文本生成低资源语言扩散模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。