arXiv:2501.02962cs.CV2025-01

生成逼真可控制的多语言场景文字,提升OCR训练效果

SceneVTG++: Controllable Multilingual Visual Text Generation in the Wild

  • 分两阶段生成:先定位合理文本区域并推荐内容,再用扩散模型生成可控文字
  • 在自然场景中生成的文本准确率高、与场景相关,且显著提升OCR任务性能
  • 支持多语言、字体颜色可调,适合需要真实场景文本数据的研究者

在自然场景图像中生成视觉文字是一项极具挑战性的任务,面临诸多未解难题。与在人工设计图像(如海报、封面、卡通等)上生成文字不同,自然场景中的文字需满足四个关键标准:(1) 真实性:生成文字应如照片般逼真,笔画无误;(2) 合理性:文字应位于合理载体(如招牌、墙壁等)上,内容与场景相关;(3) 实用性:生成图像有助于自然场景光学字符识别(OCR)任务训练;(4) 可控性:文字属性(如字体、颜色)可按需控制。本文提出两阶段方法 SceneVTG++,同时满足上述四点要求。SceneVTG++ 包含文本布局与内容生成器(TLCG)和可控局部文本扩散模型(CLTD)。前者利用多模态大语言模型的世界知识,根据场景背景图像识别合理文本位置并推荐内容;后者基于扩散模型生成可控的多语言文字。通过大量实验,分别验证了 TLCG 与 CLTD 的有效性,并展示了 SceneVTG++ 在文本生成上的最先进性能。此外,生成图像在文本检测与识别等 OCR 任务中表现出优异实用性。代码与数据集将公开。

原文摘要 · Abstract (English)

Generating visual text in natural scene images is a challenging task with many unsolved problems. Different from generating text on artificially designed images (such as posters, covers, cartoons, etc.), the text in natural scene images needs to meet the following four key criteria: (1) Fidelity: the generated text should appear as realistic as a photograph and be completely accurate, with no errors in any of the strokes. (2) Reasonability: the text should be generated on reasonable carrier areas (such as boards, signs, walls, etc.), and the generated text content should also be relevant to the scene. (3) Utility: the generated text can facilitate to the training of natural scene OCR (Optical Character Recognition) tasks. (4) Controllability: The attribute of the text (such as font and color) should be controllable as needed. In this paper, we propose a two stage method, SceneVTG++, which simultaneously satisfies the four aspects mentioned above. SceneVTG++ consists of a Text Layout and Content Generator (TLCG) and a Controllable Local Text Diffusion (CLTD). The former utilizes the world knowledge of multi modal large language models to find reasonable text areas and recommend text content according to the nature scene background images, while the latter generates controllable multilingual text based on the diffusion model. Through extensive experiments, we respectively verified the effectiveness of TLCG and CLTD, and demonstrated the state-of-the-art text generation performance of SceneVTG++. In addition, the generated images have superior utility in OCR tasks like text detection and text recognition. Codes and datasets will be available.

视觉文本生成多语言可控生成OCR增强

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。