提升视觉文本生成能力,让模型准确生成中英文可读文字。
Empowering Backbone Models for Visual Text Generation with Input Granularity Control and Glyph-Aware Training
- 采用混合粒度输入,优化文本表示方式。
- 引入字形感知损失,增强跨注意力学习效果。
- 支持中英文可读文字生成,适合多语言视觉内容创作。
基于扩散的文生图模型在多样性与美学表现上取得显著进展,但在生成可读视觉文字方面仍存在挑战。现有骨干模型普遍存在拼写错误、无法生成文字以及不支持中文等问题,但其发展潜力巨大。本文提出一系列方法,旨在提升骨干模型生成中英文视觉文字的能力。初步研究表明,字节对编码(BPE)分词方式及跨注意力模块学习不足是主要瓶颈。为此,我们设计了混合粒度输入策略,提供更合适的文本表示;同时,提出三种字形感知训练损失,增强跨注意力模块学习并引导模型关注视觉文字。实验表明,所提方法能有效生成语义相关、美观且文字准确的图像,同时保持模型原有的图像生成质量。
原文摘要 · Abstract (English)
Diffusion-based text-to-image models have demonstrated impressive achievements in diversity and aesthetics but struggle to generate images with legible visual texts. Existing backbone models have limitations such as misspelling, failing to generate texts, and lack of support for Chinese text, but their development shows promising potential. In this paper, we propose a series of methods, aiming to empower backbone models to generate visual texts in English and Chinese. We first conduct a preliminary study revealing that Byte Pair Encoding (BPE) tokenization and the insufficient learning of cross-attention modules restrict the performance of the backbone models. Based on these observations, we make the following improvements: (1) We design a mixed granularity input strategy to provide more suitable text representations; (2) We propose to augment the conventional training objective with three glyph-aware training losses, which enhance the learning of cross-attention modules and encourage the model to focus on visual texts. Through experiments, we demonstrate that our methods can effectively empower backbone models to generate semantic relevant, aesthetically appealing, and accurate visual text images, while maintaining their fundamental image generation quality.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。