提出文本绝缘与注意力机制,提升复杂图文生成的准确性。
Investigating Text Insulation and Attention Mechanisms for Complex Visual Text Generation
- 用强化学习实现多文本绝缘,不增加参数却显著提升渲染效果。
- 引入引号引导注意力门,使每段文本生成更精准。
- 在2000个复杂场景数据集上表现优异,适合小资源团队使用。
本文提出TextCrafter框架,受认知科学中选择性视觉注意启发,引入“文本绝缘-注意力”机制。为实现基于离散对象的选择性注意,提出一种无需额外参数的瓶颈感知约束强化学习方法,显著提升在强预训练模型Qwen-Image上的文本渲染性能。为模拟人类视觉的专注原则,设计面向文本的注意力模块,包含新型引号引导注意力门,进一步提升各文本实例的生成质量。基于强化学习的文本绝缘方法达到当前最优水平,结合文本导向注意力后,在已有强基线基础上获得额外增益。更重要的是,我们构建了包含2000个复杂视觉文本提示的基准测试集CVTG-2K,涵盖位置、数量、长度、属性等多样变化,覆盖多种真实场景。在CVTG-2K、CVTG-Hard、LongText-Bench和Geneval等多个数据集上的大量实验验证了TextCrafter的有效性。尽管仅使用4块GPU,远少于工业级模型(如Qwen-Image、GPT Image、Seedream),TextCrafter仍能更好缓解文本误生成、遗漏和幻觉问题。
原文摘要 · Abstract (English)
In this paper, we present TextCrafter, a Complex Visual Text Generation (CVTG) framework inspired by selective visual attention in cognitive science, and introduce the "Text Insulation-and-Attention" mechanisms. To implement the selective-attention principle that selection operates on discrete objects, we propose a novel Bottleneck-aware Constrained Reinforcement Learning for Multi-text Insulation, which substantially improves text-rendering performance on the strong Qwen-Image pretrained model without introducing additional parameters. To align with the selective concentration principle in human vision, we introduce a text-oriented attention module with a novel Quotation-guided Attention Gate that further improves generation quality for each text instance. Our Reinforcement Learning based text insulation approach attains state-of-the-art results, and incorporating text-oriented attention yields additional gains on top of an already strong baseline. More importantly, we introduce CVTG-2K, a benchmark comprising 2,000 complex visual-text prompts. These prompts vary in positions, quantities, lengths, and attributes, and span diverse real-world scenarios. Extensive evaluations on CVTG-2K, CVTG-Hard, LongText-Bench, and Geneval datasets confirm the effectiveness of TextCrafter. Despite using substantially fewer resources (i.e., 4 GPUs) than industrial-scale models (e.g., Qwen-Image, GPT Image, and Seedream), TextCrafter achieves superior performance in mitigating text misgeneration, omissions, and hallucinations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。