arXiv:2503.23461cs.CV2025-03被引 18

提出文本绝缘与注意力机制,提升复杂图文生成的准确性。

Investigating Text Insulation and Attention Mechanisms for Complex Visual Text Generation

  • 用强化学习实现多文本绝缘,不增加参数却显著提升渲染效果。
  • 引入引号引导注意力门,使每段文本生成更精准。
  • 在2000个复杂场景数据集上表现优异,适合小资源团队使用。

本文提出TextCrafter框架,受认知科学中选择性视觉注意启发,引入“文本绝缘-注意力”机制。为实现基于离散对象的选择性注意,提出一种无需额外参数的瓶颈感知约束强化学习方法,显著提升在强预训练模型Qwen-Image上的文本渲染性能。为模拟人类视觉的专注原则,设计面向文本的注意力模块,包含新型引号引导注意力门,进一步提升各文本实例的生成质量。基于强化学习的文本绝缘方法达到当前最优水平,结合文本导向注意力后,在已有强基线基础上获得额外增益。更重要的是,我们构建了包含2000个复杂视觉文本提示的基准测试集CVTG-2K,涵盖位置、数量、长度、属性等多样变化,覆盖多种真实场景。在CVTG-2K、CVTG-Hard、LongText-Bench和Geneval等多个数据集上的大量实验验证了TextCrafter的有效性。尽管仅使用4块GPU,远少于工业级模型(如Qwen-Image、GPT Image、Seedream),TextCrafter仍能更好缓解文本误生成、遗漏和幻觉问题。

原文摘要 · Abstract (English)

In this paper, we present TextCrafter, a Complex Visual Text Generation (CVTG) framework inspired by selective visual attention in cognitive science, and introduce the "Text Insulation-and-Attention" mechanisms. To implement the selective-attention principle that selection operates on discrete objects, we propose a novel Bottleneck-aware Constrained Reinforcement Learning for Multi-text Insulation, which substantially improves text-rendering performance on the strong Qwen-Image pretrained model without introducing additional parameters. To align with the selective concentration principle in human vision, we introduce a text-oriented attention module with a novel Quotation-guided Attention Gate that further improves generation quality for each text instance. Our Reinforcement Learning based text insulation approach attains state-of-the-art results, and incorporating text-oriented attention yields additional gains on top of an already strong baseline. More importantly, we introduce CVTG-2K, a benchmark comprising 2,000 complex visual-text prompts. These prompts vary in positions, quantities, lengths, and attributes, and span diverse real-world scenarios. Extensive evaluations on CVTG-2K, CVTG-Hard, LongText-Bench, and Geneval datasets confirm the effectiveness of TextCrafter. Despite using substantially fewer resources (i.e., 4 GPUs) than industrial-scale models (e.g., Qwen-Image, GPT Image, and Seedream), TextCrafter achieves superior performance in mitigating text misgeneration, omissions, and hallucinations.

图文生成注意力机制文本绝缘小模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。