arXiv:2509.12046cs.CVcs.AI2025-09被引 3

通过结构化掩码实现布局条件下的文本到图像自回归生成

Layout-Conditioned Autoregressive Text-to-Image Generation via Structured Masking

  • 用结构化掩码控制注意力机制,避免区域描述混淆
  • 在多个数据集上实现更高布局准确率与图像质量
  • 适合需要精确布局控制的生成任务开发者

尽管自回归(AR)模型在图像生成中表现优异,但将其扩展到布局条件生成仍面临布局信号稀疏和特征纠缠的问题。本文提出SMARLI框架,通过结构化掩码策略在注意力计算中整合空间布局约束,有效控制全局提示、布局与图像标记之间的交互,防止区域与描述错配,并充分注入布局信息。为缓解AR模型的暴露偏差并提升生成质量与布局准确性,引入基于组相对策略优化(GRPO)的后训练方案,设计专用布局奖励与图像质量奖励协同优化。实验表明,SMARLI能无缝融合布局、文本与图像标记,不损失生成质量,且该掩码策略与后训练方法可迁移至标准下一项预测型AR模型。所提框架在保持AR模型结构简洁与高效的同时,实现了更优的布局控制能力。

原文摘要 · Abstract (English)

Although autoregressive (AR) models have demonstrated remarkable success in image generation, extending these models to layout-conditioned generation remains challenging due to the sparse nature of layout conditions and the risk of feature entanglement. We present \textbf{S}tructured \textbf{M}asking for \textbf{AR}-based \textbf{L}ayout-to-\textbf{I}mage (SMARLI), a novel framework that effectively integrates spatial layout constraints into the AR generation process. To equip AR models with layout control, a structured masking strategy is applied to the attention computation to govern the interaction among the global prompt, layout, and image tokens. This design prevents the misassociation of different regions with their corresponding descriptions while enabling the sufficient injection of layout constraints into the generation process. To alleviate the exposure bias of AR models and further enhance generation quality and layout accuracy, we incorporate a Group Relative Policy Optimization (GRPO) post-training scheme. We adapt it to the next-set-based paradigm and introduce a specifically designed layout reward, which is coordinated with an image quality reward to guide policy optimization in a balanced manner. Experimental results demonstrate that SMARLI seamlessly integrates layout tokens with text and image tokens without compromising generation quality, and the proposed masking strategy and post-training scheme can also be transferred to standard next-token-based AR models. The proposed framework achieves superior layout control while maintaining the structural simplicity and generation efficiency of AR models.

图像生成自回归模型布局控制注意力机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。