arXiv:2603.02712cs.CVcs.MM2026-03

让AI画画先想清楚怎么画,再决定画什么。

From "What" to "How": Constrained Reasoning for Autoregressive Image Generation

  • 用视觉约束推理生成步骤,解决图像结构混乱问题。
  • 在T2I-CompBench上空间准确率提升5.41%。
  • 适合需要精确布局的图像生成任务。

自回归图像生成近期通过思维链和强化学习获得进展,但现有方法仅重写输入提示来指定“画什么”,未能深入思考“怎么画”以构建整体结构。这一根本缺陷导致空间模糊,引发物体重叠等不真实现象。为此,我们提出CoR-Painter框架,开创“如何画→画什么”的新范式,引入约束推理机制:首先从提示中推导出一组视觉约束,明确空间关系、关键属性与构图规则;这些约束引导后续生成详细描述“画什么”,为精准视觉合成提供结构化基础。此外,我们设计双目标GRPO策略,专门优化文本推理与视觉映射过程,确保整个生成流程的连贯性与质量。在T2I-CompBench、GenEval和WISE上的大量实验表明,本方法达到领先性能,空间指标显著提升(如在T2I-CompBench上+5.41%)。

原文摘要 · Abstract (English)

Autoregressive image generation has seen recent improvements with the introduction of chain-of-thought and reinforcement learning. However, current methods merely specify "What" details to depict by rewriting the input prompt, yet fundamentally fail to reason about "How" to structure the overall image. This inherent limitation gives rise to persistent issues, such as spatial ambiguity directly causing unrealistic object overlaps. To bridge this gap, we propose CoR-Painter, a novel framework that pioneers a "How-to-What" paradigm by introducing Constrained Reasoning to guide the autoregressive generation. Specifically, it first deduces "How to draw" by deriving a set of visual constraints from the input prompt, which explicitly govern spatial relationships, key attributes, and compositional rules. These constraints steer the subsequent generation of a detailed description "What to draw", providing a structurally sound and coherent basis for accurate visual synthesis. Additionally, we introduce a Dual-Objective GRPO strategy that specifically optimizes the textual constrained reasoning and visual projection processes to ensure the coherence and quality of the entire generation pipeline. Extensive experiments on T2I-CompBench, GenEval, and WISE demonstrate that our method achieves state-of-the-art performance, with significant improvements in spatial metrics (e.g., +5.41% on T2I-CompBench).

图像生成约束推理自回归结构控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。