arXiv:2607.16409cs.CVcs.AI2026-07被引 1

让AI像人一样思考布局,精准生成符合空间指令的图像

Think, Plan, Paint: Layout-Aware Reasoning for Controllable Image Generation in Unified Models

论文配图:Think, Plan, Paint: Layout-Aware Reasoning for Controllable Image Generation in Unified Models
图 1 · 摘自论文原文
  • 引入'思考-规划-作画'框架,用布局统一三阶段
  • 在基准上提升65.31%,空间任务平均增23.06%
  • 支持复杂指令生成与编辑,适合需要精确控制的场景

统一多模态大语言模型(MLLM)虽有望融合视觉理解与生成,但在遵循复杂空间指令和逻辑约束方面仍存在不足。为此,我们提出ATLAS框架,赋予MLLM类人的‘思考、规划、作画’范式。通过将布局作为三阶段共享表示,使模型能推理空间需求、规划显式对象排列并生成图像。进一步采用基于强化学习的布局对齐提升图像生成质量。我们在7B和80B规模上实现该框架,在图像生成基准上达到领先性能,较现有布局基线平均提升65.31%;在空间相关任务中较基础模型平均提升23.06%。通过同一布局接口,还支持指令引导编辑与多模态定位。我们还提出了ATLAS-Reasoning基准,用于评估复杂空间指令下的生成能力。

原文摘要 · Abstract (English)

Unified Multimodal Large Language Models (MLLMs) offer a promising paradigm for unifying visual understanding and generation, yet they still struggle to follow complex spatial instructions and logical constraints in controllable image generation. To address this gap, we present ATLAS, a unified framework that equips MLLMs with a human-like "Think, Plan, and Paint" paradigm. We adopt layout as the shared representation that connects the three stages, enabling the model to reason about spatial requirements, plan explicit object arrangements, and render the final image. We further improve plan-to-image fidelity with reinforcement-learning-based layout alignment. We instantiate ATLAS at 7B and 80B scales, achieving state-of-the-art performance among MLLMs on image generation benchmarks and an average 65.31% improvement over existing layout-based unified MLLMs. On spatially related tasks, ATLAS obtains an average 23.06% gain over the base models. Through the same layout interface, ATLAS also supports instruction-guided editing and multimodal grounding. We further introduce ATLAS-Reasoning, a benchmark for evaluating generation under complex spatial instructions.

图像生成空间推理多模态布局控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。