arXiv:2507.04451cs.CV2025-07NeurIPS被引 9

让图像生成像解题一样一步步推理,提升复杂场景的布局准确性。

CoT-lized Diffusion: Let's Reinforce T2I Generation Step-by-step

  • 用多模态大模型在每步去噪时动态调整3D场景布局
  • 复杂场景空间准确率比顶尖方法高34.7%
  • 适合需要精确空间控制的视觉生成任务

当前文本到图像(T2I)生成模型在复杂场景中难以对齐空间结构。即使基于布局的方法也因生成过程与布局规划分离,难以在合成中优化布局。我们提出CoT-Diff框架,将类思维链(CoT)式推理引入T2I生成,通过紧耦合多模态大语言模型(MLLM)驱动的3D布局规划与扩散过程,实现布局感知的实时推理。在每个去噪步骤中,MLLM评估中间生成结果,动态更新3D场景布局,并持续引导生成。更新后的布局转为语义条件与深度图,通过条件感知注意力机制融合进扩散模型,实现精准空间控制与语义注入。在3D场景基准测试中,CoT-Diff显著提升空间对齐与构图保真度,在复杂场景空间准确率上比最先进方法高出34.7%,验证了该紧密耦合生成范式的有效性。

原文摘要 · Abstract (English)

Current text-to-image (T2I) generation models struggle to align spatial composition with the input text, especially in complex scenes. Even layout-based approaches yield suboptimal spatial control, as their generation process is decoupled from layout planning, making it difficult to refine the layout during synthesis. We present CoT-Diff, a framework that brings step-by-step CoT-style reasoning into T2I generation by tightly integrating Multimodal Large Language Model (MLLM)-driven 3D layout planning with the diffusion process. CoT-Diff enables layout-aware reasoning inline within a single diffusion round: at each denoising step, the MLLM evaluates intermediate predictions, dynamically updates the 3D scene layout, and continuously guides the generation process. The updated layout is converted into semantic conditions and depth maps, which are fused into the diffusion model via a condition-aware attention mechanism, enabling precise spatial control and semantic injection. Experiments on 3D Scene benchmarks show that CoT-Diff significantly improves spatial alignment and compositional fidelity, and outperforms the state-of-the-art method by 34.7% in complex scene spatial accuracy, thereby validating the effectiveness of this entangled generation paradigm.

文本生成图像空间布局扩散模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。