arXiv:2602.11980cs.CV2026-02被引 2

用空间思维链让扩散模型更懂空间推理,生成更准。

Spatial Chain-of-Thought: Bridging Understanding and Generation Models for Spatial Reasoning Generation

  • 用文本坐标混合指令训练扩散模型,增强布局感知。
  • 调用先进多模态大模型做空间规划,直接传递给生成过程。
  • 在复杂推理和图像编辑上表现优异,适合需要精准布局的生成任务。

尽管扩散模型在美学图像合成方面表现出色,但在复杂空间理解与推理方面仍存在不足。现有方法依赖多模态大语言模型(MLLM)提升该能力,但要么因联合训练带来高计算开销,要么仅依赖文本提示导致空间信息丢失。为此,我们提出一种即插即用的空间思维链(SCoT)框架,有效连接MLLM的推理能力与扩散模型的生成能力。具体而言,我们首先通过交错文本-坐标指令格式训练扩散模型,增强其布局感知;随后利用先进的MLLM作为规划器生成全面的布局计划,并将空间规划能力直接迁移至生成流程。大量实验表明,该方法在图像生成基准上达到最先进性能,在复杂推理任务中显著优于基线,同时在图像编辑场景中也展现出强有效性。

原文摘要 · Abstract (English)

While diffusion models have shown exceptional capabilities in aesthetic image synthesis, they often struggle with complex spatial understanding and reasoning. Existing approaches resort to Multimodal Large Language Models (MLLMs) to enhance this capability. However, they either incur high computational costs through joint training or suffer from spatial information loss when relying solely on textual prompts. To alleviate these limitations, we propose a Spatial Chain-of-Thought (SCoT) framework, a plug-and-play approach that effectively bridges the reasoning capabilities of MLLMs with the generative power of diffusion models. Specifically, we first enhance the diffusion model's layout awareness by training it on an interleaved text-coordinate instruction format. We then leverage state-of-the-art MLLMs as planners to generate comprehensive layout plans, transferring their spatial planning capabilities directly to the generation process. Extensive experiments demonstrate that our method achieves state-of-the-art performance on image generation benchmarks and significantly outperforms baselines on complex reasoning tasks, while also showing strong efficacy in image editing scenarios.

空间推理扩散模型多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。