arXiv:2508.10382cs.CV2025-08AAAI被引 1

让扩散模型同时生成图像和场景结构,提升空间一致性。

Towards Spatially Consistent Image Generation: On Incorporating Intrinsic Scene Properties into Diffusion Models

  • 联合生成图像与深度、分割等内在场景属性
  • 在保持图像质量的前提下纠正布局失真
  • 适合需要真实场景布局的生成任务

大规模数据训练的图像生成模型虽能合成高质量图像,但常因缺乏底层结构与空间布局信息而产生空间不一致和扭曲。本文利用深度、分割图等内在场景属性,这些属性提供丰富的场景信息,区别于仅依赖图像-文本对或将其作为条件输入的先前方法。我们的方法旨在共同生成图像及其对应的内在属性,使模型隐式捕捉场景结构,生成更空间一致且真实的图像。具体地,首先使用预训练估计器从大规模图像数据集中提取丰富的内在属性,无需额外场景信息或显式3D表示;随后通过自编码器将多种内在属性聚合为单一潜在变量;基于预训练的大规模潜在扩散模型(LDMs),该方法通过共享互信息,同时去噪图像与内在域,使两者相互反映而不降低图像质量。实验表明,该方法有效修正空间不一致性,生成更自然的场景布局,同时保持基线模型(如Stable Diffusion)的保真度与文本对齐性。

原文摘要 · Abstract (English)

Image generation models trained on large datasets can synthesize high-quality images but often produce spatially inconsistent and distorted images due to limited information about the underlying structures and spatial layouts. In this work, we leverage intrinsic scene properties (e.g., depth, segmentation maps) that provide rich information about the underlying scene, unlike prior approaches that solely rely on image-text pairs or use intrinsics as conditional inputs. Our approach aims to co-generate both images and their corresponding intrinsics, enabling the model to implicitly capture the underlying scene structure and generate more spatially consistent and realistic images. Specifically, we first extract rich intrinsic scene properties from a large image dataset with pre-trained estimators, eliminating the need for additional scene information or explicit 3D representations. We then aggregate various intrinsic scene properties into a single latent variable using an autoencoder. Building upon pre-trained large-scale Latent Diffusion Models (LDMs), our method simultaneously denoises the image and intrinsic domains by carefully sharing mutual information so that the image and intrinsic reflect each other without degrading image quality. Experimental results demonstrate that our method corrects spatial inconsistencies and produces a more natural layout of scenes while maintaining the fidelity and textual alignment of the base model (e.g., Stable Diffusion).

扩散模型空间一致性场景理解

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。