探究扩散模型生成多物体时的失败原因,发现场景复杂度是主因。
When Do Diffusion Models learn to Generate Multiple Objects?

- 构建可控数据集框架Mosaic,分离数据分布与组合泛化影响
- 低数据下计数任务最难学,组合泛化随缺失组合增多而崩溃
- 适合关注多物体生成鲁棒性与数据设计的研究者
文本到图像扩散模型虽具备出色的视觉保真度,但在多物体生成上仍不可靠。尽管有大量实证证据表明其失败,但根本原因尚不明确。本文首先探讨这种局限性在多大程度上源于数据本身。为分离数据影响,研究了两种不同数据规模下的情形:(1) 概念泛化,即每个概念在训练中均有观测,但数据分布可能不平衡;(2) 组合泛化,即特定概念组合被系统性剔除。为此,我们引入Mosaic(多物体空间关系、属性、计数)框架,用于可控的数据生成。在Mosaic上训练扩散模型后发现,场景复杂度起主导作用,而非概念不平衡;且在低数据条件下,计数任务尤为困难。此外,随着训练中被剔除的概念组合增多,组合泛化能力迅速崩溃。这些结果揭示了扩散模型的根本局限,推动更强归纳偏置与数据设计以实现稳健的多物体组合生成。
原文摘要 · Abstract (English)
Text-to-image diffusion models achieve impressive visual fidelity, yet they remain unreliable in multi-object generation. Despite extensive empirical evidence of these failures, the underlying causes remain unclear. We begin by asking how much of this limitation arises from the data itself. To disentangle data effects, we consider two regimes across different dataset sizes: (1) concept generalization, where each individual concept is observed during training under potentially imbalanced data distributions, and (2) compositional generalization, where specific combinations of concepts are systematically held out. To study these regimes, we introduce mosaic (Multi-Object Spatial relations, AttrIbution, Counting), a controlled framework for dataset generation. By training diffusion models on mosaic, we find that scene complexity plays a dominant role rather than concept imbalance, and that counting is uniquely difficult to learn in low-data regimes. Moreover, compositional generalization collapses as more concept combinations are held out during training. These findings highlight fundamental limitations of diffusion models and motivate stronger inductive biases and data design for robust multi-object compositional generation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。