arXiv:2502.12089stat.MLcs.LG2025-02ICML被引 29

扩散模型通过层次聚类学习数据组合规则,训练数据越多,生成内容越连贯。

How Compositional Generalization and Creativity Improve as Diffusion Models are Trained

  • 基于概率上下文无关语法建模数据层级结构,揭示生成机制
  • 小样本时仅局部连贯,中等规模数据下全局连贯性显著提升
  • 机制类似物理中的重整化群,适合研究生成模型可扩展性

自然数据通常以特征的层次组合方式组织。生成模型需要多少样本才能学习这些组合规则,从而生成数量级庞大的新数据?数据中哪些信号被用来学习这些规则?我们从理论和实证两方面研究了扩散模型在这一问题上的表现。理论上,我们采用简单的概率上下文无关语法——一种用于表示语言和图像等数据层次与组合结构的树状图模型。结果表明,扩散模型能以聚类具有相似上下文特征所需的样本复杂度学习语法的组合规则,这一过程类似于word2vec算法。但聚类是分层涌现的:与更长上下文相关联的高层特征需要更多数据才能识别,导致样本复杂度随上下文长度呈多项式增长。因此,在中等数据规模下训练的扩散模型生成的数据仅在一定尺度内保持连贯,缺乏全局一致性。我们在多个领域测试了这些预测,发现惊人的一致性:随着训练时间或数据集规模增大,生成文本和图像的连贯长度逐步增加。我们还探讨了本文提出的层次聚类机制与物理学中重整化群之间的联系。

原文摘要 · Abstract (English)

Natural data is often organized as a hierarchical composition of features. How many samples do generative models need in order to learn the composition rules, so as to produce a combinatorially large number of novel data? What signal in the data is exploited to learn those rules? We investigate these questions in the context of diffusion models both theoretically and empirically. Theoretically, we consider a simple probabilistic context-free grammar - a tree-like graphical model used to represent the hierarchical and compositional structure of data such as language and images. We demonstrate that diffusion models learn the grammar's composition rules with the sample complexity required for clustering features with statistically similar context, a process similar to the word2vec algorithm. However, this clustering emerges hierarchically: higher-level features associated with longer contexts require more data to be identified. This mechanism leads to a sample complexity that scales polynomially with the said context size. As a result, diffusion models trained on an intermediate dataset size generate data coherent up to a certain scale, but lacking global coherence. We test these predictions across different domains and find remarkable agreement: both generated texts and images achieve progressively larger coherence lengths as the training time or dataset size grows. We discuss connections between the hierarchical clustering mechanism we introduce here and the renormalization group in physics.

扩散模型组合泛化层次结构生成模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。