揭示扩散模型生成与记忆的平衡机制,提出可控编辑新方法
Generalization of Diffusion Models Arises with a Balanced Representation Space
- 通过自编码器分析发现:记忆对应局部尖峰表示,泛化则依赖均衡统计表示
- 在真实扩散模型中验证了相同表示结构,表明该机制具广泛适用性
- 提出无训练编辑技术,可通过控制表示实现精准生成调控
扩散模型虽能生成高质量多样样本,但过度拟合训练目标时易出现数据记忆。本文从表征学习视角分析记忆与泛化的差异:通过两层ReLU去噪自编码器证明,记忆表现为将原始训练样本存储于权重中,产生局部尖峰表示;而泛化则源于对局部数据统计的捕捉,形成均衡表示。进一步在真实无条件及文本到图像扩散模型上验证了该理论,发现相同表示结构普遍存在,具有重要实践意义。基于此,提出基于表征的记忆检测方法和无需训练的编辑技术,实现通过表征操控进行精确控制。结果表明,学习优质表征是生成建模的核心。
原文摘要 · Abstract (English)
Diffusion models excel at generating high-quality, diverse samples, yet they risk memorizing training data when overfit to the training objective. We analyze the distinctions between memorization and generalization in diffusion models through the lens of representation learning. By investigating a two-layer ReLU denoising autoencoder (DAE), we prove that (i) memorization corresponds to the model storing raw training samples in the learned weights for encoding and decoding, yielding localized spiky representations, whereas (ii) generalization arises when the model captures local data statistics, producing balanced representations. Furthermore, we validate these theoretical findings on real-world unconditional and text-to-image diffusion models, demonstrating that the same representation structures emerge in deep generative models with significant practical implications. Building on these insights, we propose a representation-based method for detecting memorization and a training-free editing technique that allows precise control via representation steering. Together, our results highlight that learning good representations is central to novel and meaningful generative modeling.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。