深度扩散模型会放大训练数据中的偏见,引发隐私风险。
Deeper Diffusion Models Amplify Bias
- 通过理论与实证分析,揭示深度扩散模型的偏差-方差权衡机制。
- 更深的模型在生成时会加剧训练数据中的固有偏见。
- 适用于关注生成模型公平性与隐私安全的研究者。
尽管生成式扩散模型(DMs)表现出色,但其内部机制仍不清晰,存在潜在问题。本文聚焦于扩散模型中的偏差-方差权衡问题,建立系统性基础:在极端情况下,扩散模型可能放大训练数据中的固有偏见,另一方面也可能损害训练样本的隐含隐私。研究结果与生成模型的记忆-泛化理解一致,但进一步拓展了该谱系,揭示了深层模型中偏见放大的风险。相关结论通过理论推导与实证验证得到支持。
原文摘要 · Abstract (English)
Despite the remarkable performance of generative Diffusion Models (DMs), their internal working is still not well understood, which is potentially problematic. This paper focuses on exploring the important notion of bias-variance tradeoff in diffusion models. Providing a systematic foundation for this exploration, it establishes that at one extreme, the diffusion models may amplify the inherent bias in the training data, and on the other, they may compromise the presumed privacy of the training samples. Our exploration aligns with the memorization-generalization understanding of the generative models, but it also expands further along this spectrum beyond "generalization", revealing the risk of bias amplification in deeper models. Our claims are validated both theoretically and empirically.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。