用高阶动力学降低扩散模型的记忆化,保护隐私与版权。
Reducing Diffusion Model Memorization with Higher Order Langevin Dynamics

- 引入高阶拉普拉斯动态,通过速度加速度变量约束生成路径。
- 高阶模型使得分函数平滑,抑制对训练数据的直接复现。
- 实验证明高阶模型在真实数据上更少记忆原数据,适合隐私敏感场景。
扩散/基于得分的模型虽能生成高质量样本,但易重现训练数据(即‘记忆化’),引发版权与隐私问题。本文首次理论分析高阶拉普拉斯动态(HOLD)的正则化效果:其通过引入辅助变量(如速度、加速度)形成更高阶动力学,使数据变量演化受低通滤波后的得分函数控制,平滑性随阶数提升。我们分析了最优经验得分与分布坍缩可能性,揭示随着模型阶数增加,记忆化现象被有效缓解。实验在真实数据上验证理论,表明HOLD在实践中显著优于标准扩散模型。
原文摘要 · Abstract (English)
Diffusion/score-based models have emerged as powerful generative models, capable of generating high-quality samples that mimic the training data distribution. However, it has been observed that they are prone to reproducing training samples-known as "memorization"-potentially violating copyright and privacy. In this paper, we study the effect of Higher-Order Langevin Dynamics (HOLD) on this phenomenon. HOLD diffusion processes introduce auxiliary variables; if the data variable is interpreted as "position," then the auxiliary variables can be interpreted as "velocity" and "acceleration," depending on the chosen order of the model. They were originally proposed based on the intuition that they regularize the trajectories of the data variable by implicitly imposing additional dynamical constraints. Our work provides, to our knowledge, the first theoretical characterization of the regularization effect of HOLD. Specifically, we show that in HOLD, the dynamics of the data variable are governed by a low-pass-filtered version of the learned score function, with smoothness increasing with the order of HOLD. We then analyze the optimal empirical score and the possibility of distribution collapse. Together, our results explain the mitigation of memorization as the model order increases. Finally, we present an empirical study on real-world data that supports our theory and highlights this distinct advantage of HOLD over standard diffusion in practice.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。