arXiv:2503.19429cs.LGcs.CV2025-03被引 2

提出量化扩散模型复现训练数据难易程度的新方法

Quantifying the Ease of Reproducing Training Data in Unconditional Diffusion Models

  • 利用反向扩散过程中的朗之万方程建立图像与噪声的对应关系
  • 通过潜在空间体积增长率衡量数据复现概率,可识别易被记忆的样本
  • 计算高效,适合用于提升训练数据质量,避免版权风险

近年来快速发展的扩散模型可能生成与训练数据高度相似的样本,引发版权问题。本文提出一种量化无条件扩散模型中训练数据复现难易程度的方法。在反向扩散过程中,遵循朗之万方程的样本群体平均行为满足一阶常微分方程(ODE),该方程在潜在空间中建立了图像与其噪声版本的一一对应关系。由于该ODE可逆且初始噪声随机采样,图像投影区域的体积代表其生成概率。通过分析该映射的体积增长速率,我们成功量化了训练数据被复现的难易程度。该方法计算复杂度低,可用于检测并修改易被记忆的训练样本,从而提升训练数据质量。

原文摘要 · Abstract (English)

Diffusion models, which have been advancing rapidly in recent years, may generate samples that closely resemble the training data. This phenomenon, known as memorization, may lead to copyright issues. In this study, we propose a method to quantify the ease of reproducing training data in unconditional diffusion models. The average of a sample population following the Langevin equation in the reverse diffusion process moves according to a first-order ordinary differential equation (ODE). This ODE establishes a 1-to-1 correspondence between images and their noisy counterparts in the latent space. Since the ODE is reversible and the initial noisy images are sampled randomly, the volume of an image's projected area represents the probability of generating those images. We examined the ODE, which projects images to latent space, and succeeded in quantifying the ease of reproducing training data by measuring the volume growth rate in this process. Given the relatively low computational complexity of this method, it allows us to enhance the quality of training data by detecting and modifying the easily memorized training samples.

扩散模型数据复现版权风险

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。