arXiv:2411.02780cs.LGcs.CV2024-11ICLR被引 20

用少量干净数据+大量噪声数据,可达到纯干净数据训练效果。

How much is a noisy image worth? Data Scaling Laws for Ambient Diffusion

  • 混合使用少量干净数据与大量噪声数据训练扩散模型
  • 仅用噪声数据无法达到干净数据训练的性能上限
  • 适合数据采集受限的科研场景,如医学成像或天文观测

生成模型的性能依赖于训练数据的质量。构建大规模高质量数据集通常成本高昂甚至不可行,尤其在受物理或仪器限制的科学领域中难以获取干净数据。本文研究了仅使用污染数据训练扩散模型的效果,对三个数据集(3万至约130万样本)在不同污染水平下训练超过80个模型。结果表明,在现有样本规模下,仅用噪声数据无法达到纯干净数据训练模型的性能。但若结合少量干净数据(如总量的10%)和大量高度噪声数据,即可实现与纯干净数据训练相当的性能,甚至接近当前最优水平。我们通过针对异方差高斯混合模型的新样本复杂度界,提供了理论支持:在大数据量下,噪声样本的有效边际收益呈指数级低于干净样本。实验验证了少量干净数据能显著降低对噪声数据的需求。

原文摘要 · Abstract (English)

The quality of generative models depends on the quality of the data they are trained on. Creating large-scale, high-quality datasets is often expensive and sometimes impossible, e.g. in certain scientific applications where there is no access to clean data due to physical or instrumentation constraints. Ambient Diffusion and related frameworks train diffusion models with solely corrupted data (which are usually cheaper to acquire) but ambient models significantly underperform models trained on clean data. We study this phenomenon at scale by training more than $80$ models on data with different corruption levels across three datasets ranging from $30,000$ to $\approx 1.3$M samples. We show that it is impossible, at these sample sizes, to match the performance of models trained on clean data when only training on noisy data. Yet, a combination of a small set of clean data (e.g.~$10\%$ of the total dataset) and a large set of highly noisy data suffices to reach the performance of models trained solely on similar-size datasets of clean data, and in particular to achieve near state-of-the-art performance. We provide theoretical evidence for our findings by developing novel sample complexity bounds for learning from Gaussian Mixtures with heterogeneous variances. Our theoretical model suggests that, for large enough datasets, the effective marginal utility of a noisy sample is exponentially worse than that of a clean sample. Providing a small set of clean samples can significantly reduce the sample size requirements for noisy data, as we also observe in our experiments.

扩散模型数据质量噪声数据样本效率

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。