用低质合成数据训练出更优扩散模型,提升图像质量和多样性。
Ambient Diffusion Omni: Training Good Models with Bad Data
- 利用自然图像的谱衰减与局部性,从低质数据中提取有效信号。
- 在ImageNet上达到顶尖FID,文本到图像生成质量与多样性显著提升。
- 适合数据有限或需高效利用杂乱数据的研究者使用。
我们展示如何利用低质量、合成及分布外图像来提升扩散模型性能。传统扩散模型依赖经过严格筛选的网络数据集,而我们发现大量被丢弃的低质图像具有巨大价值。提出Ambient Diffusion Omni框架,通过利用自然图像的谱功率律衰减和局部性,从所有可用图像中提取信号。实验验证该框架可成功训练受高斯模糊、JPEG压缩、运动模糊等合成退化影响的图像。最终在ImageNet上实现领先水平的FID得分,并显著提升文本到图像生成的质量与多样性。核心洞察是噪声能缓解理想高质量分布与实际观测混合分布之间的初始偏差。通过分析扩散时间轴上偏倚数据与有限无偏数据的学习权衡,提供了严谨理论支持。
原文摘要 · Abstract (English)
We show how to use low-quality, synthetic, and out-of-distribution images to improve the quality of a diffusion model. Typically, diffusion models are trained on curated datasets that emerge from highly filtered data pools from the Web and other sources. We show that there is immense value in the lower-quality images that are often discarded. We present Ambient Diffusion Omni, a simple, principled framework to train diffusion models that can extract signal from all available images during training. Our framework exploits two properties of natural images -- spectral power law decay and locality. We first validate our framework by successfully training diffusion models with images synthetically corrupted by Gaussian blur, JPEG compression, and motion blur. We then use our framework to achieve state-of-the-art ImageNet FID, and we show significant improvements in both image quality and diversity for text-to-image generative modeling. The core insight is that noise dampens the initial skew between the desired high-quality distribution and the mixed distribution we actually observe. We provide rigorous theoretical justification for our approach by analyzing the trade-off between learning from biased data versus limited unbiased data across diffusion times.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。