arXiv:2601.15417cs.LGcs.AI2026-01

通过迭代优化数据与模型,提升图像和蛋白质生成质量。

Ambient Dataloops: Generative Models for Dataset Refinement

  • 让模型与数据互相进化,逐步提高数据质量
  • 在无条件与文本条件图像生成上达顶尖水平
  • 适合需要高质量生成数据的研究者使用

我们提出 Ambient Dataloops,一种用于数据集迭代精炼的框架,使扩散模型更易学习真实数据分布。现代数据集包含质量参差不齐的样本,直接训练会导致模型表现不佳。本方法实现数据与模型的协同演化:每轮迭代中,数据质量逐步提升,模型性能随之改善。为避免自毁式循环,每次生成时将合成样本视为带噪数据,但噪声水平比上一轮略低,并采用 Ambient Diffusion 技术在扰动下学习。实验表明,Ambient Dataloops 在无条件与文本条件图像生成及从头蛋白质设计任务中均达到当前最优性能。我们还提供了理论支持,阐明数据循环机制的优势。

原文摘要 · Abstract (English)

We propose Ambient Dataloops, an iterative framework for refining datasets that makes it easier for diffusion models to learn the underlying data distribution. Modern datasets contain samples of highly varying quality, and training directly on such heterogeneous data often yields suboptimal models. We propose a dataset-model co-evolution process; at each iteration of our method, the dataset becomes progressively higher quality, and the model improves accordingly. To avoid destructive self-consuming loops, at each generation, we treat the synthetically improved samples as noisy, but at a slightly lower noisy level than the previous iteration, and we use Ambient Diffusion techniques for learning under corruption. Empirically, Ambient Dataloops achieve state-of-the-art performance in unconditional and text-conditional image generation and de novo protein design. We further provide a theoretical justification for the proposed framework that captures the benefits of the data looping procedure.

扩散模型数据精炼生成模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。