让扩散模型自己当老师,用极少人工反馈实现高质量对齐。
SAIL: Self-Amplified Iterative Learning for Diffusion Model Alignment with Minimal Human Feedback
- 模型自动生成样本并自我标注偏好,闭环迭代优化。
- 仅需6%的少量数据,效果超越现有主流方法。
- 适合资源有限但需高效对齐的生成模型研究者。
将扩散模型与人类偏好对齐仍具挑战性,尤其在缺乏或难以获取奖励模型、大规模偏好数据集收集成本过高的情况下。本文提出一种新框架SAIL(Self-Amplified Iterative Learning),使扩散模型能通过迭代自提升成为自身教师。从极小规模的人工标注偏好对开始,SAIL在闭环中持续生成多样化样本,基于自身演化理解进行自我标注,并利用自增强数据集不断优化。为防止灾难性遗忘,引入排序偏好混合策略,平衡探索与初始人类先验的保持。大量实验表明,SAIL在多个基准上均优于现有最先进方法,且仅需现有方法6%的偏好数据量,揭示了扩散模型具备显著自改进能力,合理激发后可有效替代大规模人工标注与外部奖励模型。
原文摘要 · Abstract (English)
Aligning diffusion models with human preferences remains challenging, particularly when reward models are unavailable or impractical to obtain, and collecting large-scale preference datasets is prohibitively expensive. \textit{This raises a fundamental question: can we achieve effective alignment using only minimal human feedback, without auxiliary reward models, by unlocking the latent capabilities within diffusion models themselves?} In this paper, we propose \textbf{SAIL} (\textbf{S}elf-\textbf{A}mplified \textbf{I}terative \textbf{L}earning), a novel framework that enables diffusion models to act as their own teachers through iterative self-improvement. Starting from a minimal seed set of human-annotated preference pairs, SAIL operates in a closed-loop manner where the model progressively generates diverse samples, self-annotates preferences based on its evolving understanding, and refines itself using this self-augmented dataset. To ensure robust learning and prevent catastrophic forgetting, we introduce a ranked preference mixup strategy that carefully balances exploration with adherence to initial human priors. Extensive experiments demonstrate that SAIL consistently outperforms state-of-the-art methods across multiple benchmarks while using merely 6\% of the preference data required by existing approaches, revealing that diffusion models possess remarkable self-improvement capabilities that, when properly harnessed, can effectively replace both large-scale human annotation and external reward models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。