arXiv:2412.07775cs.LGcs.CV2024-12ICLR被引 30

用梯度信息提升扩散模型微调速度与生成多样性。

Efficient Diversity-Preserving Diffusion Alignment via Gradient-Informed GFlowNets

  • 基于奖励梯度设计新型生成流程网络,实现高效微调。
  • 在多个真实奖励函数下,快速保持生成多样性和先验特征。
  • 适合需要高质量可控生成的图像生成研究者使用。

尽管大型扩散模型通常通过目标下游任务数据集进行训练,但常需在预训练模型上对齐并微调,以适应由专家设计或小规模数据集学习得到的奖励函数。现有扩散模型的奖励微调方法普遍存在生成样本多样性不足、先验保留差及微调收敛慢的问题。针对此挑战,本文受生成流网络(GFlowNets)近期成功启发,提出一种强化学习微调方法——Nabla-GFlowNet($ abla$-GFlowNet),利用奖励梯度中的丰富信号实现概率性扩散模型微调。实验表明,该方法在不同真实奖励函数下,可实现对Stable Diffusion(一种大规模文本条件图像扩散模型)的快速、保多样性和保先验的微调。

原文摘要 · Abstract (English)

While one commonly trains large diffusion models by collecting datasets on target downstream tasks, it is often desired to align and finetune pretrained diffusion models with some reward functions that are either designed by experts or learned from small-scale datasets. Existing post-training methods for reward finetuning of diffusion models typically suffer from lack of diversity in generated samples, lack of prior preservation, and/or slow convergence in finetuning. In response to this challenge, we take inspiration from recent successes in generative flow networks (GFlowNets) and propose a reinforcement learning method for diffusion model finetuning, dubbed Nabla-GFlowNet (abbreviated as $\nabla$-GFlowNet), that leverages the rich signal in reward gradients for probabilistic diffusion finetuning. We show that our proposed method achieves fast yet diversity- and prior-preserving finetuning of Stable Diffusion, a large-scale text-conditioned image diffusion model, on different realistic reward functions.

扩散模型强化学习图像生成微调

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。