arXiv:2409.06493cs.CVcs.AI2024-09被引 11

提出新方法在保持图像多样性的同时优化人类偏好,解决扩散模型奖励过拟合问题。

Elucidating Optimal Reward-Diversity Tradeoffs in Text-to-Image Diffusion Models

  • 引入推理时正则化AIG,基于退火重要性采样思想
  • 在Stable Diffusion上实现奖励与多样性的帕累托最优平衡
  • 适用于多种模型和奖励函数,用户研究验证效果

文本到图像扩散模型已成为从文本提示生成高保真图像的重要工具。然而,在未过滤的互联网数据上训练时,这些模型可能生成不安全、错误或风格不符的图像,与人类偏好不一致。为解决此问题,近期方法通过人类偏好数据微调模型或优化捕捉偏好的奖励函数。尽管有效,但易出现奖励黑客现象,即模型过度拟合奖励函数,导致生成图像多样性下降。本文证明了奖励黑客的不可避免性,并研究了KL散度、LoRA缩放等自然正则化技术在扩散模型中的局限性。同时提出推理时正则化方法Annealed Importance Guidance(AIG),受退火重要性采样启发,可在保留基础模型多样性的同时实现帕累托最优的奖励-多样性权衡。实验表明,AIG在Stable Diffusion模型中能有效平衡奖励优化与图像多样性。用户研究进一步证实,AIG在不同模型架构和奖励函数下均提升了生成图像的质量与多样性。

原文摘要 · Abstract (English)

Text-to-image (T2I) diffusion models have become prominent tools for generating high-fidelity images from text prompts. However, when trained on unfiltered internet data, these models can produce unsafe, incorrect, or stylistically undesirable images that are not aligned with human preferences. To address this, recent approaches have incorporated human preference datasets to fine-tune T2I models or to optimize reward functions that capture these preferences. Although effective, these methods are vulnerable to reward hacking, where the model overfits to the reward function, leading to a loss of diversity in the generated images. In this paper, we prove the inevitability of reward hacking and study natural regularization techniques like KL divergence and LoRA scaling, and their limitations for diffusion models. We also introduce Annealed Importance Guidance (AIG), an inference-time regularization inspired by Annealed Importance Sampling, which retains the diversity of the base model while achieving Pareto-Optimal reward-diversity tradeoffs. Our experiments demonstrate the benefits of AIG for Stable Diffusion models, striking the optimal balance between reward optimization and image diversity. Furthermore, a user study confirms that AIG improves diversity and quality of generated images across different model architectures and reward functions.

扩散模型图像生成奖励对齐多样性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。