arXiv:2410.05116cs.LGcs.AI2024-10ICLR被引 7

用在线人类反馈高效微调图像生成模型,提升对齐性与效率。

HERO: Human-Feedback Efficient Reinforcement Learning for Online Diffusion Model Finetuning

  • 通过在线反馈学习人类偏好,动态生成训练信号。
  • 仅需0.5K次反馈即可完成异常修正、计数等任务,效率提升4倍。
  • 适合需要快速迭代、人力成本高的个性化图像生成场景。

通过稳定扩散(Stable Diffusion, SD)微调实现可控生成,旨在提升生成质量、安全性及与人类意图的对齐性。现有基于人类反馈的强化学习方法通常依赖预定义的启发式奖励函数或大规模数据训练的奖励模型,限制了在难以收集数据场景下的应用。为此,我们提出HERO框架,利用模型学习过程中实时采集的在线人类反馈。HERO包含两个核心机制:(1) 反馈对齐表征学习,一种在线训练方法,可捕捉人类反馈并提供有效的微调信号;(2) 反馈引导图像生成,基于优化后的初始化样本生成图像,加速收敛至评估者意图。实验表明,HERO在体部异常修正任务上比最优现有方法效率提升4倍。此外,仅需0.5K次在线反馈,HERO即可有效完成推理、计数、个性化及降低NSFW内容等任务。

原文摘要 · Abstract (English)

Controllable generation through Stable Diffusion (SD) fine-tuning aims to improve fidelity, safety, and alignment with human guidance. Existing reinforcement learning from human feedback methods usually rely on predefined heuristic reward functions or pretrained reward models built on large-scale datasets, limiting their applicability to scenarios where collecting such data is costly or difficult. To effectively and efficiently utilize human feedback, we develop a framework, HERO, which leverages online human feedback collected on the fly during model learning. Specifically, HERO features two key mechanisms: (1) Feedback-Aligned Representation Learning, an online training method that captures human feedback and provides informative learning signals for fine-tuning, and (2) Feedback-Guided Image Generation, which involves generating images from SD's refined initialization samples, enabling faster convergence towards the evaluator's intent. We demonstrate that HERO is 4x more efficient in online feedback for body part anomaly correction compared to the best existing method. Additionally, experiments show that HERO can effectively handle tasks like reasoning, counting, personalization, and reducing NSFW content with only 0.5K online feedback. The code and project page are available at https://hero-dm.github.io/.

扩散模型强化学习人类反馈微调

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。