用博弈论让扩散模型自我对抗,更好对齐人类偏好。
Towards General Preference Alignment: Diffusion Models at Nash Equilibrium

- 让模型与自身对战,通过博弈优化生成质量
- 在多个指标上超越现有对齐方法,效果更稳定
- 适合关注生成模型对齐与强化学习的读者
基于人类反馈的强化学习(RLHF)广泛用于对齐文本到图像(T2I)扩散模型与人类偏好。作为主流方法之一,直接偏好优化(DPO)无需显式建模奖励函数,计算效率高,已被广泛应用于扩散模型对齐。然而,现有基于偏好的扩散模型对齐方法仍依赖于由奖励诱导的偏好信号,并通常假设人类偏好可由布拉德利-特里(BT)模型充分描述,这可能无法捕捉人类偏好的全部复杂性。本文从博弈论视角重新审视扩散模型对齐问题,提出扩散纳什偏好优化(Diff.-NPO),一种直观且通用的扩散模型对齐框架。Diff.-NPO 促使当前策略与自身对抗,实现自我改进并达成更优对齐。实验表明,该方法在文本到图像生成任务中通过多种指标验证了有效性,其表现持续优于现有基于偏好的扩散对齐方法。
原文摘要 · Abstract (English)
Reinforcement learning from human feedback (RLHF) has been popular for aligning text-to-image (T2I) diffusion models with human preferences. As a mainstream branch of RLHF, Direct Preference Optimization (DPO) offers a computationally efficient alternative that avoids explicit reward modeling and has been widely adopted in diffusion alignment. However, existing preference-based methods for diffusion alignment still rely on reward-induced preference signals and typically assume that human preferences can be adequately modeled by the Bradley--Terry (BT) model, which may fail to capture the full complexity of human preferences. In this paper, we formulate diffusion alignment from a game-theoretic perspective. We propose Diffusion Nash Preference Optimization (Diff.-NPO), an intuitive general preference framework for diffusion alignment. Diff.-NPO encourages the current policy to play against itself to achieve self improvement and lead to a better alignment. Empirically, we demonstrate the effectiveness of Diff.-NPO on the text-to-image generation task via various metrics. Diff.-NPO consistently outperforms existing preference-based diffusion alignment methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。