用重参数技巧让扩散模型高效对齐人类偏好,400步即达顶尖效果。
InPO: Inversion Preference Optimization with Reparametrized DDIM for Efficient Diffusion Model Alignment
- 将扩散模型视为单步生成器,重参数隐变量直接赋隐式奖励。
- 仅微调与偏好数据强相关的潜变量,400步完成训练且生成质量更优。
- 适合追求高效对齐的图像生成研究者,尤其擅长少样本偏好优化。
无需显式奖励信号,直接偏好优化(DPO)通过成对的人类偏好数据微调生成模型,已在大语言模型中引发广泛关注。然而,将文本到图像(T2I)扩散模型与人类偏好对齐的研究仍有限。相较于监督微调,现有方法因长马尔可夫链过程及反向过程不可解,导致训练效率低且生成质量不佳。为此,我们提出DDIM-InPO,一种高效的扩散模型直接偏好对齐方法。该方法将扩散模型视为单步生成模型,实现对特定潜在变量输出的有选择性微调。通过重参数技术为任意潜在变量直接赋予隐式奖励,并引入反演技术估计偏好优化所需的合适潜在变量。此修改过程使扩散模型仅微调与偏好数据强相关部分。实验表明,我们的方法仅需400步微调即达到当前最优性能,在人类偏好评估任务中超越所有现有基线。
原文摘要 · Abstract (English)
Without using explicit reward, direct preference optimization (DPO) employs paired human preference data to fine-tune generative models, a method that has garnered considerable attention in large language models (LLMs). However, exploration of aligning text-to-image (T2I) diffusion models with human preferences remains limited. In comparison to supervised fine-tuning, existing methods that align diffusion model suffer from low training efficiency and subpar generation quality due to the long Markov chain process and the intractability of the reverse process. To address these limitations, we introduce DDIM-InPO, an efficient method for direct preference alignment of diffusion models. Our approach conceptualizes diffusion model as a single-step generative model, allowing us to fine-tune the outputs of specific latent variables selectively. In order to accomplish this objective, we first assign implicit rewards to any latent variable directly via a reparameterization technique. Then we construct an Inversion technique to estimate appropriate latent variables for preference optimization. This modification process enables the diffusion model to only fine-tune the outputs of latent variables that have a strong correlation with the preference dataset. Experimental results indicate that our DDIM-InPO achieves state-of-the-art performance with just 400 steps of fine-tuning, surpassing all preference aligning baselines for T2I diffusion models in human preference evaluation tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。