arXiv:2507.22604cs.CV2025-07ICCV被引 6

用短路径微调提升扩散模型与奖励函数对齐效果

ShortFT: Diffusion Model Alignment via Shortcut-based Fine-Tuning

  • 通过缩短去噪链构建快速微调路径
  • 在多个奖励函数上实现性能超越现有方法
  • 适合需要高效对齐的生成模型优化场景

基于反向传播的方法通过在去噪链中端到端反传奖励梯度来对齐扩散模型与奖励函数,前景广阔。然而,由于去噪链过长带来的计算开销和梯度爆炸风险,现有方法难以完成完整的梯度反传,导致结果欠优。本文提出一种高效的微调策略——快捷路径微调(ShortFT),利用近期提出的保轨迹少步扩散模型,在原去噪链上建立更短的快捷路径,并在此基础上构建简化的去噪链进行优化。该方法显著提升了微调的效率与有效性。实验验证表明,本方法可有效应用于多种奖励函数,大幅改善对齐性能,优于当前最先进的方法。

原文摘要 · Abstract (English)

Backpropagation-based approaches aim to align diffusion models with reward functions through end-to-end backpropagation of the reward gradient within the denoising chain, offering a promising perspective. However, due to the computational costs and the risk of gradient explosion associated with the lengthy denoising chain, existing approaches struggle to achieve complete gradient backpropagation, leading to suboptimal results. In this paper, we introduce Shortcut-based Fine-Tuning (ShortFT), an efficient fine-tuning strategy that utilizes the shorter denoising chain. More specifically, we employ the recently researched trajectory-preserving few-step diffusion model, which enables a shortcut over the original denoising chain, and construct a shortcut-based denoising chain of shorter length. The optimization on this chain notably enhances the efficiency and effectiveness of fine-tuning the foundational model. Our method has been rigorously tested and can be effectively applied to various reward functions, significantly improving alignment performance and surpassing state-of-the-art alternatives.

扩散模型微调对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。