用现成视觉语言模型生成多维可微反馈,让视频扩散模型训练更稳定高效。
Diffusion-DRF: Free, Rich, and Differentiable Reward for Video Diffusion Fine-Tuning
- 用冻结的VLM拆解提示为多维问题,生成丰富反馈
- 无需训练奖励模型,在VBench-2.0上比SOTA高4.74%
- 适合需要高质量视频生成且无标注数据的场景
视频扩散模型对齐长期依赖标量奖励,这些奖励通常来自人类偏好数据集上的学习型奖励模型,需额外训练和大量数据收集。此外,标量奖励提供粗粒度全局监督,难以精准分配提示生成不匹配的评分,易导致奖励滥用和优化不稳定。我们提出Diffusion-DRF,一种免费、丰富且可微的视频扩散微调奖励框架。该框架采用冻结的现成视觉语言模型(VLM)作为评判者,无需训练奖励模型。不同于单一标量奖励,它将每个用户提示分解为多维问题与自由形式的密集视觉问答解释查询,生成信息丰富的反馈。通过直接对这一丰富反馈进行可微优化,Diffusion-DRF实现了无需偏好数据收集的稳定奖励驱动调优。在未见过的VBench-2.0上,其整体性能相比最先进的Flow-GRPO提升4.74%。
原文摘要 · Abstract (English)
Video diffusion alignment has been heavily relied on scalar rewards. These rewards are typically derived from learned reward models in human preference datasets, requiring additional training and extensive collection. Moreover, scalar rewards provide coarse, global supervision, offering limited prompt-generation mismatch credit assignment and making models prone to reward exploitation and unstable optimization. We propose Diffusion-DRF, a free, rich, and differentiable reward framework for video diffusion fine-tuning. Diffusion-DRF employs a frozen, off-the-shelf Vision-Language Model (VLM) as the critic, eliminating the need for reward model training. Instead of relying on a single scalar reward, it decomposes each user prompt into multi-dimensional questions with freeform dense VQA explanation queries, yielding information-rich feedback. By direct differentiable optimization over this rich feedback, Diffusion-DRF achieves stable reward-based tuning without preference datasets collection. Diffusion-DRF achieves significant gains both quantitatively and qualitatively, outperforming state-of-the-art Flow-GRPO by 4.74% in overall performance on unseen VBench-2.0.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。