用噪声感知的潜空间奖励模型,让扩散模型对齐更快更省算力。
Beyond VLM-Based Rewards: Diffusion-Native Latent Reward Modeling
- 在噪声扩散状态直接做偏好学习,避免像素空间匹配问题。
- 相比传统VLM奖励,算力降低超90%且性能相当甚至更优。
- 适合需要高效对齐扩散模型的研究者和工业部署场景。
扩散模型与流匹配模型的偏好优化依赖于判别性强且计算高效的奖励函数。视觉语言模型(VLM)因其丰富的多模态先验成为主流奖励提供者,但其计算与内存开销大,且通过像素空间奖励优化潜空间扩散生成器存在领域不匹配问题。本文提出DiNa-LRM,一种扩散原生的潜空间奖励建模方法,直接在噪声扩散状态上进行偏好学习。该方法引入依赖扩散噪声的噪声校准Thurstone似然,具备扩散噪声相关的不确定性建模能力。DiNa-LRM基于预训练潜空间扩散主干网络,采用时间步条件化的奖励头,并支持推理时噪声集成,实现测试时可扩展与鲁棒奖励。在图像对齐基准上,DiNa-LRM显著优于现有基于扩散的奖励基线,在计算成本仅为先进VLM的一小部分下达到竞争力表现。在偏好优化中,它改善了优化动态,实现更快、更节省资源的模型对齐。
原文摘要 · Abstract (English)
Preference optimization for diffusion and flow-matching models relies on reward functions that are both discriminatively robust and computationally efficient. Vision-Language Models (VLMs) have emerged as the primary reward provider, leveraging their rich multimodal priors to guide alignment. However, their computation and memory cost can be substantial, and optimizing a latent diffusion generator through a pixel-space reward introduces a domain mismatch that complicates alignment. In this paper, we propose DiNa-LRM, a diffusion-native latent reward model that formulates preference learning directly on noisy diffusion states. Our method introduces a noise-calibrated Thurstone likelihood with diffusion-noise-dependent uncertainty. DiNa-LRM leverages a pretrained latent diffusion backbone with a timestep-conditioned reward head, and supports inference-time noise ensembling, providing a diffusion-native mechanism for test-time scaling and robust rewarding. Across image alignment benchmarks, DiNa-LRM substantially outperforms existing diffusion-based reward baselines and achieves performance competitive with state-of-the-art VLMs at a fraction of the computational cost. In preference optimization, we demonstrate that DiNa-LRM improves preference optimization dynamics, enabling faster and more resource-efficient model alignment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。