arXiv:2503.05236cs.CV2025-03被引 187

首个统一评估多模态理解与生成的奖励模型,提升视觉任务性能。

Unified Reward Model for Multimodal Understanding and Generation

  • 构建跨图像与视频生成/理解的大规模人类偏好数据集训练统一奖励模型。
  • 通过两阶段筛选自动生成高质量配对偏好数据,用于模型优化。
  • 在理解与生成任务中均实现一致提升,适合多模态系统研发者使用。

近期人类偏好对齐的进步显著提升了多模态生成与理解能力。核心方法是训练提供监督信号的奖励模型。然而,现有奖励模型多为任务特定,限制了在多样视觉应用中的适应性。本文提出,联合学习评估多种视觉任务的奖励模型可能产生协同效应:更好的图像理解可提升图像生成评估,更优的帧分析亦能改善视频评估。为此,本文提出UnifiedReward,首个面向多模态理解与生成评估的统一奖励模型,支持成对排序与点评分,为视觉模型偏好对齐提供有效奖励信号。具体而言:(1)在自建的大规模人类偏好数据集上训练UnifiedReward,覆盖图像与视频的生成/理解任务;(2)利用该模型通过两阶段策略(成对排序与点筛选)自动构建高质量配对偏好数据;(3)基于这些数据,采用直接偏好优化(DPO)对视觉模型进行人类偏好对齐。实验表明,联合评估多种视觉任务带来显著互惠效益。进一步将该流程应用于视觉理解与生成,各领域均实现持续改进。

原文摘要 · Abstract (English)

Recent advances in human preference alignment have significantly improved multimodal generation and understanding. A key approach is to train reward models that provide supervision signals for preference optimization. However, existing reward models are often task-specific, limiting their adaptability across diverse visual applications. We also argue that a reward model that jointly learning to assess multiple vision tasks may foster a synergistic effect, where improved image understanding enhances image generation assessment, and refined image evaluation benefits video assessment through better frame analysis. To this end, this paper proposes UnifiedReward, the first unified reward model for multimodal understanding and generation assessment. It supports both pairwise ranking and pointwise scoring, providing effective reward signals for vision model preference alignment. Specifically, (1) we first train UnifiedReward on our constructed large-scale human preference dataset, which covers both image and video generation/understanding tasks. (2) Then, we leverage it to automatically construct high-quality pairwise preference data from vision models by progressively filtering their outputs through our two-stage strategy, i.e., pair ranking and point sifting. (3) Finally, we use these data to align vision models with human preferences via Direct Preference Optimization (DPO). Experimental results show that jointly learning to assess diverse visual tasks yields substantial mutual benefits. We further apply our pipeline to both vision understanding and generation, achieving consistent improvements across each domain.

多模态奖励模型偏好对齐生成评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。