arXiv:2505.03318cs.CV2025-05NeurIPS被引 97

首个支持多模态长链思维的奖励模型,提升视觉任务判断准确性。

Unified Multimodal Chain-of-Thought Reward Model through Reinforcement Fine-Tuning

  • 通过强化微调激发模型的多步推理能力
  • 在多个视觉任务中优于现有方法,提升判断可靠性
  • 适合需要精准视觉偏好评估的研究与应用

多模态奖励模型(RMs)在对齐视觉模型与人类偏好方面展现出巨大潜力。然而,现有模型通常仅能提供直接响应或浅层推理,导致奖励信号不准确。本文提出统一多模态链式思维奖励模型UnifiedReward-Think,首次实现跨视觉理解与生成任务的多维度、分步长链推理。方法包括:1)用少量图像生成偏好数据蒸馏GPT-4o的推理过程,为模型冷启动提供CoT格式;2)基于模型先验知识构建大规模统一多模态偏好数据,通过正确推理输出进行拒绝采样优化;3)利用错误预测样本进行组相对策略优化(GRPO),推动模型探索多样化推理路径并收敛至稳健解。在多个视觉奖励任务上的实验表明,该模型性能显著领先。

原文摘要 · Abstract (English)

Recent advances in multimodal Reward Models (RMs) have shown significant promise in delivering reward signals to align vision models with human preferences. However, current RMs are generally restricted to providing direct responses or engaging in shallow reasoning processes with limited depth, often leading to inaccurate reward signals. We posit that incorporating explicit long chains of thought (CoT) into the reward reasoning process can significantly strengthen their reliability and robustness. Furthermore, we believe that once RMs internalize CoT reasoning, their direct response accuracy can also be improved through implicit reasoning capabilities. To this end, this paper proposes UnifiedReward-Think, the first unified multimodal CoT-based reward model, capable of multi-dimensional, step-by-step long-chain reasoning for both visual understanding and generation reward tasks. Specifically, we adopt an exploration-driven reinforcement fine-tuning approach to elicit and incentivize the model's latent complex reasoning ability: (1) We first use a small amount of image generation preference data to distill the reasoning process of GPT-4o, which is then used for the model's cold start to learn the format and structure of CoT reasoning. (2) Subsequently, by leveraging the model's prior knowledge and generalization capabilities, we prepare large-scale unified multimodal preference data to elicit the model's reasoning process across various vision tasks. During this phase, correct reasoning outputs are retained for rejection sampling to refine the model (3) while incorrect predicted samples are finally used for Group Relative Policy Optimization (GRPO) based reinforcement fine-tuning, enabling the model to explore diverse reasoning paths and optimize for correct and robust solutions. Extensive experiments across various vision reward tasks demonstrate the superiority of our model.

多模态链式思维奖励模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。