将生成式推理能力融入高效判别模型,提升视觉奖励模型的准确性与速度。
Joint Reward Modeling: Internalizing Chain-of-Thought for Efficient Visual Reward Models
- 共享视觉语言主干联合训练偏好学习与语言建模。
- 在MMRB2和EditReward-Bench上达到当前最优性能。
- 适合需要快速准确评估的生成式任务下游强化学习场景。
奖励模型在人类反馈的强化学习中至关重要,决定了生成模型的对齐质量和可靠性。对于图像编辑等复杂任务,奖励模型需捕捉全局语义一致性及隐含逻辑约束,而不仅限于局部相似性。现有方法存在明显局限:判别式奖励模型虽与人类偏好对齐良好,但因推理监督有限,难以处理复杂语义;生成式奖励模型具备更强的语义理解与推理能力,但推理成本高且难以直接对齐人类偏好。为此,我们提出联合奖励建模(JRM),在共享的视觉语言主干上联合优化偏好学习与语言建模。该方法将生成式模型的语义与推理能力内化为高效的判别表示,实现快速准确的评估。JRM在MMRB2和EditReward-Bench上取得领先结果,并显著提升下游在线强化学习的稳定性和性能。结果表明,联合训练有效弥合了奖励建模中效率与语义理解的鸿沟。
原文摘要 · Abstract (English)
Reward models are critical for reinforcement learning from human feedback, as they determine the alignment quality and reliability of generative models. For complex tasks such as image editing, reward models are required to capture global semantic consistency and implicit logical constraints beyond local similarity. Existing reward modeling approaches have clear limitations. Discriminative reward models align well with human preferences but struggle with complex semantics due to limited reasoning supervision. Generative reward models offer stronger semantic understanding and reasoning, but they are costly at inference time and difficult to align directly with human preferences. To this end, we propose Joint Reward Modeling (JRM), which jointly optimizes preference learning and language modeling on a shared vision-language backbone. This approach internalizes the semantic and reasoning capabilities of generative models into efficient discriminative representations, enabling fast and accurate evaluation. JRM achieves state-of-the-art results on MMRB2 and EditReward-Bench, and significantly improves stability and performance in downstream online reinforcement learning. These results show that joint training effectively bridges efficiency and semantic understanding in reward modeling.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。