arXiv:2509.16127cs.CV2025-09被引 15

提出强基准模型BaseReward,提升多模态对齐效果。

BaseReward: A Strong Baseline for Multimodal Reward Model

  • 以Qwen2.5-VL为骨干,优化双层奖励头架构。
  • 在多个基准上达到新SOTA,超越前代模型。
  • 提供可复现的构建指南,适合研究与工业应用。

多模态大语言模型(MLLMs)快速发展,如何对齐人类偏好成为关键挑战。奖励模型(RMs)是实现该目标的核心技术,但学术界与产业界尚缺乏构建顶尖多模态奖励模型(MRMs)的系统性指导。本文通过全面实验分析,系统研究了MRM开发流程中的每一关键组件:奖励建模范式(如朴素RM、基于评判器的RM、生成式RM)、奖励头结构、训练策略、数据筛选(涵盖十余个图文及纯文本偏好数据集)、主干模型与模型规模、集成方法。基于这些实证发现,我们提出 extbf{BaseReward},一种高效强大的多模态奖励建模基线。BaseReward采用简洁有效的架构,基于{Qwen2.5-VL}主干,配备优化的两层奖励头,并在精心筛选的高质量图文与纯文本偏好数据混合集上训练。结果表明,BaseReward在MM-RLHF-Reward Bench、VL-Reward Bench和Multimodal Reward Bench等多个主流基准上均取得新SOTA,优于此前模型。为进一步验证其实际价值,我们将BaseReward集成至真实强化学习管道,成功提升MLLM在感知、推理与对话任务上的表现。本工作不仅贡献了一个顶级的MRM,更向社区提供了一套基于实证的、清晰可靠的下一代MLLM奖励模型构建指南。

原文摘要 · Abstract (English)

The rapid advancement of Multimodal Large Language Models (MLLMs) has made aligning them with human preferences a critical challenge. Reward Models (RMs) are a core technology for achieving this goal, but a systematic guide for building state-of-the-art Multimodal Reward Models (MRMs) is currently lacking in both academia and industry. Through exhaustive experimental analysis, this paper aims to provide a clear ``recipe'' for constructing high-performance MRMs. We systematically investigate every crucial component in the MRM development pipeline, including \textit{reward modeling paradigms} (e.g., Naive-RM, Critic-based RM, and Generative RM), \textit{reward head architecture}, \textit{training strategies}, \textit{data curation} (covering over ten multimodal and text-only preference datasets), \textit{backbone model} and \textit{model scale}, and \textit{ensemble methods}. Based on these experimental insights, we introduce \textbf{BaseReward}, a powerful and efficient baseline for multimodal reward modeling. BaseReward adopts a simple yet effective architecture, built upon a {Qwen2.5-VL} backbone, featuring an optimized two-layer reward head, and is trained on a carefully curated mixture of high-quality multimodal and text-only preference data. Our results show that BaseReward establishes a new SOTA on major benchmarks such as MM-RLHF-Reward Bench, VL-Reward Bench, and Multimodal Reward Bench, outperforming previous models. Furthermore, to validate its practical utility beyond static benchmarks, we integrate BaseReward into a real-world reinforcement learning pipeline, successfully enhancing an MLLM's performance across various perception, reasoning, and conversational tasks. This work not only delivers a top-tier MRM but, more importantly, provides the community with a clear, empirically-backed guide for developing robust reward models for the next generation of MLLMs.

多模态奖励模型基线强化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。