构建多模态奖励模型评估基准,揭示当前模型在推理与安全上的短板
Multimodal RewardBench: Holistic Evaluation of Reward Models for Vision Language Models
- 设计涵盖六领域的专家标注数据集,含5211个三元组样本
- 顶尖模型如Gemini 1.5 Pro仅达72%整体准确率,推理与安全表现差
- 适合研究多模态对齐、奖励建模及安全评测的开发者使用
奖励模型在训练视觉语言模型(VLMs)中至关重要,通过评估输出质量实现与人类偏好的对齐。然而,社区缺乏针对多模态奖励模型的全面开放评估基准。为此,我们提出Multimodal RewardBench,一个覆盖六大领域(通用正确性、偏好、知识、推理、安全、视觉问答)的专家标注基准。数据集包含5,211个由各类VLM生成的(提示,优选回应,次优回应)三元组。对多种VLM判官模型的评估显示,即使表现最佳的Gemini 1.5 Pro和Claude 3.5 Sonnet也仅达到72%的整体准确率,且多数模型在推理与安全领域表现不佳。这些结果表明,Multimodal RewardBench为跨领域奖励模型发展提供了具有挑战性的测试平台。该基准已开源至https://github.com/facebookresearch/multimodal_rewardbench。
原文摘要 · Abstract (English)
Reward models play an essential role in training vision-language models (VLMs) by assessing output quality to enable aligning with human preferences. Despite their importance, the research community lacks comprehensive open benchmarks for evaluating multimodal reward models in VLMs. To address this gap, we introduce Multimodal RewardBench, an expert-annotated benchmark covering six domains: general correctness, preference, knowledge, reasoning, safety, and visual question-answering. Our dataset comprises 5,211 annotated (prompt, chosen response, rejected response) triplets collected from various VLMs. In evaluating a range of VLM judges, we find that even the top-performing models, Gemini 1.5 Pro and Claude 3.5 Sonnet, achieve only 72% overall accuracy. Notably, most models struggle in the reasoning and safety domains. These findings suggest that Multimodal RewardBench offers a challenging testbed for advancing reward model development across multiple domains. We release the benchmark at https://github.com/facebookresearch/multimodal_rewardbench.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。