arXiv:2512.16899cs.CLcs.CV2025-12被引 11

首个评估图文交错生成的奖励模型基准,揭示当前模型与人类判断的差距。

Multimodal RewardBench 2: Evaluating Omni Reward Models for Interleaved Text and Image

  • 构建多任务图文混合评测集,含1000组专家标注偏好对。
  • 顶尖模型准确率75%-80%,仍低于人类超90%的表现。
  • 适合研究多模态生成、奖励模型评估及对齐的开发者参考。

奖励模型(RMs)在训练大语言模型中至关重要,但在处理图文交错序列的全能模型方面仍研究不足。我们提出多模态奖励基准2.0(MMRB2),首个针对多模态理解与(交错)生成的综合性评测基准。MMRB2涵盖四类任务:文生图、图像编辑、交错生成和多模态推理(“以图思考”),每项任务包含1,000组由23个模型与智能体在21个源任务上生成的专家标注偏好对。评测设计具备:(1)贴近实际且具挑战性的提示;(2)来自最先进模型与智能体的响应;(3)经集成过滤策略筛选的强人类专家共识偏好对。基于MMRB2,我们评估了各子任务中的现有评判方法,包括多模态大模型作为评判者及基于人类偏好训练的模型。最新版Gemini 3 Pro准确率达75-80%,GPT-5与Gemini 2.5 Pro为66-75%,而人类超过90%,显著优于广泛使用的GPT-4o(59%)。最佳开源模型Qwen3-VL-32B表现与Gemini 2.5 Flash相当(64%)。我们还发现MMRB2性能与下游任务成功高度相关,通过Best-of-N采样验证,并深入分析出未来改进的关键方向。

原文摘要 · Abstract (English)

Reward models (RMs) are essential for training large language models (LLMs), but remain underexplored for omni models that handle interleaved image and text sequences. We introduce Multimodal RewardBench 2 (MMRB2), the first comprehensive benchmark for reward models on multimodal understanding and (interleaved) generation. MMRB2 spans four tasks: text-to-image, image editing, interleaved generation, and multimodal reasoning ("thinking-with-images"), providing 1,000 expert-annotated preference pairs per task from 23 models and agents across 21 source tasks. MMRB2 is designed with: (1) practical but challenging prompts; (2) responses from state-of-the-art models and agents; and (3) preference pairs with strong human-expert consensus, curated via an ensemble filtering strategy. Using MMRB2, we study existing judges for each subtask, including multimodal LLM-as-a-judge and models trained with human preferences. The latest Gemini 3 Pro attains 75-80% accuracy. GPT-5 and Gemini 2.5 Pro reach 66-75% accuracy, compared to >90% for humans, yet surpass the widely used GPT-4o (59%). The best performing open-source model Qwen3-VL-32B achieves similar accuracies as Gemini 2.5 Flash (64%). We also show that MMRB2 performance strongly correlates with downstream task success using Best-of-N sampling and conduct an in-depth analysis that shows key areas to improve the reward models going forward.

多模态奖励模型评测基准图文生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。