arXiv:2509.00484cs.CVcs.AI2025-09被引 4

首个全面评估视频多模态奖励模型的基准,覆盖感知、知识、推理与安全四大维度。

VideoRewardBench: Comprehensive Evaluation of Multimodal Reward Models for Video Understanding

  • 构建包含1563个样本的高质量偏好数据集,问题量达前人基准的15倍。
  • 顶尖模型GPT-4o在跨模态理解上准确率仅57.0%,开源模型最高53.3%。
  • 揭示奖励模型训练方式与推理扩展性的关键影响,适合视频AI研发者参考。

多模态奖励模型(MRMs)在大型视觉语言模型(LVLMs)的训练、推理与评估中起关键作用,用于判断响应质量。然而现有视频领域MRM评估基准存在题目数量与多样性不足、评估维度不全、对不同类型MRM覆盖有限等问题。为此,我们提出VideoRewardBench,首个涵盖视频理解四大核心维度——感知、知识、推理与安全的综合性基准。通过人工智能辅助的数据流水线,我们构建了一个高质量偏好数据集,含1,563个标注样本,涵盖1,482个唯一视频和1,559个不同问题,是现有最丰富基准的15倍。每个样本为视频-文本提示、优选回答与次优回答的三元组。我们在28种跨类型(生成式、判别式、半标量)多模态奖励模型上开展全面评估。结果显示,即使是最先进的模型GPT-4o整体准确率也仅达57.0%,而当前最先进的开源模型Qwen2.5-VL-72B仅为53.3%。分析揭示三大关键发现:(i) 经强化学习训练的MRM未必比非强化学习训练的具备更强跨模态泛化能力;(ii) 除判别式外,其他类型MRM在不同规模下均能从推理时扩展中获益;(iii) 输入视频帧数变化对不同类型MRM的影响各异。我们认为VideoRewardBench为推动视频领域多模态奖励模型的评估与发展提供了极具挑战性且宝贵的基准。

原文摘要 · Abstract (English)

Multimodal reward models (MRMs) play a crucial role in the training, inference, and evaluation of Large Vision Language Models (LVLMs) by assessing response quality. However, existing benchmarks for evaluating MRMs in the video domain suffer from a limited number and diversity of questions, a lack of comprehensive evaluation dimensions, and inadequate evaluation of diverse types of MRMs. To address these gaps, we introduce VideoRewardBench, the first comprehensive benchmark covering four core aspects of video understanding: perception, knowledge, reasoning, and safety. Through our AI-assisted data pipeline, we curate a high-quality preference dataset of 1,563 annotated samples, including 1,482 unique videos and 1,559 distinct questions--15 times the number found in the most question-rich prior benchmark. Each sample is a triplet consisting of a video-text prompt, a chosen response, and a rejected response. We also conduct a comprehensive evaluation across 28 multimodal reward models spanning three categories: generative, discriminative, and semi-scalar. Results show that even the top-performing model GPT-4o achieves only 57.0% overall accuracy, and the state-of-the-art open-source model Qwen2.5-VL-72B reaches merely 53.3%. Our analysis further reveals three key insights: (i) MRMs trained with reinforcement learning (RL) do not necessarily exhibit stronger cross-modal generalization than those trained without RL; (ii) except for discriminative MRMs, other types of MRMs across varying model capacities can benefit from inference-time scaling; and (iii) variations in input video frame count have different effects on different types of MRMs. We believe VideoRewardBench offers a challenging and valuable benchmark for advancing the evaluation and development of MRMs in the video domain.

多模态视频理解奖励模型评估基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。