构建首个视频理解奖励模型基准,提升模型判断力与推理能力。
Video Understanding Reward Modeling: A Robust Benchmark and Performant Reward Models

- 设计包含2100组偏好数据的VURB基准,支持长链条推理评估。
- 自动化构建3.5万条高质量视频偏好数据集,支撑模型训练。
- 提出两种新模型,在推理与测试时缩放中表现领先,适合视频智能研究者。
多模态奖励模型在文本和图像领域取得显著进展,但在视频理解奖励建模方面仍受限于缺乏稳健的评估基准和高质量偏好数据。为此,我们提出一个涵盖基准设计、数据构建和模型训练的统一框架。引入Video Understanding Reward Bench(VURB),包含2,100组偏好对,每组带有平均1,143词符的长链思维推理过程,并在通用、长视频及推理型任务上采用多数投票评估。进一步通过全自动化流程构建Video Understanding Preference Dataset(VUP-35K),提供大规模高质量监督信号用于视频奖励模型训练。基于该数据,我们训练了判别式(VideoDRM)与生成式(VideoGRM)奖励模型,两者在VURB与VideoRewardBench上均达到当前最优性能。分析表明,VUP-35K不仅提升奖励模型表现,还增强模型推理能力;VideoDRM与VideoGRM在best-of-$N$测试时缩放下均有显著增益。
原文摘要 · Abstract (English)
Multimodal reward models have advanced substantially in text and image domains, yet progress in video understanding reward modeling remains severely limited by the lack of robust evaluation benchmarks and high-quality preference data. To address this, we propose a unified framework spanning benchmark design, data construction, and reward model training. We introduce Video Understanding Reward Bench (VURB), a benchmark featuring 2,100 preference pairs with long chain-of-thought reasoning traces (averaging 1,143 tokens) and majority voting evaluation across general, long, and reasoning-oriented video tasks. We further construct Video Understanding Preference Dataset (VUP-35K) via a fully automated pipeline, providing large-scale high-quality supervision for video reward training. Building on the data, we train VideoDRM and VideoGRM, a discriminative and a generative reward model, both achieving state-of-the-art performance on VURB and VideoRewardBench. Further analysis confirms that VUP-35K enhances both reward performance and model reasoning capability, while VideoDRM and VideoGRM yield significant gains under best-of-$N$ test-time scaling.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。