arXiv:2510.10518cs.CV2025-10被引 11

让视频奖励模型像人一样边看边思考,提升长视频判断准确率。

VR-Thinker: Boosting Video Reward Models through Thinking-with-Image Reasoning

  • 引入可配置视觉记忆窗口,动态选择和更新关键帧信息。
  • 7B模型在多个基准上达到80.5%以上准确率,长视频表现更优。
  • 适合需要高精度视频生成评估的研究者与开发者使用。

近期多模态奖励模型(RMs)显著提升了视觉生成模型的后训练效果。然而现有模型存在两大瓶颈:(1)视觉输入占用大量上下文预算,导致帧数减少并丢失细粒度细节;(2)所有视觉信息集中于初始提示,加剧了链式思维过程中的幻觉与遗忘问题。为此,我们提出VideoReward Thinker(VR-Thinker),一种支持“以图思辨”的框架,赋予模型视觉推理操作(如选帧)和可配置的视觉记忆窗口。该机制使模型能在上下文限制内主动获取与更新视觉证据,提升推理的准确性与可靠性。通过强化学习微调流程实现:(i)基于精选的视觉思维链数据进行冷启动,蒸馏基础推理技能与操作格式;(ii)筛选出各维度及整体判断均正确的样本,采用拒绝采样微调进一步优化高质量推理轨迹;(iii)应用组相对策略优化(GRPO)增强推理能力。所提方法在开源模型中达到顶尖性能,尤其在长视频任务上表现突出:7B规模的VR-Thinker在VideoGen Reward上达80.5%,GenAI-Bench上为82.3%,MJ-Bench-Video上为75.6%。结果验证了“以图思辨”多模态奖励建模的有效性与前景。

原文摘要 · Abstract (English)

Recent advancements in multimodal reward models (RMs) have substantially improved post-training for visual generative models. However, current RMs face inherent limitations: (1) visual inputs consume large context budgets, forcing fewer frames and causing loss of fine-grained details; and (2) all visual information is packed into the initial prompt, exacerbating hallucination and forgetting during chain-of-thought reasoning. To overcome these issues, we introduce VideoReward Thinker (VR-Thinker), a thinking-with-image framework that equips the RM with visual reasoning operations (e.g., select frame) and a configurable visual memory window. This allows the RM to actively acquire and update visual evidence within context limits, improving reasoning fidelity and reliability. We activate visual reasoning via a reinforcement fine-tuning pipeline: (i) Cold Start with curated visual chain-of-thought data to distill basic reasoning skills and operation formatting; (ii) select samples whose per-dimension and overall judgments are all correct, then conduct Rejection sampling Fine-Tuning on these high-quality traces to further enhance reasoning; and (iii) apply Group Relative Policy Optimization (GRPO) to strengthen reasoning. Our approach delivers state-of-the-art accuracy among open-source models on video preference benchmarks, especially for longer videos: a 7B VR-Thinker achieves 80.5% on VideoGen Reward, 82.3% on GenAI-Bench, and 75.6% on MJ-Bench-Video. These results validate the effectiveness and promise of thinking-with-image multimodal reward modeling.

视频生成奖励模型多模态链式推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。