让视频智能体像人一样边看视频边查资料,自主推理。
Thinking Beyond Videos: Unifying Video Reasoning and Deep Research for Open-World Video Agents

- 视频与外部搜索交替进行,动态选择下一步动作。
- 在26K数据上训练,实测性能超越更大开源模型。
- 适合需要跨模态推理的开放世界视频任务研究者。
开放世界视频理解常需定位稀疏视觉证据并获取视频及参数记忆中缺失的外部知识。尽管思维驱动视频感知可实现主动时序感知,深度调研支持多步信息检索,但二者通常孤立发展。我们提出VideoRover,一种统一的视频深度调研框架,通过迭代协调视频裁剪、多模态搜索与网页浏览。给定视频-问题对,每一步工具结果用于选择下一动作:局部视频片段引导外部检索,检索证据又触发进一步视频检查与验证。为构建该能力,我们设计自动化数据清洗流水线,生成26,000条经验证的SFT轨迹和3,000个具有挑战性的强化学习实例。同时引入VideoRover-Bench,按视频时长与研究难度分层的基准。在VideoDR与VideoRover-Bench上的实验表明,VideoRover-8B-RL在无需工具使用的情况下,直接回答表现媲美专有模型,且优于配备相同工具套件的更大开源模型。消融实验与训练动态进一步验证了主动视频定位、外部检索与长程强化学习之间的互补作用。
原文摘要 · Abstract (English)
Open-world video understanding often requires a model to locate sparse visual evidence and acquire external knowledge that is absent from the video and its parametric memory. While Thinking-with-Videos enables active temporal perception and Deep Research supports multi-step information seeking, the two capabilities are typically developed in isolation. We introduce VideoRover, a unified Video Deep Research framework that iteratively coordinates video cropping, multimodal search, and webpage browsing. Given a video-question pair, VideoRover uses each tool result to select the next action, so localized video clips guide external retrieval and retrieved evidence triggers further video inspection and verification. To develop this capability, we construct an automated data curation pipeline, producing 26K verified SFT trajectories and 3K challenging RL instances. We also introduce VideoRover-Bench, a benchmark stratified by video duration and research difficulty. Experiments on VideoDR and VideoRover-Bench show that our VideoRover-8B-RL achieves performance comparable to proprietary models in the direct-answer setting without tool use while outperforming larger open-source models equipped with the same tool suite. Ablation studies and training dynamics further validate the complementary roles of active video grounding, external retrieval, and long-horizon reinforcement learning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。