arXiv:2601.06943cs.CVcs.AI2026-01被引 9

构建首个视频深度研究基准,测试模型跨帧找线索、联网检索与多跳推理能力。

Watching, Reasoning, and Searching: A Video Deep Research Benchmark on Open Web for Agentic Video Reasoning

  • 设计视频条件下的开放域问答任务,需跨帧提取视觉锚点
  • 实测表明代理式推理并非始终优于工作流模式,关键在保持锚点一致性
  • 适合研究视频智能体、多模态推理与开放网络搜索的学者

在真实视频问答场景中,视频仅提供局部视觉线索,而可验证答案分布在开放网络中;模型需协同完成跨帧线索提取、迭代检索与基于多跳推理的验证。为此,我们构建首个视频深度研究基准VideoDR,聚焦视频条件下的开放域视频问答,要求跨帧视觉锚点提取、交互式网络检索及联合视频-网页证据的多跳推理;通过严格的真人标注与质量控制,获得覆盖六个语义领域的高质量研究样本。我们在工作流与代理范式下评估多个闭源与开源多模态大模型,结果表明代理式并非始终优于工作流:其优势取决于模型在长检索链中维持初始视频锚点的能力。进一步分析显示,目标漂移与长程一致性是核心瓶颈。VideoDR为研究开放网络环境下视频智能体提供了系统性基准,揭示了下一代视频深度研究智能体的关键挑战。

原文摘要 · Abstract (English)

In real-world video question answering scenarios, videos often provide only localized visual cues, while verifiable answers are distributed across the open web; models therefore need to jointly perform cross-frame clue extraction, iterative retrieval, and multi-hop reasoning-based verification. To bridge this gap, we construct the first video deep research benchmark, VideoDR. VideoDR centers on video-conditioned open-domain video question answering, requiring cross-frame visual anchor extraction, interactive web retrieval, and multi-hop reasoning over joint video-web evidence; through rigorous human annotation and quality control, we obtain high-quality video deep research samples spanning six semantic domains. We evaluate multiple closed-source and open-source multimodal large language models under both the Workflow and Agentic paradigms, and the results show that Agentic is not consistently superior to Workflow: its gains depend on a model's ability to maintain the initial video anchors over long retrieval chains. Further analysis indicates that goal drift and long-horizon consistency are the core bottlenecks. In sum, VideoDR provides a systematic benchmark for studying video agents in open-web settings and reveals the key challenges for next-generation video deep research agents.

视频问答多跳推理智能体开放网络

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。