arXiv:2607.02927cs.CVcs.AI2026-07被引 2

让AI像研究员一样看视频,多工具协同推理找证据。

VideoSearcher: Empowering Video Deep Research with Multi-Tool Agentic Reasoning via Reinforcement Learning

论文配图:VideoSearcher: Empowering Video Deep Research with Multi-Tool Agentic Reasoning via Reinforcement Learning
图 1 · 摘自论文原文
  • 用多工具闭环推理,融合时间定位与视觉聚焦
  • 在VideoSearch-QA上超越现有开源模型,提升显著
  • 适合需要深度视频分析的科研与智能检索场景

视频理解正从封闭场景感知转向开放世界证据探索,这一范式被称为视频深度研究(VDR)。然而现有多模态搜索代理主要面向静态图像,当前VDR基准依赖文本中心检索,忽略了关键视觉信息。为此,我们提出VideoSearcher,一个闭环代理框架,赋予视觉语言模型多工具推理能力以支持VDR。VideoSearcher将时间定位、空间聚焦和多模态搜索统一于单一推理轨迹中,使代理能逐步定位视觉线索、检索相关证据并合成答案。为优化知识密集型推理轨迹,我们提出双分支序列策略优化(BiSPO),将工具调用优化与答案准确率优化解耦,提供稳定学习信号。此外,我们构建了首个评估开放世界视频信息定位与多模态搜索推理的基准VideoSearch-QA。大量实验表明,VideoSearcher在多个搜索导向与多模态理解基准上显著优于现有开源代理基线。

原文摘要 · Abstract (English)

Video understanding is moving beyond closed-context perception toward open-world evidence exploration, a paradigm formalized as Video Deep Research (VDR). However, existing multimodal search agents primarily target static images, and the current VDR benchmark relies on text-centric retrieval that discards crucial visual information. To address these limitations, we propose VideoSearcher, a closed-loop agentic framework that empowers Vision-Language Models with multi-tool reasoning for VDR. VideoSearcher unifies temporal localization, spatial focusing, and multimodal search within a single reasoning trajectory, enabling agents to progressively ground visual clues, retrieve relevant evidence, and synthesize answers. To optimize knowledge-intensive reasoning trajectories, we propose Bi-branch Sequence Policy Optimization (BiSPO), a reinforcement learning algorithm that decouples tool-invocation optimization from answer-accuracy optimization. This design provides stable learning signals for both evidence-grounded reasoning and purposeful tool use. Furthermore, we construct VideoSearch-QA, the first benchmark designed to evaluate open-world video information grounding and multimodal search-based reasoning. Extensive experiments demonstrate that VideoSearcher significantly outperforms prior open-source agentic baselines across various search-oriented and multimodal understanding benchmarks.

视频理解多工具推理强化学习开放世界

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。