arXiv:2605.16079cs.CVcs.AI2026-05

让视频模型主动找片段,精准定位更聪明。

VideoSeeker: Incentivizing Instance-level Video Understanding via Native Agentic Tool Invocation

论文配图:VideoSeeker: Incentivizing Instance-level Video Understanding via Native Agentic Tool Invocation
图 1 · 摘自论文原文
  • 用视觉提示驱动模型主动寻找视频关键片段
  • 比基线平均提升13.7%,超越GPT-4o等闭源模型
  • 适合需要精细定位的视频理解任务开发者

大型视觉语言模型(LVLMs)在视频理解上取得进展,但在需要实例级时空定位的任务中仍面临挑战。现有方法主要依赖文本提示进行人机交互,难以提供精确的空间与时间参考,导致体验不佳。同时,当前方法通常将视觉感知与语言推理分离,使推理聚焦于语言而非视觉内容,限制了模型主动发现细粒度视觉证据的能力。为此,我们提出VideoSeeker,一种通过视觉提示实现实例级视频理解的新范式。该方法将代理推理与实例级视频理解无缝结合,使模型能按需主动感知并检索相关视频段。我们构建了四阶段全自动数据合成管道,高效生成大规模高质量实例级视频数据。通过冷启动监督与强化学习训练,将工具调用与主动感知能力内化至模型,打造强大视频理解能力。实验表明,该模型在实例级视频理解任务上相较基线平均提升13.7%,超越GPT-4o和Gemini-2.5-Pro等强大闭源模型,并在通用视频理解基准上展现良好迁移性。相关数据集与代码将公开发布。

原文摘要 · Abstract (English)

Large Vision-Language Models (LVLMs) have shown significant progress in video understanding, yet they face substantial challenges in tasks requiring precise spatiotemporal localization at the instance level. Existing methods primarily rely on text prompts for human-model interaction, but these prompts struggle to provide precise spatial and temporal references, resulting in poor user experience. Furthermore, current approaches typically decouple visual perception from language reasoning, centering reasoning around language rather than visual content, which limits the model's ability to proactively perceive fine-grained visual evidence. To address these challenges, we propose VideoSeeker, a novel paradigm for instance-level video understanding through visual prompts. VideoSeeker seamlessly integrates agentic reasoning with instance-level video understanding tasks, enabling the model to proactively perceive and retrieve relevant video segments on demand. We construct a four-stage fully automated data synthesis pipeline to efficiently generate large-scale, high-quality instance-level video data. We internalize tool-calling and proactive perception capabilities into the model via cold-start supervision and RL training, building a powerful video understanding model. Experiments demonstrate that our model achieves an average improvement of +13.7% over baselines on instance-level video understanding tasks, surpassing powerful closed-source models such as GPT-4o and Gemini-2.5-Pro, while also showing effective transferability on general video understanding benchmarks. The relevant datasets and code will be released publicly.

视频理解智能体视觉提示定位

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。