视频智能体主动寻证,用更少帧实现更准理解
VideoSeek: Long-Horizon Video Agent with Tool-Guided Seeking
- 基于视频逻辑流主动搜寻关键证据,避免全视频扫描
- 在四基准上准确率领先,帧数减少93%且提升10.2分
- 适合长时视频理解与高效推理场景
视频智能体在复杂视频-语言任务中取得进展,但多数方法依赖密集采样帧的贪婪解析,计算开销高。我们提出 VideoSeek,一种长时视频智能体,通过视频逻辑流主动搜寻答案关键证据,而非穷尽解析整段视频。这一思路使模型使用极少帧即可保持甚至提升视频理解能力。VideoSeek 采用思考-行动-观察循环,配备多粒度观测工具包,支持基于查询的累积观测探索,实现高效的视频理解与推理。在四个挑战性视频理解与推理基准上的实验表明,VideoSeek 在显著减少帧数的同时达到优异准确率。尤其在 LVBench 上,相比基线模型 GPT-5,准确率提升 10.2 个百分点,帧数仅使用其 7%。进一步分析揭示了视频逻辑流、强推理能力及工具设计互补作用的重要性。
原文摘要 · Abstract (English)
Video agentic models have advanced challenging video-language tasks. However, most agentic approaches still heavily rely on greedy parsing over densely sampled video frames, resulting in high computational cost. We present VideoSeek, a long-horizon video agent that leverages video logic flow to actively seek answer-critical evidence instead of exhaustively parsing the full video. This insight allows the model to use far fewer frames while maintaining, or even improving, its video understanding capability. VideoSeek operates in a think-act-observe loop with a well-designed toolkit for collecting multi-granular video observations. This design enables query-aware exploration over accumulated observations and supports practical video understanding and reasoning. Experiments on four challenging video understanding and reasoning benchmarks demonstrate that VideoSeek achieves strong accuracy while using far fewer frames than prior video agents and standalone LMMs. Notably, VideoSeek achieves a 10.2 absolute points improvement on LVBench over its base model, GPT-5, while using 93% fewer frames. Further analysis highlights the significance of leveraging video logic flow, strong reasoning capability, and the complementary roles of toolkit design.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。