arXiv:2511.14446cs.CVcs.AI2025-11被引 6

提出无需训练的视频理解框架,模拟人类分阶段思考过程。

Agentic Video Intelligence: A Flexible Framework for Advanced Video Exploration and Understanding

  • 采用检索-感知-复盘三阶段机制,兼顾全局探索与局部分析。
  • 在多个长视频数据集上表现媲美顶尖模型,且结果可解释性强。
  • 开源模型组合+实体图知识库,适合研究者快速搭建视频智能系统。

视频理解不仅需要视觉识别,还需复杂推理能力。尽管视觉语言模型(VLM)表现出色,但通常以单次遍历方式处理视频,缺乏证据回溯与迭代优化支持。虽有基于智能体的方法实现长时程推理,但依赖昂贵专有模型或需大量强化学习训练。为此,我们提出无训练、灵活的智能体视频理解框架(Agentic Video Intelligence, AVI),通过系统级设计模仿人类视频认知过程。AVI引入三项创新:(1) 受人类启发的三阶段推理流程(检索-感知-复盘),兼顾全局探索与局部聚焦;(2) 基于实体图结构的视频知识库及多粒度集成工具,构成智能体交互环境;(3) 开源模型集合,融合推理型大语言模型与轻量级基础视觉模型及VLM,避免对专有API或强化学习训练的依赖。在LVBench、VideoMME-Long、LongVideoBench和Charades-STA上的实验表明,AVI实现具有竞争力的性能,同时具备更强可解释性。

原文摘要 · Abstract (English)

Video understanding requires not only visual recognition but also complex reasoning. While Vision-Language Models (VLMs) demonstrate impressive capabilities, they typically process videos largely in a single-pass manner with limited support for evidence revisit and iterative refinement. While recently emerging agent-based methods enable long-horizon reasoning, they either depend heavily on expensive proprietary models or require extensive agentic RL training. To overcome these limitations, we propose Agentic Video Intelligence (AVI), a flexible and training-free framework that can mirror human video comprehension through system-level design and optimization. AVI introduces three key innovations: (1) a human-inspired three-phase reasoning process (Retrieve-Perceive-Review) that ensures both sufficient global exploration and focused local analysis, (2) a structured video knowledge base organized through entity graphs, along with multi-granularity integrated tools, constituting the agent's interaction environment, and (3) an open-source model ensemble combining reasoning LLMs with lightweight base CV models and VLM, eliminating dependence on proprietary APIs or RL training. Experiments on LVBench, VideoMME-Long, LongVideoBench, and Charades-STA demonstrate that AVI achieves competitive performance while offering superior interpretability.

视频理解智能体多模态可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。