arXiv:2509.24943cs.CV2025-09被引 1

让AI像人一样看长视频,自适应感知+主动验证,减少幻觉。

Perceive, Verify and Understand Long Video: Multi-Granular Perception and Active Verification via Interactive Agents

  • 用可变粒度感知和主动验证双代理互动,动态调整分析策略。
  • 仅用11.2帧就达到Gemini 1.5-Pro水平,在EgoSchema上精度领先。
  • 适合需要高效准确理解长视频的场景,如视频问答、智能监控。

长视频具有时间复杂性和任务相关信息稀疏的特点,对AI系统构成显著推理挑战。尽管现有基于大语言模型(LLM)的方法已推进长视频理解,但仍受限于任务无关、固定粒度的感知流程,并易产生视觉-语言幻觉。受人类自适应感知与主动验证启发,我们提出CogniGPT框架,通过多粒度感知代理(MPA)与主动验证代理(AVA)之间的交互循环实现优化。具体而言,MPA根据上下文动态决定最优感知粒度与策略,而AVA主动挖掘多视角视觉证据,交叉验证关键观察结果,消除幻觉。该机制使CogniGPT能高效识别最少且可靠的、与任务相关的线索。在EgoSchema、Video-MME、NExT-QA和MovieChat上的大量实验表明其在准确率与效率上均具优势。特别地,在EgoSchema上,仅使用11.2帧即超越现有无训练方法,性能媲美Gemini 1.5-Pro。

原文摘要 · Abstract (English)

Long videos, characterized by temporal complexity and sparse task-relevant information, pose significant reasoning challenges for AI systems. Although existing Large Language Model (LLM)-based approaches have advanced long video understanding, they remain bottlenecked by task-agnostic, fixed-granularity perception pipelines and suffer from vision-language hallucinations. Inspired by human adaptive perception and active verification, we propose CogniGPT, a framework leveraging an interactive loop between a Multi-Granular Perception Agent (MPA) and an Active Verification Agent (AVA). Specifically, instead of predetermined heuristics, MPA adaptively determines the optimal perception granularity and strategy based on the evolving context, while AVA actively mines multi-perspective visual evidence to cross-verify key observations and eliminate hallucinations. This interaction allows CogniGPT to efficiently identify a minimal set of reliable task-related clues. Extensive experiments on EgoSchema, Video-MME, NExT-QA, and MovieChat demonstrate its superiority in accuracy and efficiency. Notably, on EgoSchema, it surpasses existing training-free methods using only 11.2 frames and achieves performance comparable to Gemini 1.5-Pro.

长视频理解多粒度感知主动验证LLM应用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。