arXiv:2512.03500cs.CV2025-12被引 1

提出可自主探索与利用的视频理解框架,高效定位关键信息

EEA: Exploration-Exploitation Agent for Long Video Understanding

  • 通过语义引导的分层树搜索实现探索与利用平衡
  • 在多个长视频基准上达到更高准确率且计算开销更低
  • 适合需要高效处理长视频的场景,如视频检索与分析

长视频理解需在海量视觉数据中高效定位稀疏但关键的信息。现有方法或因密集预处理导致严重计算开销,或无法有效平衡探索与利用,造成信息覆盖不全和效率低下。本文提出EEA,一种新型视频代理框架,通过语义引导的分层树搜索过程实现探索-利用平衡。EEA自主发现并动态更新任务相关的语义查询,收集与这些查询高度匹配的视频帧作为语义锚点。在树搜索过程中,不采用均匀扩展,而是优先探索语义相关帧,同时确保未知片段的充分覆盖。此外,EEA通过显式建模不确定性,自适应融合视觉-语言模型(VLMs)的内在奖励与语义先验,实现对视频片段稳定精确的评估。在多个长视频基准上的实验验证了该方法在性能与计算效率上的优越性。

原文摘要 · Abstract (English)

Long-form video understanding requires efficient navigation of extensive visual data to pinpoint sparse yet critical information. Current approaches to longform video understanding either suffer from severe computational overhead due to dense preprocessing, or fail to effectively balance exploration and exploitation, resulting in incomplete information coverage and inefficiency. In this work, we introduce EEA, a novel video agent framework that archives exploration-exploitation balance through semantic guidance with hierarchical tree search process. EEA autonomously discovers and dynamically updates task-relevant semantic queries, and collects video frames closely matched to these queries as semantic anchors. During the tree search process, instead of uniform expansion, EEA preferentially explores semantically relevant frames while ensuring sufficient coverage within unknown segments. Moreover, EEA adaptively combines intrinsic rewards from visionlanguage models (VLMs) with semantic priors by explicitly modeling uncertainty to achieve stable and precise evaluation of video segments. Experiments across various long-video benchmarks validate the superior performance and computational efficiency of our proposed method.

长视频理解探索利用语义搜索

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。