arXiv:2605.01662cs.CV2026-05被引 2

用主动感知提升视频问答的帧选择效率,让大模型更聪明地看视频。

Video Active Perception: Effective Inference-Time Long-Form Video Understanding with Vision-Language Models

论文配图:Video Active Perception: Effective Inference-Time Long-Form Video Understanding with Vision-Language Models
图 1 · 摘自论文原文
  • 通过主动感知理论,动态选关键帧而非均匀采样。
  • 在多个数据集上零样本表现领先,帧效率提升最高达5.6倍。
  • 无需训练,适合长视频推理任务,尤其擅长理解复杂问题。

大型视觉语言模型(VLMs)在视频问答等多模态任务中取得进展,但面临帧选择效率低下的挑战,标准均匀采样成本高且性能易饱和。受主动感知理论启发——模型通过获取与预期不同的数据来增益信息——我们提出视频主动感知(VAP),一种无需训练的方法,用于增强基于VLM的长视频问答。该方法将关键帧选择视为主动感知中的数据获取,并利用轻量级文本条件视频生成模型表征先验世界知识。实验表明,VAP在长视频或推理型视频问答数据集(如EgoSchema、NExT-QA、ActivityNet-QA、IntentQA、CLEVRER)上实现最优零样本性能,相比GPT-4o、Gemini 1.5 Pro和LLaVA-OV,帧效率最高提升5.6倍。此外,VAP展现出更强的推理能力,能有效选出与问题相关的关键帧。这些结果凸显了主动感知在提升长视频问答帧效率与有效性方面的潜力。

原文摘要 · Abstract (English)

Large vision-language models (VLMs) have advanced multimodal tasks such as video question answering (QA). However, VLMs face the challenge of selecting frames effectively and efficiently, as standard uniform sampling is expensive and performance may plateau. Inspired by active perception theory, which posits that models gain information by acquiring data that differs from their expectations, we introduce Video Active Perception (VAP), a training-free method to enhance long-form video QA using VLMs. Our approach treats keyframe selection as data acquisition in active perception and leverages a lightweight text-conditioned video generation model to represent prior world knowledge. Empirically, VAP achieves state-of-the-art zero-shot results on long-form or reasoning video QA datasets such as EgoSchema, NExT-QA, ActivityNet-QA, IntentQA, and CLEVRER, achieving an increase of up to 5.6 x frame efficiency by frames per question over standard GPT-4o, Gemini 1.5 Pro, and LLaVA-OV. Moreover, VAP shows stronger reasoning abilities than previous methods and effectively selects keyframes relevant to questions. These findings highlight the potential of leveraging active perception to improve the frame effectiveness and efficiency of long-form video QA.

视频问答主动感知高效推理多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。