让视频理解动态匹配问题需求,自动聚焦关键画面。
See What You Need: Query-Aware Visual Intelligence through Reasoning-Perception Loops
- 通过推理-感知闭环,按问题需求动态提取视觉信息。
- 在EgoSchema等数据集上最高提升6.9%,超越现有方法。
- 无需训练,适合需要精准视频分析的场景。
人类视频理解中,推理与视觉注意力动态协同,根据问题自适应聚焦相关细节。然而当前长视频问答系统采用僵化流程,将推理与感知分离,导致信息丢失或计算冗余。核心瓶颈在于视觉提取无法适配具体推理需求:同一视频内容对不同问题需提取不同视觉证据。本文提出CAVIA,一种无需训练的框架,通过推理-感知闭环重构视频理解。其创新包括:(1) 分层推理引导精确定位帧;(2) 跨模态语义桥接实现目标提取;(3) 信心驱动的迭代合成。CAVIA在挑战性基准上取得领先性能:EgoSchema(65.7%,+5.3%)、NExT-QA(76.1%,+2.6%)、IntentQA(73.8%,+6.9%),证明动态推理-感知协调是可扩展的视频理解新范式。
原文摘要 · Abstract (English)
Human video comprehension demonstrates dynamic coordination between reasoning and visual attention, adaptively focusing on query-relevant details. However, current long-form video question answering systems employ rigid pipelines that decouple reasoning from perception, leading to either information loss through premature visual abstraction or computational inefficiency through exhaustive processing. The core limitation lies in the inability to adapt visual extraction to specific reasoning requirements, different queries demand fundamentally different visual evidence from the same video content. In this work, we present CAVIA, a training-free framework that revolutionizes video understanding through reasoning, perception coordination. Unlike conventional approaches where visual processing operates independently of reasoning, CAVIA creates a closed-loop system where reasoning continuously guides visual extraction based on identified information gaps. CAVIA introduces three innovations: (1) hierarchical reasoning, guided localization to precise frames; (2) cross-modal semantic bridging for targeted extraction; (3) confidence-driven iterative synthesis. CAVIA achieves state-of-the-art performance on challenging benchmarks: EgoSchema (65.7%, +5.3%), NExT-QA (76.1%, +2.6%), and IntentQA (73.8%, +6.9%), demonstrating that dynamic reasoning-perception coordination provides a scalable paradigm for video understanding.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。