根据问题动态选择证据获取方式,提升长视频理解效果
Routing Before Looking: Query-Adaptive Evidence Acquisition for Long-form Video Understanding

- 通过路由策略动态选择全局浏览、时间定位或语义检索工具
- 在多个数据集上达到当前最优性能,且帧效率高
- 适合需要灵活应对复杂查询的长视频理解任务
长视频理解对视频智能体而言仍具挑战性,主要源于查询需求与证据获取策略之间的不匹配。尽管近期规划先于感知的方法优于无查询感知的流水线,但通常依赖单一主导策略(生成式或检索式),难以应对多样化查询需求。我们提出Route2Look,一种轻量级、模型无关的查询自适应证据获取框架。该框架采用‘路由-查看-记忆’循环,包含三个工具:全局浏览(Global Browse)获取整体上下文,时间定位(Temporal Ground)提取显式时间线索,语义检索(Semantic Retrieve)进行语义搜索。核心是基于查询动态选择工具的路由策略。该策略通过两阶段设计构建:首先从生成式与检索式轨迹的差异对比分析中蒸馏路由能力,再在推理时应用硬性路由规则与继续/停止判断。在多个挑战性长视频基准测试中,Route2Look实现了最先进性能,且在不同数据集和查询类型下均保持高效帧利用率。真值路由分析进一步揭示了查询自适应证据获取对未来长视频理解的潜力。
原文摘要 · Abstract (English)
Long-form video understanding remains challenging for video agents due to the mismatch between query demands and evidence acquisition strategies. Although recent planning-before-perception methods outperform query-agnostic pipelines, they often rely on a single dominant strategy, either generation-based strategy or retrieval-based strategy, limiting their ability to handle diverse query demands. We propose Route2Look, a lightweight and model-agnostic framework for query-adaptive evidence acquisition in long-form video understanding. Route2Look operates in a Route-Look-Memorize loop with three tools: Global Browse for holistic context, Temporal Ground for explicit temporal cues, and Semantic Retrieve for semantic search. The core component is a routing policy that dynamically selects evidence acquisition tools based on the query. To build this policy, Route2Look adopts a two-stage design: first distilling the routing skill from differential contrastive analysis between generation-based and retrieval-based trajectories, and then applying the distilled skill with hard routing rules and continue-or-stop criteria during inference. Experiments on challenging long-video benchmarks show that Route2Look achieves state-of-the-art performance while maintaining strong frame efficiency across datasets and query types. Oracle routing analysis further reveals the potential of query-adaptive evidence acquisition for future long-form video understanding.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。