arXiv:2608.31005cs.CV2026-08

根据任务需求动态调整视频证据获取策略,提升长视频理解准确性

From Intent to Evidence: Policy-Steered Multi-Strategy Retrieval for Long-Video Agents

论文配图:From Intent to Evidence: Policy-Steered Multi-Strategy Retrieval for Long-Video Agents
图 1 · 摘自论文原文
  • 通过意图路由选择聚焦、召回或对比检索策略,智能匹配不同问题的证据需求
  • 在LongVideoBench长片段上提升6.9分,全指标优于现有方法
  • 无需训练,可自动识别缺失证据并引导后续采集,适合复杂视频推理任务

现有长视频代理采用单一行为获取证据,忽视证据分布特性(集中、覆盖广或需区分假设),易导致推理前失败。为避免预设固定流程限制自主探索,本文提出VESTA——一种无需训练的长视频代理,采用路径条件化的‘获取-验证-整合’循环。在探索前,意图路由器基于共享视听场景索引,推断证据获取策略(聚焦、召回或对比)和证据管理策略。策略驱动的检索生成初步参考,多模态证据操作将其转为观测,推理器可自由验证、重查或检查未包含区域。时间证据账本将观测整合为动态压缩视图,涵盖时间位置、来源、覆盖率、冲突、验证结果与假设支持,暴露缺失与未决证据以指导后续获取;最终化优先处理已验证观测。在Video-MME-v2上,平均准确率较VideoARM提升2.7点;在LongVideoBench、EgoSchema、LVBench上,长子集提升6.9点,LVBench提升1.5点,与VideoARM在EgoSchema持平。

原文摘要 · Abstract (English)

Existing long-video agents acquire evidence through one uniform behavior, ignoring whether the required evidence is concentrated, requires broad occurrence coverage, or must discriminate competing hypotheses---which can cause failure before substantive reasoning begins. Prescribing a fine-grained solution procedure for every question is not a satisfactory remedy, as it restricts autonomous exploration. We propose VESTA, a training-free long-video agent organized as a route-conditioned acquire--verify--consolidate loop. Before exploration, an intent router infers an evidence-acquisition policy---focused, recall, or contrastive retrieval over a shared visual--speech scene index---together with an evidence-accounting policy that configures the evidence view maintained during exploration. Policy-steered retrieval yields provisional references that multimodal evidence operations convert into observations, while the Reasoner remains free to verify them, re-query using intermediate findings, or inspect regions outside the retrieved set. A temporal evidence ledger consolidates observations into an adaptive, compressed view of temporal location, provenance, coverage, conflicts, verification outcomes, and hypothesis support, exposing missing and unresolved evidence to guide subsequent acquisition; finalization prioritizes verified observations. On Video-MME-v2, VESTA improves average accuracy by 2.7 points over VideoARM and gains across all six reported metrics. On LongVideoBench, EgoSchema, and LVBench under shared query-time models, it improves by 6.9 points on the LongVideoBench long subset and 1.5 on LVBench, and matches VideoARM on EgoSchema.

视频理解推理机制多策略检索

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。