arXiv:2603.24558cs.CVcs.AI2026-03被引 4

让AI主动选择看视频的时机和方式,提升理解准确性。

LensWalk: Agentic Video Understanding by Planning How You See in Videos

  • AI通过自主规划观看视频的时间段和密度来增强感知。
  • 在长视频任务上准确率提升超5%,无需模型微调。
  • 适合需要精细视觉推理的视频分析场景。

视频内容密集且时间连续,给自动化分析带来巨大挑战。尽管使用了强大的视觉-语言模型,现有视频理解方法仍受限于推理与感知之间的脱节:依赖静态预处理信息,无法随理解进展主动获取原始视频证据。为此,我们提出LensWalk,一种灵活的智能体框架,使大语言模型推理器能够主动控制自身的视觉观察。LensWalk建立了一个紧密的“思考-规划-观察”循环,代理在每一步动态指定其观察视频的时间范围和采样密度。借助一系列由这些参数化的视觉-语言模型工具,代理可进行广域扫描以发现线索、聚焦特定片段提取事实,并整合多时刻证据实现整体验证。该设计支持逐步、按需的证据收集,直接服务于代理不断演化的思维链。无需任何模型微调,LensWalk在多种模型方案上实现显著、即插即用的性能提升,在LVBench和Video-MME等复杂长视频基准上准确率提升超过5%。分析表明,赋予代理对观看方式的控制权,是实现更准确、鲁棒和可解释视频推理的关键。

原文摘要 · Abstract (English)

The dense, temporal nature of video presents a profound challenge for automated analysis. Despite the use of powerful Vision-Language Models, prevailing methods for video understanding are limited by the inherent disconnect between reasoning and perception: they rely on static, pre-processed information and cannot actively seek raw evidence from video as their understanding evolves. To address this, we introduce LensWalk, a flexible agentic framework that empowers a Large Language Model reasoner to control its own visual observation actively. LensWalk establishes a tight reason-plan-observe loop where the agent dynamically specifies, at each step, the temporal scope and sampling density of the video it observes. Using a suite of versatile, Vision-Language Model based tools parameterized by these specifications, the agent can perform broad scans for cues, focus on specific segments for fact extraction, and stitch evidence from multiple moments for holistic verification. This design allows for progressive, on-demand evidence gathering that directly serves the agent's evolving chain of thought. Without requiring any model fine-tuning, LensWalk delivers substantial, plug-and-play performance gains on multiple model recipes, boosting their accuracy by over 5\% on challenging long-video benchmarks like LVBench and Video-MME. Our analysis reveals that enabling an agent to control how it sees is key to unlocking more accurate, robust, and interpretable video reasoning.

视频理解智能体视觉推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。