arXiv:2604.14692cs.CV2026-04被引 1

让视频理解逐步追踪关键物体,提升推理准确性和可解释性。

Chain-of-Glimpse: Search-Guided Progressive Object-Grounded Reasoning for Video Understanding

论文配图:Chain-of-Glimpse: Search-Guided Progressive Object-Grounded Reasoning for Video Understanding
图 1 · 摘自论文原文
  • 用搜索引导的控制器逐步定位视频中的关键物体区域
  • 在多个基准上实现显著性能提升,尤其在复杂场景中更鲁棒
  • 适合需要透明推理过程的视频分析任务,如司法取证或医疗诊断

视频理解需要在多帧中识别并推理语义区分性强的视觉对象,但现有无对象依赖的方法难以应对随时间变化的大量物体差异。为此,我们提出 Chain-of-Glimpse,一种搜索引导的渐进式对象锚定推理框架,将每一步推理显式锚定在具体视觉证据区域,支持组合式与多步决策。该框架将视频推理建模为逐步构建任务相关视觉对象周围空间轨迹的过程,从而减少对显著性线索的过度依赖。具体而言,Chain-of-Glimpse 设计了基于强化学习优化的搜索引导控制器,采用格式奖励机制显著激励对象锚定能力,迭代地定位视觉证据区域并形成可靠推理路径,实现准确且可解释的多步决策。在 NExTQA(域内)和 Video-Holmes、CG-Bench Reasoning、VRBench(域外)等多个基准上的广泛评估表明,Chain-of-Glimpse 在多样化视频推理任务中均展现出一致的性能提升、鲁棒性与泛化能力。

原文摘要 · Abstract (English)

Video understanding requires identifying and reasoning over semantically discriminative visual objects across frames, yet existing object-agnostic solutions struggle to effectively handle substantial object variations over time. To address this, we introduce Chain-of-Glimpse, a search-guided progressive object-grounded reasoning framework that explicitly anchors each reasoning step to specific visual evidence regions, enabling compositional and multi-step decision-making. Formally, Chain-of-Glimpse formulates video reasoning as a step-by-step process that incrementally builds spatially grounded traces around task-relevant visual objects, thereby mitigating over-reliance on saliency-driven cues. Specifically, Chain-of-Glimpse features a search-guided controller, optimized via reinforcement learning with a format reward that significantly incentivizes grounding capability, to iteratively ground visual evidence regions and form reliable reasoning trajectories, yielding accurate and interpretable multi-step decisions. Extensive evaluations on both in domain NExTQA and out-of-domain Video-Holmes, CG-Bench Reasoning, and VRBench benchmarks demonstrate consistent performance gains, robustness and generalization of Chain-of-Glimpse across diverse video reasoning tasks.

视频理解多步推理对象锚定

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。