arXiv:2605.01657cs.CV2026-05被引 1

让视觉语言模型主动看视频,边思考边调取画面。

Act2See: Emergent Active Visual Perception for Video Reasoning

论文配图:Act2See: Emergent Active Visual Perception for Video Reasoning
图 1 · 摘自论文原文
  • 在思维链中穿插调用或生成视频帧,实现动态视觉感知。
  • 在VideoEspresso等挑战性数据集上达到新最优,超越更大模型。
  • 适合需要动态理解视频的AI推理场景,如智能监控、自动驾驶。

视觉语言模型(VLMs)通常依赖静态初始帧进行视频推理,难以在推理过程中融入动态信息。现有方法虽尝试将额外帧信息加入思维链(CoT),但常导致思维链质量不佳,且缺乏对假设或反事实场景的视觉信息合成能力。本文提出Act2See框架,通过监督微调(SFT)一个由前沿VLM生成的高质量推理轨迹数据集,使VLM具备主动视觉感知能力:在文本思维链中主动插入对现有帧的检索或新帧的生成调用。这些轨迹经人工标注的思维链严格验证,确保质量。该方法催生出一种涌现能力——推理时模型可自主决定何时搜索或合成所需视觉证据。Act2See在VideoEspresso、ViTIB等挑战性基准上取得新最佳性能,并在Video-MME、EgoNormia和VCR-Bench上超越可比或更大模型,显著推进了视觉语言模型在视频推理中的主动视觉感知能力。

原文摘要 · Abstract (English)

Vision-Language Models (VLMs) typically rely on static initial frames for video reasoning, restricting their ability to incorporate essential dynamic information as the reasoning process evolves. Existing methods that augment Chain-of-Thought (CoT) with additional frame information often exhibit suboptimal CoT quality and lack the crucial ability to synthesize visual information for hypothetical or counterfactual scenarios. We introduce Act-to-See (Act2See), a novel framework that enables active visual perception by empowering VLMs to actively interleave video frames within text CoTs. Act2See is developed via Supervised Fine-Tuning (SFT) on a high-quality dataset of reasoning traces generated by a frontier VLM. These traces integrate active calls to either retrieve existing frames or generate new ones, and are rigorously verified against human-annotated CoTs to ensure quality. This approach cultivates an emergent capability: at inference time, the model actively determines when to search for or synthesize the necessary visual evidence. Act2See establishes new state-of-the-art results on challenging benchmarks, including VideoEspresso and ViTIB, and outperforms comparable or larger models on Video-MME, EgoNormia, and VCR-Bench, demonstrating an advancement in enabling VLMs with active visual perception for video reasoning.

视频推理主动感知视觉语言模型思维链

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。