arXiv:2609.06475cs.CV2026-09

让视觉模型主动调整摄像头视角,提升监控视频理解能力

Thinking with Cameras: Active Visual Reasoning via Dynamic Viewpoint Control for Surveillance Video Understanding

论文配图:Thinking with Cameras: Active Visual Reasoning via Dynamic Viewpoint Control for Surveillance Video Understanding
图 1 · 摘自论文原文
  • 通过动态控制摄像头视角,实现主动获取视觉证据
  • 在1.4万+监控视频上验证,显著优于被动观察方法
  • 适合智能安防、自动驾驶等需要主动感知的场景

大型视觉语言模型(LVLM)在通用视频理解中取得显著进展,但在监控视频应用中仍面临领域数据稀缺和固定视角被动观测的限制。关键视觉线索常因目标远、小、遮挡或移出视野而被忽略。本文提出CamVLM框架,使LVLM通过动态视角控制主动获取视觉证据。构建了包含10类异常事件的CCTV-Anomaly数据集,含14,459个视频及详细标注;并建立面向物体的视角轨迹数据集CamTrack-53K,用于学习相机动作。提出基于强化学习的视角策略优化框架,将摄像头控制建模为序列决策问题,学习长时程观测策略,超越监督轨迹模仿。大量实验表明,CamVLM在被动与动态视角设置下均达当前最优性能,验证了主动相机推理的有效性。数据集、模型与代码将公开于https://github.com/xiaozhang79/CamVLM。

原文摘要 · Abstract (English)

Large vision-language models (LVLMs) have recently achieved remarkable progress in general-purpose video understanding. However, their application to surveillance videos remains challenging due to the lack of large-scale domain-specific datasets and the limitation of passive observation from fixed viewpoints. In surveillance scenarios, critical visual evidence can be easily missed when targets are distant, small, occluded, or move beyond the current camera view. In this work, we introduce CamVLM, a new framework for Thinking with Cameras, which enables LVLMs to actively acquire visual evidence through dynamic viewpoint control rather than passively analyzing fixed video streams. We first construct CCTV-Anomaly, a large-scale surveillance video understanding dataset containing 14,459 videos across 10 anomaly categories, with detailed captions and event annotations. We further formulate viewpoint control as an active visual perception problem and build CamTrack-53K, an object-centric viewpoint trajectory dataset for learning camera actions. Moreover, we propose a reinforcement learning based viewpoint policy optimization framework, which models camera control as a sequential decision-making process and learns long-horizon observation strategies beyond supervised trajectory imitation. Extensive experiments demonstrate that CamVLM achieves state-of-the-art performance under both passive observation and dynamic viewpoint settings, validating the effectiveness of active camera-based reasoning for surveillance video understanding. Our datasets, model, and code will be available at https://github.com/xiaozhang79/CamVLM .

视觉推理监控视频主动感知

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。