arXiv:2607.00622cs.CV2026-07中稿 · ICML被引 1

让模型像人一样主动看视频,自动找异常线索。

Learning to Watch: Active Video Anomaly Understanding via Interleaved Policy Optimization

论文配图:Learning to Watch: Active Video Anomaly Understanding via Interleaved Policy Optimization
图 1 · 摘自论文原文
  • 用动态策略主动回溯、扩增时间片段来获取关键证据
  • 仅20亿参数就超越大模型,复杂场景表现更优
  • 适合需要高效理解长视频异常的场景

视频异常理解依赖稀疏且上下文相关的线索。现有被动方法存在观测混淆问题,静态采样难以区分语义不同的事件。为此,我们提出Anom-π,一个闭环框架,将视频理解重构为动态环境中的主动序列决策过程。受人类视频审阅行为启发,该框架将内部认知推理与策略性证据获取融合为交错策略,利用局部回溯、时间扩展和细粒度采样等时序原子操作,赋予模型感知主动性。为在视频级弱监督下学习复杂交互策略,我们设计交互式直接偏好优化(iDPO),以主动证据探究(AEI)效用为引导,平衡任务成功率、信息性证据获取与交互成本。该方法使代理能在排除冗余探索的同时主动消歧假设。大量实验表明,该框架仅用20亿参数,在复杂场景下显著优于当前最先进的大规模VAU模型。

原文摘要 · Abstract (English)

Video anomaly understanding (VAU) relies on sparse, context-dependent cues. However, existing passive paradigms suffer from observational aliasing, where static sampling fails to disambiguate semantically distinct events. To overcome this, we propose $Anom\text{-}π$, a closed-loop framework that reconceptualizes video understanding as an active sequential decision-making process within a dynamic environment. Inspired by human video-reviewing behavior, this framework unifies internal cognitive reasoning and strategic evidence acquisition into an interleaved policy, utilizing temporal atomic operators such as local backtracking, temporal expansion, and fine-grained sampling to endow the model with perceptual proactivity. To learn such complex interaction strategies under video-level weak supervision, we design Interactive Direct Preference Optimization (iDPO) to achieve trajectory-level policy alignment, guided by an Active Evidence Inquiry (AEI) utility that balances task success, informative evidence acquisition, and interaction cost. This approach enables the agent to learn to actively disambiguate hypotheses while suppressing redundant exploration. Extensive experiments demonstrate that our framework, with only 2B parameters, achieves highly competitive performance, significantly outperforming state-of-the-art large-scale VAU models in complex scenarios.

视频异常主动学习强化学习策略优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。