arXiv:2511.00962cs.CV2025-11NeurIPS被引 7

用链式推理实现零样本视频异常检测、定位与解释一体化

A Unified Reasoning Framework for Holistic Zero-Shot Video Anomaly Analysis

  • 通过任务间链式推理,串联时序检测、空间定位与语义解释
  • 零样本下在多个基准上达顶尖性能,无需额外训练或梯度
  • 适合需要可解释异常分析的工业场景应用

当前视频异常研究多仅限于帧级检测,缺乏对异常原因的解释,通常只输出帧级异常分数而无空间或语义上下文。近期方法虽提升可解释性,但依赖数据且任务特定。本文提出统一推理框架,连接时序检测、空间定位与文本解释,基于链式测试时推理过程,实现无需额外训练的全零样本异常分析。方法利用任务内推理优化时序检测,通过任务间链式机制实现空间与语义理解,显著提升可解释性与泛化能力。在无额外数据或梯度更新的情况下,该方法在多个视频异常检测、定位与解释基准上达到最优零样本表现。结果表明,精心设计的提示与任务链式结构可释放基础模型的推理潜力,实现实用且可解释的零样本视频异常分析。

原文摘要 · Abstract (English)

Most video-anomaly research stops at frame-wise detection, offering little insight into why an event is abnormal, typically outputting only frame-wise anomaly scores without spatial or semantic context. Recent video anomaly localization and video anomaly understanding methods improve explainability but remain data-dependent and task-specific. We propose a unified reasoning framework that bridges the gap between temporal detection, spatial localization, and textual explanation. Our approach is built upon a chained test-time reasoning process that sequentially connects these tasks, enabling holistic zero-shot anomaly analysis without any additional training. Specifically, our approach leverages intra-task reasoning to refine temporal detections and inter-task chaining for spatial and semantic understanding, yielding improved interpretability and generalization in a fully zero-shot manner. Without any additional data or gradients, our method achieves state-of-the-art zero-shot performance across multiple video anomaly detection, localization, and explanation benchmarks. The results demonstrate that careful prompt design with task-wise chaining can unlock the reasoning power of foundation models, enabling practical, interpretable video anomaly analysis in a fully zero-shot manner. Project Page: https://rathgrith.github.io/Unified_Frame_VAA/.

视频异常零样本可解释性链式推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。