让视频异常检测的解释与具体物体对齐,提升可信度。
Instance-Aligned Captions for Explainable Video Anomaly Detection

- 为每个异常物体生成带外观和运动属性的精准描述
- 在8个基准数据集上验证,揭示现有方法解释不完整
- 适合需要可验证解释的自动驾驶、安防等场景
可解释的视频异常检测对安全关键应用至关重要,但现有方法缺乏空间定位,导致解释不可验证。尤其在多主体交互中,传统方法常产生不完整或视觉错位的描述。为此,本文提出实例对齐的描述,将每条文本说明与具体物体实例关联,包含其外观、运动特征,明确指出异常责任人、行为、影响对象及位置,实现可验证的推理。我们标注了8个常用VAD基准,并在360度第一人称数据集VIEW360上新增868段视频、8个地点和4类新异常,构建VIEW360+,作为可解释VAD的综合评测基准。实验表明,该方法揭示了当前基于LLM和VLM的方法在解释完整性上的显著缺陷,为未来可信、可解释的异常检测研究提供了可靠评估标准。
原文摘要 · Abstract (English)
Explainable video anomaly detection (VAD) is crucial for safety-critical applications, yet even with recent progress, much of the research still lacks spatial grounding, making the explanations unverifiable. This limitation is especially pronounced in multi-entity interactions, where existing explainable VAD methods often produce incomplete or visually misaligned descriptions, reducing their trustworthiness. To address these challenges, we introduce instance-aligned captions that link each textual claim to specific object instances with appearance and motion attributes. Our framework captures who caused the anomaly, what each entity was doing, whom it affected, and where the explanationis grounded, enabling verifiable and actionable reasoning. We annotate eight widely used VAD benchmarks and extend the 360-degree egocentric dataset, VIEW360, with 868 additional videos, eight locations, and four new anomaly types, creating VIEW360+, a comprehensive testbed for explainable VAD. Experiments show that our instance-level spatially grounded captions reveal significant limitations in current LLM- and VLM-based methods while providing a robust benchmark for future research in trustworthy and interpretable anomaly detection.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。