通过细粒度提示提升视觉语言模型在异常视频检测中的准确性和可解释性。
Unlocking Vision-Language Models for Video Anomaly Detection via Fine-Grained Prompting
- 构建以动作为中心的结构化提示框架,细化人类与物体交互的语义描述。
- 在UCF-Crime和XD-Violence数据集上显著提升AUC,超越现有训练自由方法。
- 提供可解释的推理路径,适用于多场景、多模型的异常检测任务。
提示已成为适配冻结的视觉语言模型(VLMs)用于视频异常检测(VAD)的一种实用方法。然而,现有提示往往过于抽象,忽略了监控视频中复杂异常所依赖的细粒度人-物交互或动作语义。本文提出ASK-Hint,一种基于动作中心知识的结构化提示框架,旨在引导冻结的VLM产生更准确且可解释的推理。该方法将提示按语义归类(如暴力、财产犯罪、公共安全),并设计细粒度引导问题,使模型预测与判别性视觉线索对齐。在UCF-Crime和XD-Violence数据集上的大量实验表明,ASK-Hint持续提升AUC表现,优于以往基线,在无训练方法中达到领先水平。除准确性外,该框架还生成可解释的推理轨迹,并展现出跨数据集和VLM主干网络的强大泛化能力。结果凸显了提示粒度的关键作用,确立ASK-Hint为一种新型、无需训练且可解释的视频异常检测方案。
原文摘要 · Abstract (English)
Prompting has emerged as a practical way to adapt frozen vision-language models (VLMs) for video anomaly detection (VAD). Yet, existing prompts are often overly abstract, overlooking the fine-grained human-object interactions or action semantics that define complex anomalies in surveillance videos. We propose ASK-Hint, a structured prompting framework that leverages action-centric knowledge to elicit more accurate and interpretable reasoning from frozen VLMs. Our approach organizes prompts into semantically coherent groups (e.g. violence, property crimes, public safety) and formulates fine-grained guiding questions that align model predictions with discriminative visual cues. Extensive experiments on UCF-Crime and XD-Violence show that ASK-Hint consistently improves AUC over prior baselines, achieving state-of-the-art performance compared to both fine-tuned and training-free methods. Beyond accuracy, our framework provides interpretable reasoning traces towards anomaly and demonstrates strong generalization across datasets and VLM backbones. These results highlight the critical role of prompt granularity and establish ASK-Hint as a new training-free and generalizable solution for explainable video anomaly detection.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。