用语言理解视频异常,让检测结果可解释
Knowledge-Guided Textual Reasoning for Explainable Video Anomaly Detection via LLMs
- 通过视觉语言模型将视频转为细粒度描述文本
- 在动作、物体、上下文、环境四类语义槽中推理异常原因
- 适合需要可解释性的安防场景应用
我们提出基于文本的可解释视频异常检测(TbVAD),一种完全在文本域内进行弱监督视频异常检测的语言驱动框架。不同于依赖显式视觉特征的传统方法,TbVAD通过语言表示视频语义,实现可解释且知识引导的推理。该框架分三步运行:(1) 使用视觉语言模型将视频内容转化为细粒度描述;(2) 将描述组织成动作、物体、上下文、环境四个语义槽,构建结构化知识;(3) 生成各语义槽的解释,揭示导致异常判断的关键因素。在UCF-Crime和XD-Violence两个公开数据集上评估表明,基于文本的知识推理能为真实监控场景提供可解释且可靠的异常检测。
原文摘要 · Abstract (English)
We introduce Text-based Explainable Video Anomaly Detection (TbVAD), a language-driven framework for weakly supervised video anomaly detection that performs anomaly detection and explanation entirely within the textual domain. Unlike conventional WSVAD models that rely on explicit visual features, TbVAD represents video semantics through language, enabling interpretable and knowledge-grounded reasoning. The framework operates in three stages: (1) transforming video content into fine-grained captions using a vision-language model, (2) constructing structured knowledge by organizing the captions into four semantic slots (action, object, context, environment), and (3) generating slot-wise explanations that reveal which semantic factors contribute most to the anomaly decision. We evaluate TbVAD on two public benchmarks, UCF-Crime and XD-Violence, demonstrating that textual knowledge reasoning provides interpretable and reliable anomaly detection for real-world surveillance scenarios.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。