arXiv:2603.04727cs.CVcs.AI2026-03被引 1

测试多模态大模型在真实场景下的异常检测能力,发现其过于保守导致漏检严重。

Are Multimodal LLMs Ready for Surveillance? A Reality Check on Zero-Shot Anomaly Detection in the Wild

  • 将异常检测转为语言推理任务,用提示词控制时间窗口进行零样本评估
  • 模型在上海科技数据集上峰值F1仅0.09,召回率严重不足
  • 特定类别提示词可提升性能至F1 0.64,但仍需改进召回率

多模态大语言模型(MLLMs)在视频理解方面展现出强大泛化能力,但其在真实世界视频异常检测(VAD)中的可靠性尚未充分探索。与依赖重建或姿态线索的传统方法不同,MLLMs实现了范式转变:将异常检测视为语言引导的推理任务。本文系统评估了先进MLLMs在ShanghaiTech和CHAD基准上的表现,将VAD重构为弱时序监督下的二分类任务。研究了提示词具体性与时长窗口(1-3秒)对性能的影响,重点关注精确率与召回率权衡。结果表明,零样本设置下模型存在明显保守偏差:尽管置信度高,却过度倾向'正常'类,导致精确率高而召回率崩溃,限制实际应用。我们证明,使用类别特定指令可显著调整决策边界,使ShanghaiTech上的峰值F1从0.09提升至0.64,但召回仍是关键瓶颈。这些结果揭示了MLLMs在噪声环境中的显著性能差距,并为未来面向开放世界监控的召回导向提示设计与模型校准研究提供了基础。

原文摘要 · Abstract (English)

Multimodal large language models (MLLMs) have demonstrated impressive general competence in video understanding, yet their reliability for real-world Video Anomaly Detection (VAD) remains largely unexplored. Unlike conventional pipelines relying on reconstruction or pose-based cues, MLLMs enable a paradigm shift: treating anomaly detection as a language-guided reasoning task. In this work, we systematically evaluate state-of-the-art MLLMs on the ShanghaiTech and CHAD benchmarks by reformulating VAD as a binary classification task under weak temporal supervision. We investigate how prompt specificity and temporal window lengths (1s--3s) influence performance, focusing on the precision--recall trade-off. Our findings reveal a pronounced conservative bias in zero-shot settings; while models exhibit high confidence, they disproportionately favor the 'normal' class, resulting in high precision but a recall collapse that limits practical utility. We demonstrate that class-specific instructions can significantly shift this decision boundary, improving the peak F1-score on ShanghaiTech from 0.09 to 0.64, yet recall remains a critical bottleneck. These results highlight a significant performance gap for MLLMs in noisy environments and provide a foundation for future work in recall-oriented prompting and model calibration for open-world surveillance, which demands complex video understanding and reasoning.

多模态大模型异常检测零样本视频理解

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。