现有视频异常检测忽视场景特异性,导致误判率上升。
Is Video Anomaly Detection Misframed? Evidence from LLM-Based and Multi-Scene Models

- 回归单场景建模,强调空间感知与上下文理解
- 多场景通用模型易引入语义偏差,降低定位精度
- 适合需要可解释性的实际监控场景使用
当前视频异常检测研究倾向于构建跨场景通用的正常行为模型,虽提升了可扩展性与多场景泛化能力,却弱化了对特定场景中正常行为的上下文依赖性建模。现有方法常依赖视频级弱监督和多模态大语言模型(MLLM)的黑箱预训练表示,使模型更关注熟悉语义类别而非具体环境中的正常模式偏离。这种趋势抑制了空间定位能力,引入语义偏见,将异常检测简化为动作识别任务。本文通过视觉分析与实证评估,揭示此类范式在真实场景中的局限性,指出有意义的进展需回归单场景、空间敏感且可解释的建模方式,以捕捉各环境内正常的复杂结构。
原文摘要 · Abstract (English)
Recent video anomaly detection research has expanded rapidly with an emphasis on general models of normality intended to work across many different scenes. While this focus has led to improvements in scalability and multi-scene generalization, it has also shifted the field away from modeling the scene-specific and context-dependent nature of normal behavior. Contemporary approaches frequently rely on video-level weak supervision and opaque pretrained representations from multi-modal large language models (MLLMs), which encourage models to respond to familiar semantic anomaly categories rather than to deviations from the normal patterns of a particular environment. This trend suppresses spatial localization, introduces semantic bias, and reduces anomaly detection to a form of action recognition. In this paper, we examine whether these prevailing formulations align with the core requirements of real-world VAD, which is typically performed within a single scene where normality is determined by local geometry, semantics, and activity patterns. Through targeted visual analyses and empirical evaluations, we demonstrate the practical consequences of these limitations and show that meaningful progress in VAD requires renewed focus on single-scene, spatially-aware, and explainable formulations that capture the nuanced structure of normality within individual environments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。