arXiv:2603.13306cs.CVcs.LG2026-03被引 6

小模型也能高效准确检测监控异常,适合弱监督场景。

Benchmarking Compact VLMs for Clip-Level Surveillance Anomaly Detection Under Weak Supervision

  • 用轻量视觉语言模型+参数高效微调,适配弱监督监控任务。
  • 准确率媲美甚至超过主流方法,单片段推理延迟可控。
  • 统一评测标准让结果可比,适合工程落地部署。

CCTV安全监控要求异常检测器在弱监督下兼具可靠的片段级准确率和可预测的单片段延迟。本文研究紧凑型视觉-语言模型(VLMs)在此场景下的实用性。建立统一评估协议,标准化预处理、提示设计、数据集划分、评价指标与运行环境,用于对比参数高效微调的紧凑VLMs、无需训练的VLM流水线及弱监督基线。评估涵盖准确率、精确率、召回率、F1、ROC-AUC与平均单片段延迟,综合衡量检测质量与效率。经参数高效适应后,紧凑VLMs性能达到甚至超过现有方法,同时保持良好的单片段延迟表现。适应还降低提示敏感性,在统一协议下展现出更一致的行为。结果表明,参数高效微调使紧凑VLMs成为可靠、高效的片段级异常检测器,在透明一致的实验设置中实现优异的精度-效率权衡。

原文摘要 · Abstract (English)

CCTV safety monitoring demands anomaly detectors combine reliable clip-level accuracy with predictable per-clip latency despite weak supervision. This work investigates compact vision-language models (VLMs) as practical detectors for this regime. A unified evaluation protocol standardizes preprocessing, prompting, dataset splits, metrics, and runtime settings to compare parameter-efficiently adapted compact VLMs against training-free VLM pipelines and weakly supervised baselines. Evaluation spans accuracy, precision, recall, F1, ROC-AUC, and average per-clip latency to jointly quantify detection quality and efficiency. With parameter-efficient adaptation, compact VLMs achieve performance on par with, and in several cases exceeding, established approaches while retaining competitive per-clip latency. Adaptation further reduces prompt sensitivity, producing more consistent behavior across prompt regimes under the shared protocol. These results show that parameter-efficient fine-tuning enables compact VLMs to serve as dependable clip-level anomaly detectors, yielding a favorable accuracy-efficiency trade-off within a transparent and consistent experimental setup.

视频异常检测小模型弱监督VLM

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。