arXiv:2410.21113cs.CVcs.CL2024-10

用大模型零样本识别监控视频动作,性能提升20%以上。

Zero-Shot Action Recognition in Surveillance Videos

  • 用视觉语言大模型+改进采样法处理监控视频
  • 零样本准确率达44.6%,比基线高20%
  • 适合数据少、场景复杂的监控场景应用

公共空间监控需求增长,但人力不足。现有AI系统依赖需大量微调的视觉模型,在监控场景中因数据有限、视角和画质差而难以应用。本文提出利用具备强零样本泛化能力的大型视觉语言模型(LVLM)解决这一问题。我们采用VideoLLaMA2,并引入改进的逐标记采样方法——自省采样(Self-Reflective Sampling, Self-ReS)。在UCF-Crime数据集上的实验表明,VideoLLaMA2相比基线实现20%的零样本性能提升,结合Self-ReS后,零样本动作识别准确率达到44.6%。结果表明,结合优化采样技术的LVLM在多样化监控场景中具有显著潜力。

原文摘要 · Abstract (English)

The growing demand for surveillance in public spaces presents significant challenges due to the shortage of human resources. Current AI-based video surveillance systems heavily rely on core computer vision models that require extensive finetuning, which is particularly difficult in surveillance settings due to limited datasets and difficult setting (viewpoint, low quality, etc.). In this work, we propose leveraging Large Vision-Language Models (LVLMs), known for their strong zero and few-shot generalization, to tackle video understanding tasks in surveillance. Specifically, we explore VideoLLaMA2, a state-of-the-art LVLM, and an improved token-level sampling method, Self-Reflective Sampling (Self-ReS). Our experiments on the UCF-Crime dataset show that VideoLLaMA2 represents a significant leap in zero-shot performance, with 20% boost over the baseline. Self-ReS additionally increases zero-shot action recognition performance to 44.6%. These results highlight the potential of LVLMs, paired with improved sampling techniques, for advancing surveillance video analysis in diverse scenarios.

零样本识别监控视频大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。