用视觉语言模型自动生成监控视频摘要,省空间还提效率
Large Language Models for Video Surveillance Applications
- 用视觉语言模型按用户提问生成精准视频摘要
- 时间与空间信息准确率分别达80%和70%
- 适合需要快速检索监控事件的安防场景
视频内容激增带来海量数据,分析与管理面临挑战。本文提出一种基于生成式人工智能(GenAI)的创新概念验证,利用视觉语言模型提升下游视频分析效率。该工具可根据用户自定义查询生成定制化文本摘要,从大量监控视频中提取关键信息,相比传统方法的通用摘要或有限动作识别,显著提升分析精度与效率。所提方法能从长时序CCTV画面生成文本摘要,存储体积远小于原始视频,可长期保存,便于快速定位和验证重要事件。定性评估显示,该流程在时间与空间质量一致性上分别达到80%和70%的准确率。
原文摘要 · Abstract (English)
The rapid increase in video content production has resulted in enormous data volumes, creating significant challenges for efficient analysis and resource management. To address this, robust video analysis tools are essential. This paper presents an innovative proof of concept using Generative Artificial Intelligence (GenAI) in the form of Vision Language Models to enhance the downstream video analysis process. Our tool generates customized textual summaries based on user-defined queries, providing focused insights within extensive video datasets. Unlike traditional methods that offer generic summaries or limited action recognition, our approach utilizes Vision Language Models to extract relevant information, improving analysis precision and efficiency. The proposed method produces textual summaries from extensive CCTV footage, which can then be stored for an indefinite time in a very small storage space compared to videos, allowing users to quickly navigate and verify significant events without exhaustive manual review. Qualitative evaluations result in 80% and 70% accuracy in temporal and spatial quality and consistency of the pipeline respectively.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。