arXiv:2501.04764cs.CVcs.MM2025-01被引 1

用生成式AI从监控视频中自动提炼关键事件摘要

Video Summarisation with Incident and Context Information using Generative AI

  • 结合YOLO-V8检测物体,Gemini分析视频与文本
  • 生成摘要与真实内容相似度达72.8%,准确率85%
  • 适合需要快速筛查大量监控视频的安防人员

视频内容的爆炸式增长带来了分析效率和资源利用的巨大挑战。为此,本文提出一种基于生成式人工智能(GenAI)的新方法,实现针对用户查询的定制化视频摘要,提升分析精度与效率。不同于传统框架仅提供通用摘要或有限动作识别,本方法利用生成式AI提取关键信息。通过YOLO-V8进行目标检测,结合Gemini完成视频与文本的综合分析,实现了从大规模监控视频中提取精准文本摘要的能力,使用户无需逐帧查看即可快速定位关键事件。定量评估显示摘要相似度为72.8%,定性评估准确率达85%,验证了该方法的有效性。

原文摘要 · Abstract (English)

The proliferation of video content production has led to vast amounts of data, posing substantial challenges in terms of analysis efficiency and resource utilization. Addressing this issue calls for the development of robust video analysis tools. This paper proposes a novel approach leveraging Generative Artificial Intelligence (GenAI) to facilitate streamlined video analysis. Our tool aims to deliver tailored textual summaries of user-defined queries, offering a focused insight amidst extensive video datasets. Unlike conventional frameworks that offer generic summaries or limited action recognition, our method harnesses the power of GenAI to distil relevant information, enhancing analysis precision and efficiency. Employing YOLO-V8 for object detection and Gemini for comprehensive video and text analysis, our solution achieves heightened contextual accuracy. By combining YOLO with Gemini, our approach furnishes textual summaries extracted from extensive CCTV footage, enabling users to swiftly navigate and verify pertinent events without the need for exhaustive manual review. The quantitative evaluation revealed a similarity of 72.8%, while the qualitative assessment rated an accuracy of 85%, demonstrating the capability of the proposed method.

视频摘要生成式AI监控分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。