arXiv:2505.00254cs.CVcs.AI2025-05中稿 · NDSI 2026, 19pages…被引 5

用视觉语言模型实现超长视频的智能分析,支持复杂问题问答。

AVA: Towards Agentic Video Analytics with Vision Language Models

  • 构建事件知识图谱实时索引长视频,提升处理效率
  • 在长视频任务中达到64.1%准确率,超现有系统表现
  • 适合需要开放场景、复杂推理的视频分析应用

基于AI的视频分析在多个领域日益重要,但现有系统多局限于预设任务,难以应对开放式分析需求。视觉语言模型(VLM)虽有望实现开放式的视频理解与推理,但其有限的上下文窗口制约了对超长视频的处理能力。为此,我们提出AVA系统,利用VLM实现开放式的高级视频分析。核心创新包括:(1) 近实时构建事件知识图谱(EKG),高效索引长或连续视频流;(2) 基于EKG的智能检索-生成机制,应对复杂多样查询。在公开基准LVBench和VideoMME-Long上的评估显示,AVA分别取得62.3%和64.1%的准确率,显著优于现有VLM与视频RAG系统。为评估超长视频与开放世界场景下的分析能力,我们引入新基准AVA-100,包含8段超10小时的视频及120个手动标注的复杂问答对。在该基准上,AVA达到75.8%的准确率,表现顶尖。代码与数据集已开源。

原文摘要 · Abstract (English)

AI-driven video analytics has become increasingly important across diverse domains. However, existing systems are often constrained to specific, predefined tasks, limiting their adaptability in open-ended analytical scenarios. The recent emergence of Vision Language Models (VLMs) as transformative technologies offers significant potential for enabling open-ended video understanding, reasoning, and analytics. Nevertheless, their limited context windows present challenges when processing ultra-long video content, which is prevalent in real-world applications. To address this, we introduce AVA, a VLM-powered system designed for open-ended, advanced video analytics. AVA incorporates two key innovations: (1) the near real-time construction of Event Knowledge Graphs (EKGs) for efficient indexing of long or continuous video streams, and (2) an agentic retrieval-generation mechanism that leverages EKGs to handle complex and diverse queries. Comprehensive evaluations on public benchmarks, LVBench and VideoMME-Long, demonstrate that AVA achieves state-of-the-art performance, attaining 62.3% and 64.1% accuracy, respectively-significantly surpassing existing VLM and video Retrieval-Augmented Generation (RAG) systems. Furthermore, to evaluate video analytics in ultra-long and open-world video scenarios, we introduce a new benchmark, AVA-100. This benchmark comprises 8 videos, each exceeding 10 hours in duration, along with 120 manually annotated, diverse, and complex question-answer pairs. On AVA-100, AVA achieves top-tier performance with an accuracy of 75.8%. The source code of AVA is available at https://github.com/I-ESC/Project-Ava. The AVA-100 benchmark can be accessed at https://huggingface.co/datasets/iesc/Ava-100.

视频分析视觉语言模型知识图谱长视频

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。