arXiv:2503.24008cs.CVcs.AI2025-03被引 1

新基准H2VU评估模型对长视频和实时视频的综合理解能力

H2VU-Benchmark: A Comprehensive Benchmark for Hierarchical Holistic Video Understanding

  • 构建分层全景视频理解框架,覆盖3秒到1.5小时视频
  • 新增反常识推理与轨迹追踪任务,检验深层理解能力
  • 专为第一人称流媒体视频设计,适合研究实时视觉理解

随着多模态模型快速发展,视频理解评估需求日益增长。然而现有基准在覆盖范围、任务多样性和场景适应性方面存在显著局限,难以准确评估模型的综合视频理解能力。为此,我们提出分层全景视频理解(H2VU)基准,用于评估通用视频与在线流媒体视频的理解能力。该基准具有三大特性:扩展视频时长,涵盖3秒至1.5小时的视频,填补当前基准的时间跨度空白;全面评估任务,除传统感知与推理任务外,新增反常识理解与轨迹状态追踪模块,测试模型超越常识知识的深层理解能力;丰富视频数据,扩充第一人称流媒体数据集,以适应当前AI代理的发展,支持从第一人称视角探索多模态模型在流媒体视频中的表现。大量实验表明,现有多模态大语言模型(MLLMs)在新提出的评估任务中仍有巨大提升空间。我们期望H2VU能推动视频理解研究的发展,为MLLMs提供全面深入的分析工具。

原文摘要 · Abstract (English)

With the rapid development of multimodal models, the demand for assessing video understanding capabilities has been steadily increasing. However, existing benchmarks for evaluating video understanding exhibit significant limitations in coverage, task diversity, and scene adaptability. These shortcomings hinder the accurate assessment of models' comprehensive video understanding capabilities. To tackle this challenge, we propose a hierarchical and holistic video understanding (H2VU) benchmark designed to evaluate both general video and online streaming video comprehension. This benchmark contributes three key features: Extended video duration: Spanning videos from brief 3-second clips to comprehensive 1.5-hour recordings, thereby bridging the temporal gaps found in current benchmarks. Comprehensive assessment tasks: Beyond traditional perceptual and reasoning tasks, we have introduced modules for countercommonsense comprehension and trajectory state tracking. These additions test the models' deep understanding capabilities beyond mere prior knowledge. Enriched video data: To keep pace with the rapid evolution of current AI agents, we have expanded first-person streaming video datasets. This expansion allows for the exploration of multimodal models' performance in understanding streaming videos from a first-person perspective. Extensive results from H2VU reveal that existing multimodal large language models (MLLMs) possess substantial potential for improvement in our newly proposed evaluation tasks. We expect that H2VU will facilitate advancements in video understanding research by offering a comprehensive and in-depth analysis of MLLMs.

视频理解多模态基准测试流媒体

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。