arXiv:2411.04998cs.CVcs.AI2024-11NeurIPS被引 139

构建小时级视频理解基准,揭示大模型与人类差距

HourVideo: 1-Hour Video-Language Understanding

  • 设计涵盖多任务的小时级视频理解评测体系
  • 大模型在复杂任务上仅达37.3%准确率,远低于人类85%
  • 适合研究长时序多模态理解、人机能力对比的学者

我们提出HourVideo,一个用于小时级视频-语言理解的基准数据集。该数据集包含500段来自Ego4D的主观视角视频,时长20至120分钟,涵盖总结、感知(回忆、追踪)、视觉推理(空间、时间、预测、因果、反事实)和导航(室间、物品检索)等新任务。共包含12,976个高质量五选一多项选择题。基准测试显示,多模态模型如GPT-4和LLaVA-NeXT仅略高于随机猜测水平。相比之下,人类专家显著优于当前最先进的长上下文多模态模型Gemini Pro 1.5(85.0% vs. 37.3%),暴露出多模态理解能力的巨大鸿沟。我们的基准、评估工具、提示模板及文档已公开于https://hourvideo.stanford.edu

原文摘要 · Abstract (English)

We present HourVideo, a benchmark dataset for hour-long video-language understanding. Our dataset consists of a novel task suite comprising summarization, perception (recall, tracking), visual reasoning (spatial, temporal, predictive, causal, counterfactual), and navigation (room-to-room, object retrieval) tasks. HourVideo includes 500 manually curated egocentric videos from the Ego4D dataset, spanning durations of 20 to 120 minutes, and features 12,976 high-quality, five-way multiple-choice questions. Benchmarking results reveal that multimodal models, including GPT-4 and LLaVA-NeXT, achieve marginal improvements over random chance. In stark contrast, human experts significantly outperform the state-of-the-art long-context multimodal model, Gemini Pro 1.5 (85.0% vs. 37.3%), highlighting a substantial gap in multimodal capabilities. Our benchmark, evaluation toolkit, prompts, and documentation are available at https://hourvideo.stanford.edu

视频理解多模态长时序基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。