构建首个千小时级长视频基准,推动长视频理解研究
HLV-1K: A Large-scale Hour-Long Video Benchmark for Time-Specific Long Video Understanding
- 构建1009段时长超1小时的视频数据集,含近1.5万条带时间感知的问答对
- 覆盖帧级、事件内、事件间及长期推理等多层级任务,全面评估模型能力
- 适用于会议记录、直播、电影等场景的细粒度长视频理解研究
多模态大语言模型在深度视觉理解中日益重要,但时长超过一小时、包含数万帧画面的长视频理解仍面临挑战,主要源于长期分析困难、大模型效率低以及缺乏大规模基准数据集。本文聚焦于构建首个大规模小时级长视频基准数据集HLV-1K,包含1009段时长为一小时的视频,共14,847条高质量问题回答(QA)与多选题(MCQA)对,涵盖帧级、事件内、跨事件及长期推理等任务类型,并支持时间感知查询与多样化标注。我们使用现有先进方法对该基准进行评估,验证其在多层次、多任务下测试深度长视频理解能力的价值,有助于推动会议记录、直播、电影等场景下的细粒度长视频理解研究。
原文摘要 · Abstract (English)
Multimodal large language models have become a popular topic in deep visual understanding due to many promising real-world applications. However, hour-long video understanding, spanning over one hour and containing tens of thousands of visual frames, remains under-explored because of 1) challenging long-term video analyses, 2) inefficient large-model approaches, and 3) lack of large-scale benchmark datasets. Among them, in this paper, we focus on building a large-scale hour-long long video benchmark, HLV-1K, designed to evaluate long video understanding models. HLV-1K comprises 1009 hour-long videos with 14,847 high-quality question answering (QA) and multi-choice question asnwering (MCQA) pairs with time-aware query and diverse annotations, covering frame-level, within-event-level, cross-event-level, and long-term reasoning tasks. We evaluate our benchmark using existing state-of-the-art methods and demonstrate its value for testing deep long video understanding capabilities at different levels and for various tasks. This includes promoting future long video understanding tasks at a granular level, such as deep understanding of long live videos, meeting recordings, and movies.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。