arXiv:2506.05332cs.CVcs.CL2025-06NeurIPS被引 21

构建小时级视频数据集,实现长视频理解模型高效训练。

Unleashing Hour-Scale Video Training for Long Video-Language Understanding

  • 构建9700小时跨领域长视频指令数据集,支持1小时视频训练。
  • 提出Hour-LLaVA模型,1帧/秒下完成小时级视频推理。
  • 适合做长视频理解、多任务视频语言模型研究者使用。

近期长视频-语言理解基准推动了视频大模型的发展,但高质量长视频标注数据稀缺,导致小时级视频大模型训练仍不充分。为此,我们提出VideoMarathon,一个大规模小时级视频指令跟随数据集,包含约9700小时的长视频,单段时长3至60分钟,涵盖330万条高质量问答对,覆盖时间、空间、物体、动作、场景、事件六大核心主题。相比现有数据集,VideoMarathon将训练视频时长扩展至1小时,支持22项需短时与长时理解的任务。基于此,我们提出Hour-LLaVA,一种高效视频大模型,通过记忆增强模块自适应融合问题相关及时空信息,实现1帧/秒下的小时级视频训练与推理。实验表明,Hour-LLaVA在多个主流长视频-语言基准上表现最佳,验证了数据集质量与模型优势。

原文摘要 · Abstract (English)

Recent long-form video-language understanding benchmarks have driven progress in video large multimodal models (Video-LMMs). However, the scarcity of well-annotated long videos has left the training of hour-long Video-LMMs underexplored. To close this gap, we present VideoMarathon, a large-scale hour-long video instruction-following dataset. This dataset includes around 9,700 hours of long videos sourced from diverse domains, ranging from 3 to 60 minutes per video. Specifically, it contains 3.3M high-quality QA pairs, spanning six fundamental topics: temporality, spatiality, object, action, scene, and event. Compared to existing video instruction datasets, VideoMarathon significantly extends training video durations up to 1 hour, and supports 22 diverse tasks requiring both short- and long-term video comprehension. Building on VideoMarathon, we propose Hour-LLaVA, a powerful and efficient Video-LMM for hour-scale video-language modeling. It enables hour-long video training and inference at 1-FPS sampling by leveraging a memory augmentation module, which adaptively integrates question-relevant and spatiotemporally informative semantics from the cached full video context. In our experiments, Hour-LLaVA achieves the best performance on multiple representative long video-language benchmarks, demonstrating the high quality of the VideoMarathon dataset and the superiority of the Hour-LLaVA model.

长视频理解视频大模型数据集构建多任务学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。