arXiv:2501.05037cs.CVcs.LG2025-01被引 14

构建超长视频理解数据集,提升模型长期上下文与推理能力

LongViTU: Instruction Tuning for Long-Form Video Understanding

  • 将视频分层组织并自修正生成高质量问答对
  • 平均4.6分钟长上下文,涵盖因果、规划等复杂推理
  • 适合研究长视频理解与指令微调的学者和工程师

本文提出LongViTU,一个大规模(约12.1万组问答对,约900小时视频)的自动构建长视频理解数据集。我们采用系统性方法,将视频组织为分层树结构以生成问答对,并引入自修正机制保障质量。每组问答对具备:1)长期上下文(平均证书长度4.6分钟);2)丰富知识与凝练推理(常识、因果、规划等)。每个问题均附带明确的时间戳标注。通过大规模人工评估验证了数据集质量。为评估长期上下文与凝练推理的挑战,我们手动构建子集作为基准测试。使用先进开源模型(LongVU)、专有模型(Gemini-1.5-Pro)及人工标注者进行评估,GPT-4得分分别为49.9、52.3和81.0,凸显任务难度。在LongViTU上对LongVU和LLaVA-Video进行监督微调,分别在多个长视频理解基准(EgoSchema、VideoMME-Long、MLVU、LVBench)上平均提升2.5%和3.7%。

原文摘要 · Abstract (English)

This paper introduces LongViTU, a large-scale (~121k QA pairs, ~900h videos), automatically generated dataset for long-form video understanding. We propose a systematic approach that organizes videos into a hierarchical tree structure for QA generation and incorporates self-revision mechanisms to ensure high-quality QA pairs. Each QA pair in LongViTU features: 1) long-term context (average certificate length of 4.6 minutes); 2) rich knowledge and condensed reasoning (commonsense, causality, planning, etc.)). We also offer explicit timestamp annotations of relevant events for each QA pair. We have conducted extensive human studies on LongViTU, and the results prove the quality of our dataset. To better evaluate the challenges posed by LongViTU's emphasis on long-term context and condensed reasoning, we manually curate a subset of LongViTU into a benchmark. Evaluations using a state-of-the-art open-source model (LongVU), a proprietary model (Gemini-1.5-Pro), and human annotators yield GPT-4 scores of 49.9, 52.3, and 81.0, respectively, underscoring the substantial difficulty presented by LongViTU questions. Performing supervised fine-tuning (SFT) of LongVU and LLaVA-Video on LongViTU data results in average performance gains of 2.5% and 3.7%, respectively, across a suite of long video understanding benchmarks (EgoSchema, VideoMME-Long, MLVU, LVBench).

视频理解长上下文指令微调数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。