构建首个教学视频理解基准,测试模型对步骤化内容的时序推理能力。
InstructionBench: An Instructional Video Understanding Benchmark
- 用GPT-4生成开闭卷问答对,聚焦视频中的细粒度物体与事件推理
- 5000+问题覆盖700+视频,最佳模型准确率仅53.42%,暴露出时序理解短板
- 开源19000+问答对数据集,支持自动化生成,助力社区研究
尽管视频大语言模型(Video-LLMs)取得进展,但教学视频理解研究仍不足。为此,我们提出InstructionBench,一个教学视频理解基准,挑战模型在严格步骤流程视频中的高级时序推理能力。基于GPT-4,我们构建了开闭卷问答对,评估粗粒度事件级与细粒度物体级推理。通过筛选策略排除仅靠常识可答的问题,聚焦视觉感知与分析。最终基准包含5000+问题,覆盖700多段视频。我们在最新Video-LLMs上进行评估,发现闭源模型优于开源模型,但最优模型GPT-4o准确率仅为53.42%,表明时序推理仍有巨大差距。为推动研究,我们还构建了一个包含近2.5千段视频、19000+问答对的综合性教学视频数据集,采用自动化生成框架,丰富社区资源。所有数据已公开于https://huggingface.co/datasets/sunwhw/InstructionBench。
原文摘要 · Abstract (English)
Despite progress in video large language models (Video-LLMs), research on instructional video understanding, crucial for enhancing access to instructional content, remains insufficient. To address this, we introduce InstructionBench, an Instructional video understanding Benchmark, which challenges models' advanced temporal reasoning within instructional videos characterized by their strict step-by-step flow. Employing GPT-4, we formulate Q&A pairs in open-ended and multiple-choice formats to assess both Coarse-Grained event-level and Fine-Grained object-level reasoning. Our filtering strategies exclude questions answerable purely by common-sense knowledge, focusing on visual perception and analysis when evaluating Video-LLM models. The benchmark finally contains 5k questions across over 700 videos. We evaluate the latest Video-LLMs on our InstructionBench, finding that closed-source models outperform open-source ones. However, even the best model, GPT-4o, achieves only 53.42% accuracy, indicating significant gaps in temporal reasoning. To advance the field, we also develop a comprehensive instructional video dataset with over 19k Q&A pairs from nearly 2.5k videos, using an automated data generation framework, thereby enriching the community's research resources. All data are available at https://huggingface.co/datasets/sunwhw/InstructionBench.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。