arXiv:2412.12075cs.CV2024-12被引 82

为长视频理解设计新评测基准,防止模型靠猜题得分

CG-Bench: Clue-grounded Question Answering Benchmark for Long Video Understanding

  • 用线索定位机制评估模型是否真看懂长视频
  • 含1219段视频、12129个问答对,覆盖三类问题类型
  • 适合研究长视频理解与多模态大模型可信评估的学者

现有视频理解评测多聚焦短视频,长视频评测普遍依赖单选题,导致模型可仅通过短片段理解与排除法作答,无需真正理解内容。为此,我们提出CG-Bench,一个面向长视频的线索引导式问答评测基准。该基准强调模型检索相关线索的能力,提升评估可信度。包含1,219段人工标注视频,按14个主类、171个次类、638个细类进行精细分类,是目前最大的长视频分析基准。涵盖12,129个问答对,分感知、推理、幻觉三类问题。为弥补纯单选题评估缺陷,设计两种新型线索评估方法:线索引导白盒与黑盒评估,检验模型答案是否基于正确视频理解。在多个闭源与开源多模态大模型上测试显示,当前模型在长视频理解上表现显著低于短视频,且开源与商用模型间存在明显差距。所有标注数据与视频已公开于https://cg-bench.github.io/leaderboard/。

原文摘要 · Abstract (English)

Most existing video understanding benchmarks for multimodal large language models (MLLMs) focus only on short videos. The limited number of benchmarks for long video understanding often rely solely on multiple-choice questions (MCQs). However, because of the inherent limitation of MCQ-based evaluation and the increasing reasoning ability of MLLMs, models can give the current answer purely by combining short video understanding with elimination, without genuinely understanding the video content. To address this gap, we introduce CG-Bench, a novel benchmark designed for clue-grounded question answering in long videos. CG-Bench emphasizes the model's ability to retrieve relevant clues for questions, enhancing evaluation credibility. It features 1,219 manually curated videos categorized by a granular system with 14 primary categories, 171 secondary categories, and 638 tertiary categories, making it the largest benchmark for long video analysis. The benchmark includes 12,129 QA pairs in three major question types: perception, reasoning, and hallucination. Compensating the drawbacks of pure MCQ-based evaluation, we design two novel clue-based evaluation methods: clue-grounded white box and black box evaluations, to assess whether the model generates answers based on the correct understanding of the video. We evaluate multiple closed-source and open-source MLLMs on CG-Bench. Results indicate that current models significantly underperform in understanding long videos compared to short ones, and a significant gap exists between open-source and commercial models. We hope CG-Bench can advance the development of more trustworthy and capable MLLMs for long video understanding. All annotations and video data are released at https://cg-bench.github.io/leaderboard/.

长视频理解多模态模型评测基准线索引导

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。