新评测基准让长视频理解模型真实能力暴露无遗
VideoEval-Pro: Robust and Realistic Long Video Understanding Evaluation
- 改用开放式问答,杜绝猜测和先验干扰
- 模型在开放题上性能下降超25%,真实能力显现
- 更适合作为评估长视频理解能力的可靠标准
大型多模态模型(LMMs)在长视频理解(LVU)任务中表现强劲,推动了标准化评测基准的发展。然而,我们发现现有基准存在严重问题:多数依赖多项选择题(MCQs),易被猜中导致结果虚高;大量题目有强烈先验信息,使模型无需观看视频即可作答(如Gemini-1.5-Pro在Video-MME上仅凭随机帧即可达50%以上准确率)。此外,增加帧数并未带来预期性能提升,违背直觉。为此,我们提出VideoEval-Pro,一个更真实的长视频理解评测基准,采用开放式短答案问题,真正要求模型理解全视频内容。该基准通过感知与推理任务评估片段级与整体视频理解能力。对21个专有及开源视频LMMs的评估显示:(1) 模型在开放题上的性能相比MCQ大幅下降超过25%;(2) 高MCQ得分不意味着高开放题得分;(3) 相较于其他MCQ基准,VideoEval-Pro更能从增加输入帧数中获益。结果表明,VideoEval-Pro提供了更真实、可靠的长视频理解评估方式。
原文摘要 · Abstract (English)
Large multimodal models (LMMs) have recently emerged as a powerful tool for long video understanding (LVU), prompting the development of standardized LVU benchmarks to evaluate their performance. However, our investigation reveals a rather sober lesson for existing LVU benchmarks. First, most existing benchmarks rely heavily on multiple-choice questions (MCQs), whose evaluation results are inflated due to the possibility of guessing the correct answer; Second, a significant portion of questions in these benchmarks have strong priors to allow models to answer directly without even reading the input video. For example, Gemini-1.5-Pro can achieve over 50\% accuracy given a random frame from a long video on Video-MME. We also observe that increasing the number of frames does not necessarily lead to improvement on existing benchmarks, which is counterintuitive. As a result, the validity and robustness of current LVU benchmarks are undermined, impeding a faithful assessment of LMMs' long-video understanding capability. To tackle this problem, we propose VideoEval-Pro, a realistic LVU benchmark containing questions with open-ended short-answer, which truly require understanding the entire video. VideoEval-Pro assesses both segment-level and full-video understanding through perception and reasoning tasks. By evaluating 21 proprietary and open-source video LMMs, we conclude the following findings: (1) video LMMs show drastic performance ($>$25\%) drops on open-ended questions compared with MCQs; (2) surprisingly, higher MCQ scores do not lead to higher open-ended scores on VideoEval-Pro; (3) compared to other MCQ benchmarks, VideoEval-Pro benefits more from increasing the number of input frames. Our results show that VideoEval-Pro offers a more realistic and reliable measure of long video understanding, providing a clearer view of progress in this domain.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。