arXiv:2506.09987cs.CVcs.LG2025-06被引 23

构建最小变化视频对基准,精准评估模型物理理解能力

A Shortcut-aware Video-QA Benchmark for Physical Understanding via Minimal Video Pairs

  • 设计最小变化视频对,迫使模型理解物理因果而非依赖表面线索
  • 人类在5.5万题上的准确率达92.9%,顶尖模型仅40.2%
  • 适合评估视频语言模型的物理推理能力,尤其对抗捷径学习

现有视频问答基准易因模型依赖表面视觉或文本线索导致评分虚高。本文提出最小视频对(MVP)基准,用于评估视频语言模型的物理理解能力。该基准包含5.5万个高质量多选题,覆盖9个数据源,涵盖第一人称和第三人称视频、机器人交互数据及认知科学直觉物理基准。每个样本均配有视觉相似但答案相反的最小变化对,要求模型在两对中均正确作答,否则将低于随机水平(25%)。人类在MVP上表现达92.9%,而最佳开源模型仅为40.2%。

原文摘要 · Abstract (English)

Existing benchmarks for assessing the spatio-temporal understanding and reasoning abilities of video language models are susceptible to score inflation due to the presence of shortcut solutions based on superficial visual or textual cues. This paper mitigates the challenges in accurately assessing model performance by introducing the Minimal Video Pairs (MVP) benchmark, a simple shortcut-aware video QA benchmark for assessing the physical understanding of video language models. The benchmark is comprised of 55K high-quality multiple-choice video QA examples focusing on physical world understanding. Examples are curated from nine video data sources, spanning first-person egocentric and exocentric videos, robotic interaction data, and cognitive science intuitive physics benchmarks. To mitigate shortcut solutions that rely on superficial visual or textual cues and biases, each sample in MVP has a minimal-change pair -- a visually similar video accompanied by an identical question but an opposing answer. To answer a question correctly, a model must provide correct answers for both examples in the minimal-change pair; as such, models that solely rely on visual or textual biases would achieve below random performance. Human performance on MVP is 92.9\%, while the best open-source state-of-the-art video-language model achieves 40.2\% compared to random performance at 25\%.

视频问答物理理解基准测试模型评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。