新基准揭示视频大模型缺乏真实时空证据推理能力
VideoZeroBench: Probing the Limits of Video MLLMs with Spatio-Temporal Evidence Verification
- 设计分层评估体系,严格验证模型回答的时空证据
- 顶级模型在需精准定位时准确率低于17%,双重约束下不足1%
- 适合研究视频理解、可信AI和多模态推理的学者
近期视频多模态大模型在各类评测中表现优异,但现有评估存在两大缺陷:(1)分数虚高掩盖了细粒度视觉理解与推理缺陷;(2)答案正确性未验证模型是否识别支持预测的精确时空证据。为此,我们提出VideoZeroBench,一个针对长视频问答的分层基准,严格验证时空证据。包含500个手工标注问题,覆盖13个领域,每个问题配有时间区间和空间边界框作为证据。为分离回答生成、时间定位与空间定位能力,引入五级评估协议,逐步收紧证据要求。实验显示,即使Gemini-3-Pro在标准端到端问答(Level-3)下正确率也低于17%;当同时要求答案正确与精准时空定位时,性能骤降——无模型准确率超过1%,多数无法完成任何正确定位。结果揭示表面正确答案与真实证据推理间存在显著差距,表明基于证据的视频理解仍是长视频问答的核心瓶颈。我们进一步分析最小证据跨度、原子能力与推理范式,为未来研究提供洞见。基准与代码将公开。
原文摘要 · Abstract (English)
Recent video multimodal large language models achieve impressive results across various benchmarks. However, current evaluations suffer from two critical limitations: (1) inflated scores can mask deficiencies in fine-grained visual understanding and reasoning, and (2) answer correctness is often measured without verifying whether models identify the precise spatio-temporal evidence supporting their predictions. To address this, we present VideoZeroBench, a hierarchical benchmark designed for challenging long-video question answering that rigorously verifies spatio-temporal evidence. It comprises 500 manually annotated questions across 13 domains, paired with temporal intervals and spatial bounding boxes as evidence. To disentangle answering generation, temporal grounding, and spatial grounding, we introduce a five-level evaluation protocol that progressively tightens evidence requirements. Experiments show that even Gemini-3-Pro correctly answers fewer than 17% of questions under the standard end-to-end QA setting (Level-3). When grounding constraints are imposed, performance drops sharply: No model exceeds 1% accuracy when both correct answering and accurate spatio-temporal localization are required (Level-5), with most failing to achieve any correct grounded predictions. These results expose a significant gap between surface-level answer correctness and genuine evidence-based reasoning, revealing that grounded video understanding remains a bottleneck for long-video QA. We further analyze performance across minimal evidence spans, atomic abilities, and inference paradigms, providing insights for future research in grounded video reasoning. The benchmark and code will be made publicly available.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。