新视频理解基准测试揭示模型真实能力与排行榜分数的差距
Video-MME-v2: Towards the Next Stage in Benchmarks for Comprehensive Video Understanding

- 构建三级递进评估体系,从视觉信息整合到多模态推理逐步提升难度
- 采用分组非线性评估策略,只认可有逻辑支撑的答案,杜绝猜题得分
- 基于3300小时人工标注和5轮质检,适合评估下一代视频多模态大模型
随着视频理解技术快速发展,现有基准测试日趋饱和,导致排行榜分数虚高与实际模型能力之间出现显著差距。为此,我们提出 Video-MME-v2,一个全面评估视频理解鲁棒性与忠实度的基准。设计了渐进式三级评估体系,依次涵盖多点视觉信息聚合、时序动态建模与复杂多模态推理。不同于传统单题准确率,引入分组非线性评估策略,强制相关问题间的一致性与多步推理连贯性,仅对有合理依据的答案赋分。数据通过12名标注员与50名独立审校者完成,累计投入3300人小时,经最多5轮质量控制。实验显示,当前最佳模型Gemini-3-Pro与人类专家间仍有显著差距,且视觉信息聚合与时序建模错误会传导至高层推理造成瓶颈。研究还发现,基于思维链的推理高度依赖文本线索,在含字幕场景表现提升,但在纯视觉环境下反而下降。Video-MME-v2为下一代视频多模态大模型的发展提供了严苛的新测试平台。
原文摘要 · Abstract (English)
With the rapid advancement of video understanding, existing benchmarks are becoming increasingly saturated, exposing a critical discrepancy between inflated leaderboard scores and real-world model capabilities. To address this widening gap, we introduce Video-MME-v2, a comprehensive benchmark designed to rigorously evaluate the robustness and faithfulness of video understanding. To systematically evaluate model capabilities, we design a \textbf{progressive tri-level hierarchy} that incrementally increases the complexity of video comprehension, ranging from multi-point visual information aggregation, to temporal dynamics modeling, and ultimately to complex multimodal reasoning. Besides, in contrast to conventional per-question accuracy, we propose a \textbf{group-based non-linear evaluation} strategy that enforces both consistency across related queries and coherence in multi-step reasoning. It penalizes fragmented or guess-based correctness and assigns credit only to answers supported by valid reasoning. To guarantee data quality, Video-MME-v2 is constructed through a rigorously controlled human annotation pipeline, involving 12 annotators and 50 independent reviewers. Backed by \textbf{3,300 human-hours} and up to \textbf{5 rounds} of quality assurance, Video-MME-v2 aims to serve as one of the most authoritative video benchmarks. Extensive experiments reveal a substantial gap between current best model Gemini-3-Pro and human experts, and uncover a clear hierarchical bottleneck where errors in visual information aggregation and temporal modeling propagate to limit high-level reasoning. We further find that thinking-based reasoning is highly dependent on textual cues, improving performance with subtitles but sometimes degrading it in purely visual settings. By exposing these limitations, Video-MME-v2 establishes a demanding new testbed for the development of next-generation video MLLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。