发现现有视频理解评估存在55%可不看视频就能答对,模型实际能力远低于预期。
Video-Oasis: Rethinking Evaluation of Video Understanding
- 构建诊断工具Video-Oasis,系统筛查评估数据中的作弊线索
- 过滤后剩余挑战中,顶尖模型表现仅略高于随机猜测
- 揭示评估标准缺陷,为未来视频大模型测试提供新基准
视频理解的固有复杂性导致难以判断模型性能源自视觉感知、语言推理还是知识先验。尽管已有诸多高阶推理评测基准,但统一的评估标准仍被忽视。本文不新增基准,而是重新审视评估准则,提出Video-Oasis——一个可持续的诊断套件,用于系统审计现有视频理解基准。审计发现,55%的基准样本可在无视觉输入或时间上下文的情况下被解决。剔除这些捷径后,剩下的视频原生挑战暴露出显著的能力差距:当前最优模型的表现仅略高于随机猜测。基于此,我们利用提炼出的挑战作为测试平台,探究哪些算法设计选择有助于实现鲁棒的视频理解。期望本工作为构建严谨的视频评测基准和评估未来视频大模型提供实践基础。代码已开源:https://github.com/sejong-rcv/Video-Oasis。
原文摘要 · Abstract (English)
The inherent complexity of video understanding makes it difficult to determine whether Video-LLM benchmark performance stems from visual perception, linguistic reasoning, or knowledge priors. While many benchmarks have emerged to assess high-level reasoning, shared criteria for evaluating video understanding remain largely overlooked. Instead of introducing yet another benchmark, we take a step back to re-examine the criteria for evaluating video understanding. In this work, we introduce Video-Oasis, a sustainable diagnostic suite for systematically auditing existing video understanding benchmarks. This audit reveals that 55\% of existing benchmark samples are solvable without visual input or temporal context. After filtering these shortcuts, the remaining video-native challenges expose a substantial capability gap: state-of-the-art models perform only marginally above random guessing. Building on these findings, we use the distilled challenges as a testbed to investigate which algorithmic design choices contribute to robust video understanding. We hope our work provides a practical foundation for constructing rigorous video benchmarks and evaluating future Video-LLMs. Code is available at https://github.com/sejong-rcv/Video-Oasis.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。