arXiv:2502.09696cs.CV2025-02中稿 · ICML被引 49

打造一个当前大模型几乎无法通过的视觉推理基准,检验真实视觉理解能力。

ZeroBench: An Impossible Visual Benchmark for Contemporary Large Multimodal Models

  • 用对抗筛选构建极难视觉任务,初始时顶尖模型全军覆没。
  • 一年后最强模型仅达6%通过率,远低于人类水平。
  • 适合评估模型真实视觉认知能力,研究者必看。

大型多模态模型在图像理解方面存在明显缺陷,某些衡量标准下其空间认知能力甚至不如幼儿或动物。尽管如此,它们在多数主流视觉基准上得分很高,且随着模型进步,这些基准的区分度迅速下降。为此,我们提出 ZeroBench——一个通过对抗筛选构建的轻量级视觉推理基准,设计为在发布之初对前沿大模型而言‘不可能’完成,初始最优表现仅为0%通过率(pass@1)和 pass^5。我们追踪了后续一年的进展,观察到当前最优模型达到6% pass^5 和19% pass@5,表明该基准具备持久有效性。我们在 ZeroBench 上评估了46个 LMMs,与人类基线对比,分析优劣势,描绘了一年来的视觉能力演进图景,并公开发布该基准:https://zerobench.github.io。

原文摘要 · Abstract (English)

Large Multimodal Models (LMMs) exhibit shortfalls when interpreting images and, by some measures, have poorer spatial cognition than young children or animals. Despite this, they attain high scores on many popular visual benchmarks, with headroom rapidly eroded by model progress. This creates a need for difficult benchmarks that remain relevant for longer. We introduce ZeroBench - a lightweight visual reasoning benchmark curated using adversarial filtering to be "impossible" for frontier LMMs at its original release, with initial SotA scores of 0% pass@1 and pass^5. We track progress on ZeroBench over the subsequent year, observing SotA reaching 6% pass^5 and 19% pass@5, indicating the potential longevity of the benchmark. We evaluate 46 LMMs on ZeroBench, compare performance to a human baseline, analyse strengths and weaknesses, chart a year of progress in visual capabilities, and publicly release ZeroBench at https://zerobench.github.io.

多模态视觉推理基准测试大模型评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。