构建细粒度视觉感知评测基准,揭示大模型在基础视觉理解上的短板。
PerceptionBench: Evaluating Atomic Visual Perception in Multimodal Large Language Models

- 基于错误诊断构建10项原子级视觉能力分类体系
- 3000道题仅测试单一感知能力,准确率普遍低于60%
- 发现感知幻觉是最弱环节,各模型能力差异显著
我们提出PerceptionBench,一个专门评估多模态大语言模型(MLLMs)基础视觉感知能力的基准。现有评测常混淆感知误差与推理或知识缺陷,而应用导向的基准又局限于碎片化领域。为此,PerceptionBench采用自下而上的方法:通过分析16个前沿模型在42个现有基准中的早期失败点,构建包含感知分支的错误分类体系,定义出10项原子级感知能力。据此设计3000道经验证的问题,每题仅考察单一能力,难度源于感知而非推理或知识。对16个前沿模型的评测显示,原子级感知仍基本未解决——无模型准确率超过60%,感知幻觉为平均最弱能力,且总体得分相似的模型表现出截然不同的能力分布。PerceptionBench因此提供了一个可测量、可诊断的视觉感知能力标准。
原文摘要 · Abstract (English)
We introduce PerceptionBench, a benchmark specifically designed to evaluate the atomic visual perception capabilities of Multimodal Large Language Models (MLLMs). Existing benchmarks often fail to isolate perception: holistic evaluations conflate perceptual errors with failures in reasoning or domain knowledge, while application-driven benchmarks only cover narrow, fragmented domains shaped by heuristic designs. To address these limitations, PerceptionBench adopts a bottom-up approach: by diagnosing the earliest failure points in the responses of frontier MLLMs across 42 existing benchmarks, we construct an error taxonomy whose perception branch defines ten atomic perceptual capabilities. Guided by this taxonomy, we construct 3,000 verified questions with short, unambiguous answers, each isolating a single capability, with difficulty stemming from perception rather than reasoning or knowledge. Benchmark results across sixteen frontier MLLMs reveal that atomic perception remains largely unsolved---no model reaches 60\% accuracy, perception-related hallucination is the weakest capability on average, and similar overall scores conceal sharply divergent capability profiles. PerceptionBench thus provides a capability-level standard for measuring and diagnosing the visual perception boundaries of MLLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。