评测大模型在基础视觉感知上的表现,发现其远不如人类。
The Percept-V Challenge: Can Multimodal LLMs Crack Simple Perception Problems?
- 用6000张程序生成的图像构建测评集,聚焦7类基础感知能力。
- 顶尖大模型在多物体场景下性能骤降,准确率远低于人类。
- 适合关注视觉理解缺陷的研究者和开发者参考。
认知科学将视觉感知视为智能发展的早期标志,其TVPS-4框架将人类感知能力分为七类,如视觉区分与形状恒常性。现有基准多聚焦高级推理与知识,缺乏对基础感知的系统评估。为此,我们提出Percept-V数据集,包含6000张程序生成的无污染图像,覆盖30个领域,每个领域测试一个或多个TVPS-4感知技能。由于任务设计简单且推理需求低,我们预期现代多模态大模型(MLLMs)应轻松应对。然而实验显示,主流开源与专有模型在Percept-V上表现较弱,远低于人类水平;随着图像中物体数量增加,模型性能迅速下降。此外,我们识别出所有模型普遍难以掌握的感知技能。
原文摘要 · Abstract (English)
Cognitive science research treats visual perception, the ability to understand and make sense of a visual input, as one of the early developmental signs of intelligence. Its TVPS-4 framework categorizes and tests human perception into seven skills such as visual discrimination, and form constancy. Do Multimodal Large Language Models (MLLMs) match up to humans in basic perception? Even though there are many benchmarks that evaluate MLLMs on advanced reasoning and knowledge skills, there is limited research that focuses evaluation on simple perception. In response, we introduce Percept-V, a dataset containing 6000 program-generated uncontaminated images divided into 30 domains, where each domain tests one or more TVPS-4 skills. Our focus is on perception, so we make our domains quite simple and the reasoning and knowledge required for solving them are minimal. Since modern-day MLLMs can solve much more complex tasks, our a-priori expectation is that they will solve these domains very easily. Contrary to our belief, our experiments show a weak performance of SoTA proprietary and open-source MLLMs compared to very high human performance on Percept-V. We find that as number of objects in the image increases, performance goes down rather fast. Our experiments also identify the perception skills that are considerably harder for all models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。