测试大模型视觉推理能力,发现差距显著。
Mind's Eye: A Benchmark of Visual Abstraction, Transformation and Composition for Multimodal LLMs

- 设计八项视觉认知任务,按抽象、关系、变换分类。
- 人类准确率80%,顶尖模型不足50%。
- 适合研究多模态模型认知能力的学者参考。
多模态大语言模型(MLLM)在视觉语言基准上取得了显著进展,但其在视觉认知与空间推理方面的能力仍不明确。我们提出“Mind's Eye”,一个包含八项受经典人类智力测试启发的多选题基准,基于全新的“A-R-T”分类体系:抽象(Abstraction)、关系(Relation)、变换(Transformation)。这些任务考察流体智能的核心过程,如模式归纳、类比关系映射和心理变换。我们评估了多种闭源与开源的MLLM,并与人类参与者进行对比。人类平均准确率为80%,而表现最佳的MLLM仍低于50%。错误分析揭示其失败原因包括:(i) 视觉注意力分配不当,(ii) 内部感知操作能力弱,(iii) 对底层视觉概念抽象不足。结果表明,当前MLLM在视觉空间推理方面与人类存在明显差距,凸显了建立更贴近认知机制的评估框架的必要性。
原文摘要 · Abstract (English)
Multimodal large language models (MLLMs) have achieved impressive progress on vision language benchmarks, yet their capacity for visual cognitive and visuospatial reasoning remains less understood. We introduce "Mind's Eye", a multiple-choice benchmark of eight visuo-cognitive tasks inspired by classic human intelligence tests and organized under a novel "A-R-T" taxonomy: Abstraction, Relation, and Transformation. The tasks probe core processes of fluid intelligence such as pattern induction, analogical relation mapping, and mental transformation. We evaluate a diverse suite of closed-source and open-source MLLMs and compare their performance with human participants. Humans achieve 80% accuracy, while top performing MLLMs remain below 50%. Error analysis reveals failures in: (i) visual attention allocation, (ii) internal perceptual manipulation, and (iii) weak abstraction of underlying visual concepts. Our findings suggest that current MLLMs exhibit limited visuospatial reasoning capabilities, when compared with human participants, highlighting the need for more cognitively grounded evaluation frameworks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。