arXiv:2502.16435cs.CVcs.CL2025-02被引 6

测试发现大模型在基础视觉认知上远逊于人类,存在根本性缺陷。

Human Cognitive Benchmarks Reveal Foundational Visual Gaps in MLLMs

  • 构建可自动扩展的视觉认知基准VisFactor,涵盖20项人类视觉任务
  • 顶尖模型最高仅得54.0%,在空间推理等任务上普遍表现差
  • 揭示现有评测可能高估模型能力,适合关注视觉认知本质的研究者

人类感知发展遵循从基本视觉原语和格式塔原则到高层语义的自下而上层级。相比之下,当前多模态大模型直接在复杂下游任务上训练,往往跳过这些基础视觉能力。为系统研究此差距,我们提出VisFactor基准,将来自FRCT(一项成熟的认知心理学评估)的20个以视觉为中心的子测试数字化,覆盖人类视觉认知的四个领域。我们还设计算法,可自动生成并验证无限数量、难度可控的测试用例。利用VisFactor,我们评估了39个前沿多模态大模型,包括专有模型(如GPT、Gemini)和开源模型(如LLaMA、Qwen)。最佳模型得分仅为54.0%。分析显示该基准具有良好的内部一致性(Cronbach's alpha = 0.94)和结构效度(与现有视觉基准对比)。模型在心理旋转、空间关系推断和图形-背景区分等任务上持续表现不佳,无论模型规模或提示策略如何。这些发现表明,现有通用基准上的性能提升可能只是空中楼阁,而非真正掌握类人视觉认知。

原文摘要 · Abstract (English)

Humans develop perception through a bottom-up hierarchy: from basic primitives and Gestalt principles to high-level semantics. In contrast, current Multimodal Large Language Models (MLLMs) are trained directly on complex downstream tasks, often bypassing these foundational visual capabilities. To systematically investigate this gap, we introduce VisFactor, a benchmark that digitizes 20 vision-centric subtests from FRCT, a well-established cognitive psychology assessment spanning four domains of human visual cognition. Furthermore, we design algorithms to automatically construct and validate unlimited test cases with controllable difficulty. Using VisFactor, we evaluate 39 frontier MLLMs, including both proprietary (e.g., GPT, Gemini) and open-source (e.g., LLaMA, Qwen) models. The best model achieves a score of only 54.0%. Analysis reveals good internal consistency (Cronbach's alpha = 0.94) and construct validity (compared to existing vision benchmarks). Models consistently fail on tasks such as mental rotation, spatial relation inference, and figure-ground discrimination, regardless of model size or prompting strategy. These findings suggest that performance improvements on existing general benchmarks might represent castles in the air instead of a genuine mastery of human-like visual cognition.

视觉认知大模型评测基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。