arXiv:2508.20068cs.CLcs.CV2025-08被引 9

用真实认知测试评估多模态大模型的空间推理能力,发现其表现接近人类但仍有随机性。

11Plus-Bench: Demystifying Multimodal LLM Spatial Reasoning with Cognitive-Inspired Analysis

  • 基于真实标准化测试构建11Plus-Bench,含感知复杂度与推理过程细粒度标注
  • 14个主流多模态模型在该基准上表现远低于人类,但认知努力与任务复杂度相关
  • 模型推理结果随机性强,而人类表现可由抽象模式复杂度精准预测

人类的认知过程中,空间推理与感知紧密交织,但在多模态大语言模型(MLLMs)的评估中,这种相互作用仍鲜有研究。尽管近期的MLLM在推理任务上表现优异,但其是否具备类人空间认知能力仍是未解之谜。本文提出系统化评估框架,对比先进MLLMs与人类在空间推理上的表现。核心是11Plus-Bench,一个源自真实标准化空间能力测试的高质量基准,包含细粒度的专家标注——涵盖感知复杂度与推理过程,支持实例级行为分析。通过在14个MLLMs上开展大规模实验并结合人类评估,我们发现当前MLLM已显现空间认知的早期迹象:虽整体性能远逊于人类,但其认知努力与任务复杂度高度相关。然而,模型在具体实例上的表现近乎随机,而人类正确率则高度可预测,且受抽象模式复杂度显著影响。这些发现揭示了现有模型在空间推理上的潜力与局限,并为模型设计提供了可操作的改进方向。

原文摘要 · Abstract (English)

For human cognitive process, spatial reasoning and perception are closely entangled, yet the nature of this interplay remains underexplored in the evaluation of multimodal large language models (MLLMs). While recent MLLM advancements show impressive performance on reasoning, their capacity for human-like spatial cognition remains an open question. In this work, we introduce a systematic evaluation framework to assess the spatial reasoning abilities of state-of-the-art MLLMs relative to human performance. Central to our work is 11Plus-Bench, a high-quality benchmark derived from realistic standardized spatial aptitude tests. 11Plus-Bench also features fine-grained expert annotations of both perceptual complexity and reasoning process, enabling detailed instance-level analysis of model behavior. Through extensive experiments across 14 MLLMs and human evaluation, we find that current MLLMs exhibit early signs of spatial cognition. Despite a large performance gap compared to humans, MLLMs' cognitive profiles resemble those of humans in that cognitive effort correlates strongly with reasoning-related complexity. However, instance-level performance in MLLMs remains largely random, whereas human correctness is highly predictable and shaped by abstract pattern complexity. These findings highlight both emerging capabilities and limitations in current MLLMs' spatial reasoning capabilities and provide actionable insights for advancing model design.

空间推理多模态认知评估基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。