arXiv:2602.19357cs.CVcs.LG2026-02

测试大模型空间想象能力,发现其在对称变换和旋转任务上表现差。

MentalBlackboard: Evaluating Spatial Visualization via Mathematical Transformations

  • 构建开放题空间想象评测基准,涵盖折纸与打孔任务
  • 最高仅25%准确率,对称与旋转是主要难点
  • 适合研究视觉语言模型认知能力的学者参考

空间可视化是人类心智中对物体与动作的空间特征进行想象、变换和操作的能力,涉及感知与行为的内在联结。为探究当前最先进的视觉-语言模型(VLMs)是否具备此能力,我们开发了MentalBlackboard,一个用于纸张折叠与打孔测试的开放式空间可视化评估基准,包含预测与规划两个核心任务。预测实验显示,即使模型能正确预测展开步骤序列,仍难以处理对称变换;旋转操作显著削弱模型的物理情境理解能力。规划任务揭示模型在分析对称关系及执行多阶段对称过程方面存在局限,其中Claude Opus 4.1在规划任务中表现最佳,准确率为10%。表现最好的模型o3在无需空间可视化的泛化任务上达到71.6%峰值性能,但在基于文本的预测任务中准确率仅为25%。

原文摘要 · Abstract (English)

Spatial visualization is the mental ability to imagine, transform, and manipulate the spatial characteristics of objects and actions. This intelligence is a part of human cognition where actions and perception are connected on a mental level. To explore whether state-of-the-art Vision-Language Models (VLMs) exhibit this ability, we develop MentalBlackboard, an open-ended spatial visualization benchmark for Paper Folding and Hole Punching tests within two core tasks: prediction and planning. Our prediction experiments reveal that models struggle with applying symmetrical transformations, even when they predict the sequence of unfolding steps correctly. Also, rotations introduce a significant challenge to the physical situational awareness for models. The planning task reveals limitations of models in analyzing symmetrical relationships and in implementing the multi-stage symmetry process, with Claude Opus 4.1 achieving the highest planning score at an accuracy of 10\%. The top-performing model, o3, attains a peak performance of 71.6\% on the generalization task, which does not require spatial visualization but transfers spatial data; however, it achieves only 25\% accuracy on text-based prediction tasks.

空间推理视觉语言模型认知评测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。