arXiv:2603.26779cs.CVcs.AI2026-03被引 1

给大模型加视觉模块,仍难解决空间推理难题。

Limits of Spatial Imagery Reasoning in Frontier LLM Models

  • 用外部3D渲染模块辅助大模型进行空间旋转推理
  • 最高准确率仅62.5%,远低于预期表现
  • 模型缺乏基础空间感知与动态视觉注意力机制

大型语言模型(LLMs)虽具备出色推理能力,但在需要心理模拟的空间任务(如心理旋转)上表现不佳。本文探究将外部‘图像模块’——一个可渲染和旋转3D模型的工具——赋予大模型,以充当‘认知假体’是否能弥补这一差距。我们采用双模块架构,让推理模块(一种多模态大模型)与图像模块协同完成3D模型旋转任务。实验结果显示,性能低于预期,准确率最高仅达62.5%。深入分析表明,即便将整体3D状态的维护与操作交由外部模块承担,系统依然失败。这揭示当前前沿模型缺乏接口图像所需的底层视觉-空间原语:一是对深度、运动及短时程动态预测等低层空间信号的敏感性不足;二是无法在图像上进行反思式推理,动态调整视觉焦点,并平衡图像信息与符号、联想内容之间的关系。

原文摘要 · Abstract (English)

Large Language Models (LLMs) have demonstrated impressive reasoning capabilities, yet they struggle with spatial tasks that require mental simulation, such as mental rotation. This paper investigates whether equipping an LLM with an external ``Imagery Module'' -- a tool capable of rendering and rotating 3D models -- can bridge this gap, functioning as a ``cognitive prosthetic.'' We conducted experiments using a dual-module architecture in which a reasoning module (an MLLM) interacts with an imagery module on 3D model rotation tasks. Performance was lower than expected, with accuracy reaching at most 62.5%. Further investigation suggests that even when the burden of maintaining and manipulating a holistic 3D state is outsourced, the system still fails. This reveals that current frontier models lack the foundational visual-spatial primitives required to interface with imagery. Specifically, they lack: (1) the low-level sensitivity to extract spatial signals such as (a) depth, (b) motion, and (c) short-horizon dynamic prediction; and (2) the capacity to reason contemplatively over images, dynamically shifting visual focus and balancing imagery with symbolic and associative information.

空间推理视觉感知大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。