评测发现大模型空间推理能力远低于人类,仅60%准确率。
Do MLLMs Really Understand Space? A Mathematical Reasoning Evaluation
- 构建首个系统性数学空间推理数据集MathSpatial,含2000道题
- 16个主流模型平均准确率不足60%,GPT-5仍落后人类35个百分点
- 训练集可提升模型表现,适合研究多模态推理与认知建模者
多模态大模型在感知任务上表现优异,但在数学空间推理(即解析和操作二维、三维关系的能力)方面仍不清晰。人类解决教材类空间推理问题准确率超95%,而多数领先模型在相同任务上准确率未达60%。为探究此差距,本文提出首个大规模系统性数据集MathSpatial,包含两个互补子集:(i) MathSpatial-Bench,2000道经严格筛选的问题,覆盖3类11子类,剥离感知干扰;(ii) MathSpatial-Corpus,8000道带验证解和结构化推理链的问题,来源自真实教育材料,经多阶段质量控制。在MathSpatial-Bench上对16个主流模型的基准测试显示,空间推理仍是核心瓶颈:即使GPT-5也比人类低超过35个百分点,尤其在抽象推理任务表现差。进一步实验表明,在MathSpatial-Corpus上训练能跨模型族带来一致提升,证明其实际价值。数据集已公开:https://shuolucs.github.io/MathSpatial。
原文摘要 · Abstract (English)
Multimodal large language models (MLLMs) have achieved strong performance on perception-oriented tasks, yet their ability to perform mathematical spatial reasoning, defined as the capacity to parse and manipulate two- and three-dimensional relations, remains unclear. Humans easily solve textbook-style spatial reasoning problems with over 95\% accuracy, but we find that most leading MLLMs fail to reach even 60\% on the same tasks. This striking gap highlights spatial reasoning as a fundamental weakness of current models. To investigate this gap, we present \emph{MathSpatial}, the first large-scale and systematic dataset resource dedicated to mathematical spatial reasoning in MLLMs. \emph{MathSpatial} provides two complementary subsets: (i)~\emph{MathSpatial-Bench}, a rigorously curated evaluation set of 2{,}000 problems spanning 3 categories and 11 subtypes, designed to isolate spatial reasoning from perceptual noise; and (ii)~\emph{MathSpatial-Corpus}, a training set of 8{,}000 problems equipped with verified solutions and structured reasoning traces. All problems are sourced from authentic educational materials and undergo multi-stage quality control including deduplication, geometric consistency checking, and cross-validated solution verification. Benchmarking 16 leading MLLMs on \emph{MathSpatial-Bench} reveals that spatial reasoning remains a fundamental bottleneck: even GPT-5 lags behind human performance by over 35 percentage points, with particularly poor results on abstract deduction tasks. We further show that training on \emph{MathSpatial-Corpus} yields consistent improvements across model families, demonstrating the dataset's practical value for advancing spatial reasoning capabilities. \emph{MathSpatial} is publicly available at https://shuolucs.github.io/MathSpatial.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。