arXiv:2512.22207cs.AI2025-12被引 4

用折纸任务评估多模态大模型的空间推理与2D转3D规划能力

GamiBench: Evaluating Spatial Reasoning and 2D-to-3D Planning Capabilities of MLLMs with Origami Folding Tasks

  • 设计折纸式折叠任务,通过多视角动态过程评估空间推理
  • 提出视点一致性与不可能折叠识别率新指标,覆盖186个真实与186个不可能图案
  • 揭示当前顶级模型在单步空间理解上仍存在显著不足,适合研究视觉-语言模型的学者

多模态大语言模型在感知和指令遵循方面表现良好,但在空间推理(即跨视角、跨时间的心理对象追踪与操作)方面仍存在困难。空间推理是人类智能的关键成分,但现有基准大多关注静态图像或最终输出,未能反映该技能的序列性和视角依赖性。为此,我们提出GamiBench,一个基于折纸任务的基准,用于评估多模态大模型在空间推理与2D到3D规划方面的能力。GamiBench包含186个真实与186个不可能的二维折痕图样,及其对应的三维折叠形态,涵盖六个不同视角下的三类视觉问答任务:预测三维折叠结构、区分有效视角、检测不可能图样。不同于仅评估最终结果的基准,GamiBench全面评估整个推理过程——包括跨视角一致性、通过不可能折叠检测实现的物理可行性判断,以及对中间折叠步骤的理解。它进一步引入新诊断指标:视点一致性(VC)与不可能折叠选择率(IFSR),以衡量模型对不同复杂度折叠任务的处理能力。实验表明,即使是最先进的模型如GPT-5和Gemini-2.5-Pro,在单步空间理解上依然表现不佳。这些贡献建立了评估多模态大模型几何理解与空间推理能力的标准框架。数据集与代码:https://github.com/stvngo/GamiBench。

原文摘要 · Abstract (English)

Multimodal large language models (MLLMs) are proficient in perception and instruction-following, but they still struggle with spatial reasoning: the ability to mentally track and manipulate objects across multiple views and over time. Spatial reasoning is a key component of human intelligence, but most existing benchmarks focus on static images or final outputs, failing to account for the sequential and viewpoint-dependent nature of this skill. To close this gap, we introduce GamiBench, a benchmark designed to evaluate spatial reasoning and 2D-to-3D planning in MLLMs through origami-inspired folding tasks. GamiBench includes 186 regular and 186 impossible 2D crease patterns paired with their corresponding 3D folded shapes, produced from six distinct viewpoints across three visual question-answering (VQA) tasks: predicting 3D fold configurations, distinguishing valid viewpoints, and detecting impossible patterns. Unlike previous benchmarks that assess only final predictions, GamiBench holistically evaluates the entire reasoning process--measuring cross-view consistency, physical feasibility through impossible-fold detection, and interpretation of intermediate folding steps. It further introduces new diagnostic metrics--viewpoint consistency (VC) and impossible fold selection rate (IFSR)--to measure how well models handle folds of varying complexity. Our experiments show that even leading models such as GPT-5 and Gemini-2.5-Pro struggle on single-step spatial understanding. These contributions establish a standardized framework for evaluating geometric understanding and spatial reasoning in MLLMs. Dataset and code: https://github.com/stvngo/GamiBench.

空间推理折纸任务多模态模型评估基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。