arXiv:2511.18450cs.AI2025-11NeurIPS被引 6

用折纸任务评测多模态大模型的空间推理与数学约束能力

ORIGAMISPACE: Benchmarking Multimodal LLMs in Multi-Step Spatial Reasoning with Mathematical Constraints

  • 通过折纸图谱设计多步骤空间推理任务
  • 350个实例验证模型在数学约束下的表现
  • 适合研究视觉-语言模型与机器人规划的学者

空间推理是人工智能的关键能力,尤其在机器人、计算机视觉和自然语言理解中至关重要。然而,评估多模态大语言模型(MLLMs)在复杂空间推理任务中的表现仍面临挑战,尤其是在需要多步推理和精确数学约束的场景下。本文提出ORIGAMISPACE,一个全新的数据集与基准,通过折纸任务评估MLLMs的多步空间推理能力及处理数学约束的能力。数据集包含350个实例,每个实例包括严格格式化的折痕图(CP图)、展开图、完整折叠过程和最终折叠形状图像。我们设计了四项评估任务:模式预测、多步空间推理、空间关系预测和端到端CP代码生成。针对CP代码生成任务,我们构建了交互式环境,并探索使用强化学习训练MLLMs的可行性。通过对现有MLLMs的实验,初步揭示了这些模型在处理复杂空间推理任务时的优势与局限。

原文摘要 · Abstract (English)

Spatial reasoning is a key capability in the field of artificial intelligence, especially crucial in areas such as robotics, computer vision, and natural language understanding. However, evaluating the ability of multimodal large language models(MLLMs) in complex spatial reasoning still faces challenges, particularly in scenarios requiring multi-step reasoning and precise mathematical constraints. This paper introduces ORIGAMISPACE, a new dataset and benchmark designed to evaluate the multi-step spatial reasoning ability and the capacity to handle mathematical constraints of MLLMs through origami tasks. The dataset contains 350 data instances,each comprising a strictly formatted crease pattern (CP diagram), the Compiled Flat Pattern, the complete Folding Process, and the final Folded Shape Image. We propose four evaluation tasks: Pattern Prediction, Multi-step Spatial Reasoning, Spatial Relationship Prediction, and End-to-End CP Code Generation. For the CP code generation task, we design an interactive environment and explore the possibility of using reinforcement learning methods to train MLLMs. Through experiments on existing MLLMs, we initially reveal the strengths and weaknesses of these models in handling complex spatial reasoning tasks.

空间推理多模态模型折纸任务数学约束

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。