测试大模型多步空间推理能力,发现其表现远低于人类。
LEGO-Puzzles: How Good Are MLLMs at Multi-Step Spatial Reasoning?
- 用乐高积木构建任务评估多步空间推理能力
- 最强模型在基础任务上仍比人类低20%以上
- 随着步骤增加,模型规划准确率迅速降为0%
许多现实世界应用如机器人控制、自动驾驶和自动化装配,需要跨多步骤的空间推理能力。然而,当前多模态大模型(MLLMs)是否具备这种能力尚不明确。受乐高积木搭建启发,我们提出LEGO-Puzzles基准,系统评估MLLMs从基础空间理解到多步规划的能力。该基准包含两个任务集:Elementary集涵盖11个视觉问答任务,共1,100个精心设计样本,用于测试基本空间推理能力;Planning集要求模型生成组装目标结构的分步计划,任务按规划视野分为不同子集,最长可达8步。对29个顶尖MLLMs的评估显示,即使最强模型在基础任务中也至少落后人类20%,规划准确率随步骤增加迅速降至0%,而人类参与者全部正确完成。将输出格式从选择题改为图像生成,模型性能进一步恶化,3步规划即达零准确率。总体表明,当前MLLMs在空间推理方面存在严重局限,亟需重大改进。
原文摘要 · Abstract (English)
Many real-world applications of spatial intelligence, such as robotic control, autonomous driving, and automated assembly, require spatial reasoning across multiple sequential steps. However, the extent to which current Multimodal Large Language Models (MLLMs) possess this capability remains largely unexplored. Inspired by LEGO construction, a recreational activity that critically relies on multi-step spatial reasoning, we introduce LEGO-Puzzles: a benchmark designed to systematically evaluate the spatial reasoning capabilities of MLLMs from basic spatial understanding to multi-step planning. LEGO-Puzzles contains two task sets. The Elementary set covers 11 visual question-answering (VQA) tasks with 1,100 carefully curated samples to test elementary spatial reasoning skills that are cruical for LEGO assembly. The Planning set directly requires the model to generate a step-by-step plan for assembling a target LEGO structure, where the tasks are organized into subsets with different planning horizons ranging up to 8. Our evaluation of 29 state-of-the-art MLLMs shows that even the strongest models struggle with elementary reasoning tasks in LEGO construction, falling at least 20% behind human performance. The planning accuracy also quickly drops to 0% as the number of planning steps increases, whereas our human participants solve all the tasks perfectly. Switching the output format from multiple choice to image generation degrades model performance even further, leading to zero accuracy even for planning 3 steps. Overall, LEGO-Puzzles reveals critical limitations in current MLLMs' spatial reasoning capabilities and highlights the need for substantial advances.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。