arXiv:2506.22992cs.AIcs.CL2025-06被引 3

MARBLE挑战多模态模型的复杂空间推理能力,揭示当前模型严重不足。

MARBLE: A Hard Benchmark for Multimodal Spatial Reasoning and Planning

  • 设计两个需多步规划的视觉空间任务,强制模型逐步推理。
  • 12个主流多模态模型在任务中表现接近随机,M-Cube全错(0%准确率)。
  • 暴露感知与推理双重瓶颈,适合研究下一代多模态智能体。

多模态信息处理与分步推理仍是人工智能发展的关键挑战。现有基准大多聚焦纯文本推理,或仅需从非文本模态直接检索答案,难以检验真正的复杂推理能力。本文提出MARBLE,一个针对多模态语言模型(MLLMs)的高难度推理基准,包含两个复杂任务:M-Portal和M-Cube,要求模型在空间、视觉和物理约束下制定并理解多步计划。实验发现,当前12个先进模型在M-Portal上表现接近随机,在M-Cube上准确率为0%。仅在简化子任务中部分模型优于随机基线,表明复杂推理仍属重大挑战。此外,模型在视觉输入信息提取上存在明显瓶颈。MARBLE揭示了当前多模态模型的局限性,旨在推动具备跨模态多步推理与规划能力的新一代模型发展。

原文摘要 · Abstract (English)

The ability to process information from multiple modalities and to reason through it step-by-step remains a critical challenge in advancing artificial intelligence. However, existing reasoning benchmarks focus on text-only reasoning, or employ multimodal questions that can be answered by directly retrieving information from a non-text modality. Thus, complex reasoning remains poorly understood in multimodal domains. Here, we present MARBLE, a challenging multimodal reasoning benchmark that is designed to scrutinize multimodal language models (MLLMs) in their ability to carefully reason step-by-step through complex multimodal problems and environments. MARBLE is composed of two highly challenging tasks, M-Portal and M-Cube, that require the crafting and understanding of multistep plans under spatial, visual, and physical constraints. We find that current MLLMs perform poorly on MARBLE -- all the 12 advanced models obtain near-random performance on M-Portal and 0% accuracy on M-Cube. Only in simplified subtasks some models outperform the random baseline, indicating that complex reasoning is still a challenge for existing MLLMs. Moreover, we show that perception remains a bottleneck, where MLLMs occasionally fail to extract information from the visual inputs. By shedding a light on the limitations of MLLMs, we hope that MARBLE will spur the development of the next generation of models with the ability to reason and plan across many, multimodal reasoning steps.

多模态推理空间规划模型评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。