构建3D密室逃脱环境,评估模型多模态推理全过程。
EscapeCraft: A 3D Room Escape Environment for Benchmarking Complex Multimodal Reasoning Ability
- 设计可定制的3D环境,支持自由探索与过程追踪。
- 小模型可完成简单任务,难度提升后性能显著下降。
- 揭示模型在空间感知、道具使用等方面的差异性缺陷。
多模态大语言模型(MLLMs)的发展推动了真实世界与虚拟环境中复杂多模态推理任务的研究,这类任务需整合视觉感知、视觉推理、空间意识和目标推断等多种能力。然而,现有评估主要关注最终任务完成情况,常将评测简化为视觉定位或视觉问答等单一能力。对多模态环境中推理过程的全面、量化分析仍被忽视,而这对于理解模型行为与内在推理机制至关重要。为此,我们提出MM-Escape,一个可扩展的基准,受现实密室逃脱游戏启发,强调中间行为与最终任务完成并重。为此,我们开发了EscapeCraft,一个可自定义、开源的3D环境,使模型能进行自由探索以评估多模态推理能力。大量实验表明,无论规模大小,MLLMs均能成功完成最简单的密室逃脱任务,部分模型表现出类人探索策略。但随着任务难度增加,性能急剧下降。此外,我们观察到不同模型存在各异的性能瓶颈,暴露出多模态推理中的不同失败模式,如重复轨迹、因空间感知差而困于角落、以及无效使用已获取道具(如钥匙)。我们希望该工作揭示多模态推理的新挑战,并为提升MLLMs能力提供改进方向。
原文摘要 · Abstract (English)
The rapid advancing of Multimodal Large Language Models (MLLMs) has spurred interest in complex multimodal reasoning tasks in the real-world and virtual environment, which require coordinating multiple abilities, including visual perception, visual reasoning, spatial awareness, and target deduction. However, existing evaluations primarily assess the final task completion, often degrading assessments to isolated abilities such as visual grounding and visual question answering. Less attention is given to comprehensively and quantitatively analyzing reasoning process in multimodal environments, which is crucial for understanding model behaviors and underlying reasoning mechanisms beyond merely task success. To address this, we introduce MM-Escape, an extensible benchmark for investigating multimodal reasoning, inspired by real-world escape games. MM-Escape emphasizes intermediate model behaviors alongside final task completion. To achieve this, we develop EscapeCraft, a customizable and open environment that enables models to engage in free-form exploration for assessing multimodal reasoning. Extensive experiments show that MLLMs, regardless of scale, can successfully complete the simplest room escape tasks, with some exhibiting human-like exploration strategies. Yet, performance dramatically drops as task difficulty increases. Moreover, we observe that performance bottlenecks vary across models, revealing distinct failure modes and limitations in their multimodal reasoning abilities, such as repetitive trajectories without adaptive exploration, getting stuck in corners due to poor visual spatial awareness, and ineffective use of acquired props, such as the key. We hope our work sheds light on new challenges in multimodal reasoning, and uncovers potential improvements in MLLMs capabilities.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。