提出新基准,揭示大模型在导航中空间推理能力不足。
VLN-MME: Diagnosing MLLMs as Language-guided Visual Navigation agents
- 构建统一评估框架,将视觉导航数据转为标准化测试集。
- 引入思维链和自省后性能反而下降,暴露模型上下文感知缺陷。
- 适合研究多模态大模型在具身智能中应用的学者参考。
多模态大语言模型(MLLMs)在众多视觉-语言任务中表现卓越,但在需要多轮对话、空间推理与序列动作预测的具身代理场景中,其性能仍需深入探索。本文通过引入一个统一且可扩展的评估框架,将传统导航数据集转化为标准化基准 VLN-MME,以零样本方式评估 MLLMs 作为具身代理的能力。该框架设计模块化、易用,支持跨架构、代理设计与任务的结构化对比与组件级消融实验。关键发现:在基线代理中引入思维链(CoT)推理与自省机制后,性能意外下降。这表明 MLLMs 在具身导航任务中存在严重的上下文感知缺陷——虽能遵循指令并组织输出,但其三维空间推理准确性极低。VLN-MME 为通用 MLLMs 在具身导航中的系统评估奠定基础,并揭示其序列决策能力的局限性。这些发现对 MLLM 后训练为具身智能体具有重要指导意义。
原文摘要 · Abstract (English)
Multimodal Large Language Models (MLLMs) have demonstrated remarkable capabilities across a wide range of vision-language tasks. However, their performance as embodied agents, which requires multi-round dialogue spatial reasoning and sequential action prediction, needs further exploration. Our work investigates this potential in the context of Vision-and-Language Navigation (VLN) by introducing a unified and extensible evaluation framework to probe MLLMs as zero-shot agents by bridging traditional navigation datasets into a standardized benchmark, named VLN-MME. We simplify the evaluation with a highly modular and accessible design. This flexibility streamlines experiments, enabling structured comparisons and component-level ablations across diverse MLLM architectures, agent designs, and navigation tasks. Crucially, enabled by our framework, we observe that enhancing our baseline agent with Chain-of-Thought (CoT) reasoning and self-reflection leads to an unexpected performance decrease. This suggests MLLMs exhibit poor context awareness in embodied navigation tasks; although they can follow instructions and structure their output, their 3D spatial reasoning fidelity is low. VLN-MME lays the groundwork for systematic evaluation of general-purpose MLLMs in embodied navigation settings and reveals limitations in their sequential decision-making capabilities. We believe these findings offer crucial guidance for MLLM post-training as embodied agents.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。