评测大模型在真实场景中一步步推理的能力,发现其潜力与短板。
How Good are Foundation Models in Step-by-Step Embodied Reasoning?
- 构建了涵盖10类任务、8种机器人形态的推理基准测试
- 超1100个样本验证模型在物理约束下的决策能力
- 分离感知与推理评估,帮助定位模型弱点
在物理世界中运行的具身智能体需做出既有效又安全、空间连贯且情境贴合的决策。尽管大型多模态模型(LMMs)在视觉理解与语言生成方面取得进展,其在真实具身任务中进行结构化推理的能力仍不明确。本文提出基础模型具身推理基准(FoMER),用于评估LMM在复杂具身决策场景中的推理能力。该基准包含超过1.1k个样本,覆盖10项任务与8种机器人形态,涉及三种不同机器人类型。我们设计了一个可分离感知基底与动作推理的新评估框架,并对多个主流LMM进行了实证分析。结果揭示了当前模型在具身推理中的潜力与局限,指明未来研究的关键挑战与方向。数据与代码将公开共享。
原文摘要 · Abstract (English)
Embodied agents operating in the physical world must make decisions that are not only effective but also safe, spatially coherent, and grounded in context. While recent advances in large multimodal models (LMMs) have shown promising capabilities in visual understanding and language generation, their ability to perform structured reasoning for real-world embodied tasks remains underexplored. In this work, we aim to understand how well foundation models can perform step-by-step reasoning in embodied environments. To this end, we propose the Foundation Model Embodied Reasoning (FoMER) benchmark, designed to evaluate the reasoning capabilities of LMMs in complex embodied decision-making scenarios. Our benchmark spans a diverse set of tasks that require agents to interpret multimodal observations, reason about physical constraints and safety, and generate valid next actions in natural language. We present (i) a large-scale, curated suite of embodied reasoning tasks, (ii) a novel evaluation framework that disentangles perceptual grounding from action reasoning, and (iii) empirical analysis of several leading LMMs under this setting. Our benchmark includes over 1.1k samples with detailed step-by-step reasoning across 10 tasks and 8 embodiments, covering three different robot types. Our results highlight both the potential and current limitations of LMMs in embodied reasoning, pointing towards key challenges and opportunities for future research in robot intelligence. Our data and code will be made publicly available.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。