arXiv:2410.07166cs.CLcs.AI2024-10NeurIPS被引 175

构建统一接口,系统评估大模型在具身决策中的表现。

Embodied Agent Interface: Benchmarking LLMs for Embodied Decision Making

论文配图:Embodied Agent Interface: Benchmarking LLMs for Embodied Decision Making
图 1 · 摘自论文原文
  • 设计通用接口,统一不同任务与输入输出规范。
  • 引入细粒度指标,识别幻觉、规划、动作等错误类型。
  • 帮助研究人员精准定位大模型短板,指导有效使用。

我们旨在评估大语言模型(LLMs)在具身决策中的表现。尽管已有大量研究将LLMs应用于具身环境的决策,但因任务领域、目标和输入输出方式各异,缺乏系统性理解。现有评估多依赖最终成功率,难以定位模型缺失能力的具体环节,阻碍了具身智能体对LLMs的有效利用。为此,我们提出通用接口(Embodied Agent Interface),可形式化多种任务及基于LLM的模块(如目标解析、子目标分解、动作序列生成、状态转移建模),并支持细粒度评估指标,涵盖幻觉、可操作性、规划等各类错误。该基准全面评估了LLMs在不同子任务中的表现,揭示其优劣势,为具身决策中选择性地使用LLMs提供依据。

原文摘要 · Abstract (English)

We aim to evaluate Large Language Models (LLMs) for embodied decision making. While a significant body of work has been leveraging LLMs for decision making in embodied environments, we still lack a systematic understanding of their performance because they are usually applied in different domains, for different purposes, and built based on different inputs and outputs. Furthermore, existing evaluations tend to rely solely on a final success rate, making it difficult to pinpoint what ability is missing in LLMs and where the problem lies, which in turn blocks embodied agents from leveraging LLMs effectively and selectively. To address these limitations, we propose a generalized interface (Embodied Agent Interface) that supports the formalization of various types of tasks and input-output specifications of LLM-based modules. Specifically, it allows us to unify 1) a broad set of embodied decision-making tasks involving both state and temporally extended goals, 2) four commonly-used LLM-based modules for decision making: goal interpretation, subgoal decomposition, action sequencing, and transition modeling, and 3) a collection of fine-grained metrics which break down evaluation into various types of errors, such as hallucination errors, affordance errors, various types of planning errors, etc. Overall, our benchmark offers a comprehensive assessment of LLMs' performance for different subtasks, pinpointing the strengths and weaknesses in LLM-powered embodied AI systems, and providing insights for effective and selective use of LLMs in embodied decision making.

具身智能大模型评估决策系统

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。