arXiv:2410.14682cs.ROcs.AI2024-10被引 15

评测大模型在具身任务中的时空因果理解能力,发现其在复杂任务中表现显著下降。

ET-Plan-Bench: Embodied Task-level Planning Benchmark Towards Spatial-Temporal Cognition with Foundation Models

  • 构建多源仿真环境,让大模型动态交互并实时重规划。
  • 复杂任务中模型表现大幅下滑,暴露空间与因果推理短板。
  • 适合研究具身智能、具身规划与大模型泛化能力的学者使用。

近期大型语言模型(LLMs)的发展推动了其在具身任务中的应用,尤其聚焦于高层级任务规划与分解。为此,我们提出新的具身任务规划基准ET-Plan-Bench,专门用于评估LLMs在具身任务中的表现。该基准包含可控且多样化的任务,涵盖不同难度与复杂度,旨在检验模型在空间(目标物体的位置约束、遮挡关系)与时间/因果(动作序列逻辑)理解方面的双重能力。通过多源仿真器作为后端,可即时提供环境反馈,支持模型动态交互与必要时的重规划。我们在GPT-4、LLAMA和Mistral等开源与闭源基础模型上进行了评估。尽管在简单导航任务中表现良好,但面对需深层时空因果理解的任务时,性能显著下降。因此,本基准成为大规模、可量化、高度自动化、细粒度诊断的挑战框架,有望推动具身任务规划领域的发展。

原文摘要 · Abstract (English)

Recent advancements in Large Language Models (LLMs) have spurred numerous attempts to apply these technologies to embodied tasks, particularly focusing on high-level task planning and task decomposition. To further explore this area, we introduce a new embodied task planning benchmark, ET-Plan-Bench, which specifically targets embodied task planning using LLMs. It features a controllable and diverse set of embodied tasks varying in different levels of difficulties and complexities, and is designed to evaluate two critical dimensions of LLMs' application in embodied task understanding: spatial (relation constraint, occlusion for target objects) and temporal & causal understanding of the sequence of actions in the environment. By using multi-source simulators as the backend simulator, it can provide immediate environment feedback to LLMs, which enables LLMs to interact dynamically with the environment and re-plan as necessary. We evaluated the state-of-the-art open source and closed source foundation models, including GPT-4, LLAMA and Mistral on our proposed benchmark. While they perform adequately well on simple navigation tasks, their performance can significantly deteriorate when faced with tasks that require a deeper understanding of spatial, temporal, and causal relationships. Thus, our benchmark distinguishes itself as a large-scale, quantifiable, highly automated, and fine-grained diagnostic framework that presents a significant challenge to the latest foundation models. We hope it can spark and drive further research in embodied task planning using foundation models.

具身智能任务规划大模型评测时空认知

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。