arXiv:2411.00081cs.ROcs.AI2024-11被引 83

构建10万条家庭任务数据集,评测机器人与人协作的规划推理能力。

PARTNR: A Benchmark for Planning and Reasoning in Embodied Multi-agent Tasks

  • 用大模型生成带物理约束的日常任务,结合仿真验证。
  • 现有模型协作效率仅为人协同的67%,任务追踪能力差。
  • 小模型微调后性能达大模型水平,推理速度快8.6倍。

我们提出一个面向人机协作家庭任务的规划与推理基准(PARTNR),旨在研究日常家务中的人机协同。该基准任务具备空间、时间及异构智能体能力约束等现实特征。采用基于大语言模型的半自动化任务生成流程,结合仿真进行任务扎根与验证。PARTNR是同类中规模最大的基准,包含10万条自然语言任务,覆盖60套房屋和5,819种独特物体。我们对当前主流大模型在规划、感知与技能执行三个维度进行了分析,发现其存在显著缺陷,如协作能力弱、任务追踪失败及错误恢复能力差。当大模型与真人协作时,所需步骤数为两人协作的1.5倍,比单人多1.1倍,表明仍有巨大改进空间。进一步研究表明,使用规划数据微调小型大模型,可达到比其大9倍的模型相当的性能,且推理速度提升8.6倍。总体而言,PARTNR揭示了协作式具身智能体面临的关键挑战,并旨在推动该方向的研究进展。

原文摘要 · Abstract (English)

We present a benchmark for Planning And Reasoning Tasks in humaN-Robot collaboration (PARTNR) designed to study human-robot coordination in household activities. PARTNR tasks exhibit characteristics of everyday tasks, such as spatial, temporal, and heterogeneous agent capability constraints. We employ a semi-automated task generation pipeline using Large Language Models (LLMs), incorporating simulation in the loop for grounding and verification. PARTNR stands as the largest benchmark of its kind, comprising 100,000 natural language tasks, spanning 60 houses and 5,819 unique objects. We analyze state-of-the-art LLMs on PARTNR tasks, across the axes of planning, perception and skill execution. The analysis reveals significant limitations in SoTA models, such as poor coordination and failures in task tracking and recovery from errors. When LLMs are paired with real humans, they require 1.5x as many steps as two humans collaborating and 1.1x more steps than a single human, underscoring the potential for improvement in these models. We further show that fine-tuning smaller LLMs with planning data can achieve performance on par with models 9 times larger, while being 8.6x faster at inference. Overall, PARTNR highlights significant challenges facing collaborative embodied agents and aims to drive research in this direction.

人机协作任务规划基准测试大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。