arXiv:2508.03232cs.RO2025-08被引 12

构建复杂烹饪场景的长时序具身规划基准,推动智能体执行精细物理操作。

CookBench: A Long-Horizon Embodied Planning Benchmark for Complex Cooking Scenarios

  • 分两阶段:意图识别与具身交互,实现复杂任务理解与执行
  • 采用高保真仿真环境,支持空间级细粒度动作控制
  • 适合研究长时序规划、多模态大模型与具身智能的学者

具身规划致力于让智能体在复杂物理世界中完成长时序任务。然而,现有基准普遍局限于短时序任务和粗粒度动作。为此,我们提出CookBench,一个面向复杂烹饪场景的长时序规划基准。基于强大的Unity游戏引擎构建高保真仿真环境,定义前沿人工智能挑战。核心任务分为两阶段:第一阶段为意图识别,要求智能体准确解析用户复杂指令;第二阶段为具身交互,智能体需通过长时序、细粒度的物理动作序列完成烹饪目标。与现有基准不同,我们将动作粒度细化至空间级别,保留关键操作信息,同时抽象低层机器人控制。此外,我们提供完整工具集,包含统一API,支持宏观操作(如下单、购料)与丰富的微观具身动作,使研究者聚焦高层规划与决策。我们还对当前闭源的大语言模型与视觉-语言模型进行了深入分析,揭示其在复杂长时序任务中的主要缺陷与挑战。完整基准将开源,以促进未来研究。

原文摘要 · Abstract (English)

Embodied Planning is dedicated to the goal of creating agents capable of executing long-horizon tasks in complex physical worlds. However, existing embodied planning benchmarks frequently feature short-horizon tasks and coarse-grained action primitives. To address this challenge, we introduce CookBench, a benchmark for long-horizon planning in complex cooking scenarios. By leveraging a high-fidelity simulation environment built upon the powerful Unity game engine, we define frontier AI challenges in a complex, realistic environment. The core task in CookBench is designed as a two-stage process. First, in Intention Recognition, an agent needs to accurately parse a user's complex intent. Second, in Embodied Interaction, the agent should execute the identified cooking goal through a long-horizon, fine-grained sequence of physical actions. Unlike existing embodied planning benchmarks, we refine the action granularity to a spatial level that considers crucial operational information while abstracting away low-level robotic control. Besides, We provide a comprehensive toolset that encapsulates the simulator. Its unified API supports both macro-level operations, such as placing orders and purchasing ingredients, and a rich set of fine-grained embodied actions for physical interaction, enabling researchers to focus on high-level planning and decision-making. Furthermore, we present an in-depth analysis of state-of-the-art, closed-source Large Language Model and Vision-Language Model, revealing their major shortcomings and challenges posed by complex, long-horizon tasks. The full benchmark will be open-sourced to facilitate future research.

具身智能长时序规划多模态仿真基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。