构建可同时评估任务规划与底层控制的机器人基准测试
Mind and Motion Aligned: A Joint Evaluation IsaacSim Benchmark for Task Planning and Low-Level Policies in Mobile Manipulation
- 用模拟厨房环境统一评估任务规划与执行能力
- 支持500+复杂语言指令,含移动操作机器人
- 适合研究端到端机器人系统与多模态交互的学者
基准测试对机器人与具身AI研究至关重要。然而,现有基准要么侧重高层语言指令执行(假设底层执行完美),要么仅关注低层控制(使用简单单步指令),导致无法全面评估任务规划与物理执行并重的集成系统。为此,我们提出Kitchen-R——一个基于Isaac Sim模拟器构建的数字孪生基准,包含超过500条复杂语言指令,支持移动操作机器人。提供基于视觉-语言模型的任务规划策略和基于扩散策略的底层控制基线方法,并配备轨迹采集系统。该基准支持三种评估模式:独立评估规划模块、独立评估控制策略,以及关键的全流程系统联合评估。Kitchen-R填补了具身AI研究中的重要空白,实现更全面、真实的语言引导机器人智能评估。
原文摘要 · Abstract (English)
Benchmarks are crucial for evaluating progress in robotics and embodied AI. However, a significant gap exists between benchmarks designed for high-level language instruction following, which often assume perfect low-level execution, and those for low-level robot control, which rely on simple, one-step commands. This disconnect prevents a comprehensive evaluation of integrated systems where both task planning and physical execution are critical. To address this, we propose Kitchen-R, a novel benchmark that unifies the evaluation of task planning and low-level control within a simulated kitchen environment. Built as a digital twin using the Isaac Sim simulator and featuring more than 500 complex language instructions, Kitchen-R supports a mobile manipulator robot. We provide baseline methods for our benchmark, including a task-planning strategy based on a vision-language model and a low-level control policy based on diffusion policy. We also provide a trajectory collection system. Our benchmark offers a flexible framework for three evaluation modes: independent assessment of the planning module, independent assessment of the control policy, and, crucially, an integrated evaluation of the whole system. Kitchen-R bridges a key gap in embodied AI research, enabling more holistic and realistic benchmarking of language-guided robotic agents.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。