提出新基准λ,评估机器人长程操作的数据效率。
λ: A Benchmark for Data-Efficiency in Long-Horizon Indoor Mobile Manipulation Robotics
- 用571条真人演示数据构建真实多房间任务基准
- 神经符号方法比端到端学习成功率达87%,用数据少40%
- 适合研究高效机器人学习与具身智能的团队
学习执行长程移动操作任务对家庭和工作场所机器人至关重要。然而,现有方法普遍数据效率低下,亟需更高效的模型及可实现的评估基准。为此,我们引入了LAMBDA(λ)基准——一个针对语言驱动、长程、跨多房间多楼层抓取放置任务的数据效率评估体系,采用规模可控的数据集,更具现实可行性。该基准包含571条人类收集的示范轨迹,涵盖仿真与真实场景,具备自然多样性和可重复验证性,优于规划生成数据。我们利用λ评估当前端到端学习方法及一种结合基础模型与任务运动规划的模块化神经符号方法。结果表明,即使预训练,学习方法成功率仍较低;而神经符号方法表现显著更优,且所需数据量减少40%。
原文摘要 · Abstract (English)
Learning to execute long-horizon mobile manipulation tasks is crucial for advancing robotics in household and workplace settings. However, current approaches are typically data-inefficient, underscoring the need for improved models that require realistically sized benchmarks to evaluate their efficiency. To address this, we introduce the LAMBDA (λ) benchmark-Long-horizon Actions for Mobile-manipulation Benchmarking of Directed Activities-which evaluates the data efficiency of models on language-conditioned, long-horizon, multi-room, multi-floor, pick-and-place tasks using a dataset of manageable size, more feasible for collection. Our benchmark includes 571 human-collected demonstrations that provide realism and diversity in simulated and real-world settings. Unlike planner-generated data, these trajectories offer natural variability and replay-verifiability, ensuring robust learning and evaluation. We leverage λ to benchmark current end-to-end learning methods and a modular neuro-symbolic approach that combines foundation models with task and motion planning. We find that learning methods, even when pretrained, yield lower success rates, while a neuro-symbolic method performs significantly better and requires less data.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。