arXiv:2505.17389cs.ROcs.AI2025-05被引 3

通过分层数据收集空间,用少量示范数据训练出更可靠的长时序机器人操作策略。

Bootstrapping Imitation Learning for Long-horizon Manipulation via Hierarchical Data Collection Space

  • 将复杂操作分解为多个原子任务,构建分层数据采集空间。
  • 仅需少量示范数据,即在5个真实场景中实现显著更高的成功率。
  • 适合需要高效数据利用的长时序机器人操作研究者。

模仿学习(IL)通过人类示范进行机器人操作是一种有前景的方法。尽管少量示范即可实现基本动作执行,但要达到高成功率和良好泛化能力,通常需付出高昂代价,如持续增加数据或反复进行人机协同的复杂软硬件系统迭代。本文重新思考数据采集流程中的状态/动作空间及导致非鲁棒动作预测的根本因素,提出一种分层数据收集空间(HD-Space),设计了一种简洁的数据采集方案,使模型能够主动获取高质量数据。具体而言,从高层视角将精细操作任务划分为多个关键原子任务,并为人类示范设计对应的原子状态/动作空间,以生成更具鲁棒性的模仿学习数据。我们在两个仿真环境和五个真实世界长时序操作任务上进行了实证评估,结果表明基于HD-Space的数据训练出的策略性能显著提升。该方法仅用少量示范数据即可训练出更强的策略,尤其适用于长时序操作任务。我们希望HD-Space能为优化数据质量与指导数据扩展提供新思路。

原文摘要 · Abstract (English)

Imitation learning (IL) with human demonstrations is a promising method for robotic manipulation tasks. While minimal demonstrations enable robotic action execution, achieving high success rates and generalization requires high cost, e.g., continuously adding data or incrementally conducting human-in-loop processes with complex hardware/software systems. In this paper, we rethink the state/action space of the data collection pipeline as well as the underlying factors responsible for the prediction of non-robust actions. To this end, we introduce a Hierarchical Data Collection Space (HD-Space) for robotic imitation learning, a simple data collection scheme, endowing the model to train with proactive and high-quality data. Specifically, We segment the fine manipulation task into multiple key atomic tasks from a high-level perspective and design atomic state/action spaces for human demonstrations, aiming to generate robust IL data. We conduct empirical evaluations across two simulated and five real-world long-horizon manipulation tasks and demonstrate that IL policy training with HD-Space-based data can achieve significantly enhanced policy performance. HD-Space allows the use of a small amount of demonstration data to train a more powerful policy, particularly for long-horizon manipulation tasks. We aim for HD-Space to offer insights into optimizing data quality and guiding data scaling. project page: https://hd-space-robotics.github.io.

模仿学习机器人操作数据效率分层结构

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。