用少量人类示范生成多样化数据,让机器人学会复杂长时序操作。
LodeStar: Long-horizon Dexterity via Synthetic Data Augmentation from Human Demonstrations
- 用基础模型自动分解人类示范为有意义技能
- 仅需少量真人演示即可生成丰富合成数据
- 适合需要高精度长时序操作的机器人研发
开发能够以人类级灵巧性稳健执行长时序操纵任务的机器人系统极具挑战,此类任务不仅要求物理灵巧性,还需无缝组合操纵技能并适应环境变化。尽管模仿学习具有潜力,但获取全面数据集成本高昂。本文提出LodeStar学习框架与系统,利用现成基础模型自动将任务示范分解为语义明确的技能,并通过强化学习从少量真人示范中生成多样化的合成示范数据集。这些模拟增强数据集支持鲁棒技能训练,结合技能路由变压器(SRT)策略可有效串联学习到的技能,完成复杂长时序操纵任务。在三个具有挑战性的真实世界长时序灵巧操纵任务上的实验表明,该方法显著优于以往基线,大幅提升了任务性能与鲁棒性。视频见 lodestar-robot.github.io。
原文摘要 · Abstract (English)
Developing robotic systems capable of robustly executing long-horizon manipulation tasks with human-level dexterity is challenging, as such tasks require both physical dexterity and seamless sequencing of manipulation skills while robustly handling environment variations. While imitation learning offers a promising approach, acquiring comprehensive datasets is resource-intensive. In this work, we propose a learning framework and system LodeStar that automatically decomposes task demonstrations into semantically meaningful skills using off-the-shelf foundation models, and generates diverse synthetic demonstration datasets from a few human demos through reinforcement learning. These sim-augmented datasets enable robust skill training, with a Skill Routing Transformer (SRT) policy effectively chaining the learned skills together to execute complex long-horizon manipulation tasks. Experimental evaluations on three challenging real-world long-horizon dexterous manipulation tasks demonstrate that our approach significantly improves task performance and robustness compared to previous baselines. Videos are available at lodestar-robot.github.io.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。