无需专家数据,让四足机器人自主学习多样技能并自然切换。
Continuous Control of Diverse Skills in Quadruped Robots Without Complete Expert Datasets
- 基于目标姿态自动筛选高质量轨迹,替代依赖完整专家数据。
- 在仿真和Solo 8机器人上实现平滑技能过渡与高保真动作再现。
- 适合希望降低数据成本、提升机器人泛化能力的研究者。
为四足机器人学习多样化技能面临显著挑战,如掌握复杂技能间的过渡以及应对不同难度任务。现有模仿学习方法虽有效,但依赖昂贵的专家数据集来复现专家行为。受内省学习启发,我们提出渐进式对抗自模仿技能迁移(PASIST),该方法无需完整专家数据集。PASIST通过预设目标姿态,自主探索并选择高质量轨迹,基于生成对抗自模仿学习(GASIL)框架实现。为进一步提升学习效果,我们设计了技能选择模块,通过平衡不同难度技能的权重以缓解模式崩溃问题。借助这些方法,PASIST能够生成对应目标姿态的技能,并实现平滑自然的技能间过渡。在仿真平台及Solo 8机器人上的评估验证了PASIST的有效性,为专家驱动学习提供高效替代方案。
原文摘要 · Abstract (English)
Learning diverse skills for quadruped robots presents significant challenges, such as mastering complex transitions between different skills and handling tasks of varying difficulty. Existing imitation learning methods, while successful, rely on expensive datasets to reproduce expert behaviors. Inspired by introspective learning, we propose Progressive Adversarial Self-Imitation Skill Transition (PASIST), a novel method that eliminates the need for complete expert datasets. PASIST autonomously explores and selects high-quality trajectories based on predefined target poses instead of demonstrations, leveraging the Generative Adversarial Self-Imitation Learning (GASIL) framework. To further enhance learning, We develop a skill selection module to mitigate mode collapse by balancing the weights of skills with varying levels of difficulty. Through these methods, PASIST is able to reproduce skills corresponding to the target pose while achieving smooth and natural transitions between them. Evaluations on both simulation platforms and the Solo 8 robot confirm the effectiveness of PASIST, offering an efficient alternative to expert-driven learning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。