无监督学习让四足机器人自动发现多样运动技能,效率更高且不乱学。
Diverse Skill Discovery for Quadruped Robots via Unsupervised Learning
- 用正交专家混合架构防止不同动作重叠,提升学习效率
- 多判别器设计避免奖励作弊,技能多样性提升18.3%
- 适合对机器人自主行为探索感兴趣的开发者
强化学习需专家精心设计奖励函数以诱导目标行为,而模仿学习依赖昂贵的任务专属数据。相比之下,无监督技能发现可通过内在动机驱动,学习一组多样且有用的行为,减轻负担。然而,现有方法存在两大缺陷:通常仅用单一策略掌握多样化行为,未建模行为间的共性与差异,导致学习效率低;且易出现奖励劫持,即奖励信号快速上升并收敛,但实际技能多样性不足。本文提出正交专家混合(OMoE)架构,防止不同行为向量坍缩至重叠表示,使单个策略能掌握广泛运动技能。此外,设计多判别器框架,各判别器在不同观测空间运作,有效缓解奖励劫持问题。在12自由度的Unitree A1四足机器人上验证,展示了多样化的运动技能。实验表明,该框架提升了训练效率,相比基线,状态空间覆盖率扩大18.3%。
原文摘要 · Abstract (English)
Reinforcement learning necessitates meticulous reward shaping by specialists to elicit target behaviors, while imitation learning relies on costly task-specific data. In contrast, unsupervised skill discovery can potentially reduce these burdens by learning a diverse repertoire of useful skills driven by intrinsic motivation. However, existing methods exhibit two key limitations: they typically rely on a single policy to master a versatile repertoire of behaviors without modeling the shared structure or distinctions among them, which results in low learning efficiency; moreover, they are susceptible to reward hacking, where the reward signal increases and converges rapidly while the learned skills display insufficient actual diversity. In this work, we introduce an Orthogonal Mixture-of-Experts (OMoE) architecture that prevents diverse behaviors from collapsing into overlapping representations, enabling a single policy to master a wide spectrum of locomotion skills. In addition, we design a multi-discriminator framework in which different discriminators operate on distinct observation spaces, effectively mitigating reward hacking. We evaluated our method on the 12-DOF Unitree A1 quadruped robot, demonstrating a diverse set of locomotion skills. Our experiments demonstrate that the proposed framework boosts training efficiency and yields an 18.3\% expansion in state-space coverage compared to the baseline.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。