用对比学习提升人形机器人在复杂地形中的专家分工能力。
CMoE: Contrastive Mixture of Experts for Motion Control and Terrain Adaptation of Humanoid Robots
- 引入对比学习优化专家激活分布,增强地形特异性
- 实测可跨越20厘米高台阶与80厘米宽缝隙
- 适合研究机器人运动控制与自适应导航的开发者
为实现人形机器人在真实环境中的有效部署,必须自主应对多样且复杂的地形,包括突变场景。尽管原始的混合专家(MoE)框架理论上能建模不同地形特征,但实际中门控网络在各类地形上呈现近似均匀的专家激活,削弱了专家专属性,限制了模型表达能力。为此,我们提出CMoE——一种结合对比学习的单阶段强化学习框架,通过施加对比约束,最大化同一地形内专家激活的一致性,同时最小化不同地形间的相似性,从而促进专家对特定地形类型的分化。我们在Unitree G1人形机器人上进行了系列挑战性实验。结果表明,CMoE使机器人能够连续跨越高达20厘米的台阶和最大80厘米的缝隙,并在多种混合地形中实现稳定自然的步态,超越现有方法极限。为支持后续研究并推动社区发展,我们已公开代码。
原文摘要 · Abstract (English)
For effective deployment in real-world environments, humanoid robots must autonomously navigate a diverse range of complex terrains with abrupt transitions. While the Vanilla mixture of experts (MoE) framework is theoretically capable of modeling diverse terrain features, in practice, the gating network exhibits nearly uniform expert activations across different terrains, weakening the expert specialization and limiting the model's expressive power. To address this limitation, we introduce CMoE, a novel single-stage reinforcement learning framework that integrates contrastive learning to refine expert activation distributions. By imposing contrastive constraints, CMoE maximizes the consistency of expert activations within the same terrain while minimizing their similarity across different terrains, thereby encouraging experts to specialize in distinct terrain types. We validated our approach on the Unitree G1 humanoid robot through a series of challenging experiments. Results demonstrate that CMoE enables the robot to traverse continuous steps up to 20 cm high and gaps up to 80 cm wide, while achieving robust and natural gait across diverse mixed terrains, surpassing the limits of existing methods. To support further research and foster community development, we release our code publicly.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。