让机器人像人一样在复杂地形上自然切换走跑,保持稳定。
CoRe-MoE: Contrastive Reweighted Mixture of Experts for Multi-Terrain Humanoid Locomotion with Gait Adaptation

- 分两阶段训练:先学自然走跑,再用对比学习强化地形适应能力。
- 在楼梯、斜坡等10种地形上成功率超90%,动态稳定性显著提升。
- 适合需要跨地形自主行走的机器人研发人员参考。
人类主要依靠行走和奔跑穿越复杂地形,类似地,类人机器人也应能平滑切换走跑并保持自然稳定的运动。然而,将步态切换与多地形适应统一于单一策略中仍面临挑战,主要源于任务间梯度干扰及地形变化带来的分布偏移。尽管混合专家(MoE)架构可缓解多技能干扰,但直接联合训练常导致专家缺乏明确分工。为此,我们提出CoRe-MoE,一种两阶段强化学习框架,将步态生成与地形适应解耦。第一阶段学习稳定运动策略,生成自然的走跑行为及平滑过渡;第二阶段引入地形感知的MoE分支,通过对比学习训练门控网络,以构建结构化地形表征并促进专家分化。最终动作通过基线步态策略与地形感知分支加权融合获得,使策略在保持运动稳定性的同时适应复杂地形。大量仿真结果表明,该方法在成功率、运动稳定性和多地形适应性方面均优于基线。此外,在Unitree G1类人机器人上进行零样本部署验证,成功实现对台阶、斜坡、障碍物及非结构化户外地形的稳健走跑,同时保持精确落脚控制与动态稳定。
原文摘要 · Abstract (English)
Humans primarily rely on walking and running to traverse complex terrains. Similarly, humanoid robots should be able to smoothly transition between walking and running while maintaining natural and stable locomotion. However, unifying gait transition and multi-terrain adaptation within a single policy remains challenging due to gradient interference between tasks and the distribution shift caused by terrain variations. Although Mixture-of-Experts (MoE) architectures can mitigate multi-skill interference, direct joint training often fails to achieve clear expert specialization. To address these challenges, we propose CoRe-MoE, a two-stage reinforcement learning framework that decouples gait generation from terrain adaptation. In the first stage, a stable locomotion policy is learned to produce natural walking and running behaviors with smooth transitions. In the second stage, a terrain-aware MoE branch is introduced, and the gating network is trained with a contrastive objective to learn structured terrain representations and promote expert specialization. The final action is obtained through weighted fusion of the base gait policy and the terrain-aware branch, enabling the policy to preserve stable locomotion while adapting to complex terrains. Extensive simulation results demonstrate that the proposed method outperforms baseline approaches in terms of success rate, locomotion stability, and multi-terrain adaptability. Furthermore, zero-shot deployment on a Unitree G1 humanoid robot validates the effectiveness of our framework, achieving robust walking and running across stairs, slopes, steps, obstacles, and unstructured outdoor terrains while maintaining accurate foothold control and dynamic stability.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。