arXiv:2602.00678cs.RO2026-02中稿 · Robotics Science a…被引 4

用专家混合模型和评估工具,让四足机器人在真实环境稳定跑动。

Toward Reliable Sim-to-Real Predictability for MoE-based Robust Quadrupedal Locomotion

  • 采用专家混合策略分解地形与指令建模,仅靠本体感知实现鲁棒运动
  • 在雪地、沙地、台阶等未见地形上成功通行,高速达4米/秒
  • 通过仿真测试评估迁移可靠性,减少物理试验风险

强化学习在仅依赖本体感知的四足敏捷运动中展现出巨大潜力。然而,复杂地形下的仿真-现实差距和奖励过拟合会导致策略无法迁移,而实际验证又存在风险且效率低。为此,我们提出一个统一框架:采用基于门控机制的专家混合(MoE)运动策略,实现多地形的鲁棒表征,并引入RoboGauge评估套件,量化仿真到现实的可迁移性。该方法通过一组专家专门处理不同地形与指令,仅凭本体感知即实现优异的部署鲁棒性和泛化能力。RoboGauge利用仿真内多维度测试(包含地形、难度等级与域随机化),提供基于本体感知的综合指标,无需大量物理实验即可可靠选择最佳策略。在Unitree Go2机器人上的实验表明,其可在未见过的复杂地形(如雪地、沙地、台阶、斜坡及30厘米障碍物)上稳定行走;高速测试中最高速度达4米/秒,并涌现出一种提升高速稳定性的小步态。

原文摘要 · Abstract (English)

Reinforcement learning has shown strong promise for quadrupedal agile locomotion, even with proprioception-only sensing. In practice, however, sim-to-real gap and reward overfitting in complex terrains can produce policies that fail to transfer, while physical validation remains risky and inefficient. To address these challenges, we introduce a unified framework encompassing a Mixture-of-Experts (MoE) locomotion policy for robust multi-terrain representation with RoboGauge, a predictive assessment suite that quantifies sim-to-real transferability. The MoE policy employs a gated set of specialist experts to decompose latent terrain and command modeling, achieving superior deployment robustness and generalization via proprioception alone. RoboGauge further provides multi-dimensional proprioception-based metrics via sim-to-sim tests over terrains, difficulty levels, and domain randomizations, enabling reliable MoE policy selection without extensive physical trials. Experiments on a Unitree Go2 demonstrate robust locomotion on unseen challenging terrains, including snow, sand, stairs, slopes, and 30 cm obstacles. In dedicated high-speed tests, the robot reaches 4 m/s and exhibits an emergent narrow-width gait associated with improved stability at high velocity.

四足机器人强化学习仿真迁移专家混合

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。