arXiv:2606.31807cs.RO2026-06中稿 · IROS 2026

用强化学习让机器人滑轮滑,比走路省50%能耗。

Reinforcement Learning-Based Control for an Inline Skating Humanoid Robot

  • 用奖励机制训练机器人自主学会轮滑推进策略。
  • 能耗降低50%,实机零样本迁移成功。
  • 适合对机器人动态平衡与高效运动感兴趣的团队。

随着类人机器人越来越动态,将其与强化学习结合可有效应对被动轮滑中的复杂、欠驱动力学问题。本文为一台改装了消费级轮滑鞋的类人机器人训练并部署了强化学习控制策略,实现了新型轮滑运动方式。不同于以往局限于四足机器人或主动驱动轮子的工作,本系统实现对轮滑的精确6自由度控制,执行动态的边缘推进策略。所有滑行策略均由奖励函数自动生成,无需人类运动数据、模仿学习或运动学先验。通过在训练和验证中使用球形与椭球形轮子模型,结合定制的成功导向指令课程与特殊滚动奖励,克服了被动轮子的固有不稳定性及仿真接触伪影。结果表明,该策略相比标准步行步态,能量消耗降低最高达50%。所获策略成功实现零样本迁移至真实机器人Booster T1硬件平台。实地部署展示出动态平衡能力、抵抗主动物理扰动的能力,以及高速转弯等灵巧运动策略。视频演示见:https://www.youtube.com/watch?v=-_APcOS7uFo。

原文摘要 · Abstract (English)

As humanoid robots become increasingly dynamic, coupling them with reinforcement learning offers a promising approach to solving the complex, underactuated mechanics of passive inline skating. Equipping a humanoid robot with passive inline skating wheels presents an opportunity to combine the versatile agility of humanoids with the high-speed, energy-efficient locomotion strategies utilized by human skaters. In this paper, we train and deploy a reinforcement learning control policy that enables novel locomotion strategies for a humanoid robot modified to equip consumer inline skates instead of conventional feet. Unlike previous work limited to quadrupedal robots or actively driven wheels, our system allows for precise 6-DoF control of the skates to execute dynamic, edge-driven propulsion strategies. Our skating strategies emerge entirely from our reward structure, without reliance on human motion data, imitation learning, or kinematic priors. We overcome the inherent instability of passive wheels and simulation contact artifacts by utilizing different geometric wheel models (spherical and ellipsoidal) during training and validation, along with a custom success-based command curriculum and a specialized rolling reward. Consequently, our policy demonstrates up to a 50% reduction in Cost of Transport (CoT) compared to standard walking gaits. The resulting policy successfully transfers zero-shot to the physical Booster T1 hardware. Real-world deployments demonstrate dynamic balance, the ability to reject active physical perturbations, and agile locomotion strategies capable of turning at speed. A video of our results can be found at https://www.youtube.com/watch?v=-_APcOS7uFo.

轮滑机器人强化学习能耗优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。