arXiv:2603.15013cs.RO2026-03中稿 · ICRA

用强化学习让自行车在仿真中学会平衡,直接搬上真实车体

CycleRL: Sim-to-Real Deep Reinforcement Learning for Robust Autonomous Bicycle Control

  • 在高保真仿真环境里端到端训练控制策略
  • 仿真中平衡成功率99.90%,速度误差仅0.18米/秒
  • 适合想做机器人控制或仿真实战落地的研究者

自主自行车为城市出行与最后一公里物流提供了灵活解决方案。然而,传统控制方法常因系统欠驱动和非线性动力学,在模型不匹配时表现敏感,且对现实世界不确定性适应能力有限。为此,我们提出CycleRL,一种面向鲁棒自主自行车控制的完整仿真到现实框架。该方法在高保真NVIDIA Isaac Sim环境中建立感知到动作的直接映射,采用近端策略优化(PPO)优化控制策略,并设计包含平衡维持、速度追踪与转向控制的复合奖励函数。关键在于,通过系统性的领域随机化减少对精确建模的依赖,缩小仿真与现实差距,实现直接迁移。在仿真中,CycleRL表现优异:平衡成功率达99.90%,航向追踪误差1.15°,速度追踪误差0.18米/秒。结合硬件部署成功验证,证明深度强化学习在自主自行车控制中的有效性,其适应性优于传统方法。视频演示见https://cpnt-lab.github.io/CycleRL/。

原文摘要 · Abstract (English)

Autonomous bicycles offer a promising agile solution for urban mobility and last-mile logistics. However, conventional control strategies often struggle with underactuated nonlinear dynamics, suffering from sensitivity to model mismatches and limited adaptability to real-world uncertainties. To address this, we develop CycleRL, a comprehensive sim-to-real framework for robust autonomous bicycle control. Our approach establishes a direct perception-to-action mapping within the high-fidelity NVIDIA Isaac Sim environment, leveraging Proximal Policy Optimization (PPO) to optimize the control policy. The framework features a composite reward function tailored for concurrent balance maintenance, velocity tracking, and steering control. Crucially, systematic domain randomization is employed to reduce the reliance on precise system modeling, bridge the simulation-to-reality gap and facilitate direct transfer. In simulation, CycleRL achieves promising performance, including a 99.90% balance success rate, a heading tracking error of 1.15°, and a velocity tracking error of 0.18 m/s. These quantitative results, coupled with successful hardware deployment, validate DRL as an effective paradigm for autonomous bicycle control, offering superior adaptability over traditional methods. Video demonstrations are available at https://cpnt-lab.github.io/CycleRL/.

强化学习机器人控制仿真到现实自行车

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。