通过动态奖励与轮腿协同,让摔倒机器人快速自愈。
Learning to Recover: Dynamic Reward Shaping with Wheel-Leg Coordination for Fallen Robots
- 用动态奖励和课程学习,让机器人自主学会多样恢复动作。
- 在双足平台上实现99.1%和97.8%的恢复成功率,无需调参。
- 轮腿协作降低15.8%~26.2%关节扭矩,提升能量利用效率。
轮腿机器人兼具腿部灵活性与轮子高速性,但传统方法依赖预设动作、简化动力学或稀疏奖励,难以生成鲁棒恢复策略。本文提出基于阶段性动态奖励塑形与课程学习的端到端训练框架,通过非对称演员-评论家结构利用仿真中的特权信息加速训练,并引入噪声观测增强对不确定性的鲁棒性。实验表明,轮腿协同可减少15.8%至26.2%的关节扭矩消耗,并通过能量传递机制提升稳定性能。在两个不同四足平台上的大量测试中,恢复成功率分别达到99.1%和97.8%,且无需针对平台进行调参。补充材料见https://boyuandeng.github.io/L2R-WheelLegCoordination/
原文摘要 · Abstract (English)
Adaptive recovery from fall incidents are essential skills for the practical deployment of wheeled-legged robots, which uniquely combine the agility of legs with the speed of wheels for rapid recovery. However, traditional methods relying on preplanned recovery motions, simplified dynamics or sparse rewards often fail to produce robust recovery policies. This paper presents a learning-based framework integrating Episode-based Dynamic Reward Shaping and curriculum learning, which dynamically balances exploration of diverse recovery maneuvers with precise posture refinement. An asymmetric actor-critic architecture accelerates training by leveraging privileged information in simulation, while noise-injected observations enhance robustness against uncertainties. We further demonstrate that synergistic wheel-leg coordination reduces joint torque consumption by 15.8% and 26.2% and improves stabilization through energy transfer mechanisms. Extensive evaluations on two distinct quadruped platforms achieve recovery success rates up to 99.1% and 97.8% without platform-specific tuning. The supplementary material is available at https://boyuandeng.github.io/L2R-WheelLegCoordination/
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。