arXiv:2510.09543cs.RO2025-10

通过撞击缓解奖励提升机器人运动能效,更贴近动物自然步态。

Guiding Energy-Efficient Locomotion through Impact Mitigation Rewards

  • 引入物理启发的撞击缓解因子作为奖励,引导强化学习捕捉被动动态
  • 在多种奖励结构下实现最高32%的能耗降低(以运输成本衡量)
  • 适合关注仿生机器人能效与自然运动控制的研究者

动物通过隐含的被动动力学实现高效运动,这一现象吸引了机器人领域数十年的研究。近年来,结合对抗性运动先验(AMP)与强化学习(RL)的方法在复现动物自然运动方面取得进展,但这类模仿学习方法主要捕捉显式的运动模式(如步态),忽略了隐含的被动动力学。本文通过引入由撞击缓解因子(IMF)指导的奖励项,弥补了这一空白。IMF是一种基于物理的度量,用于量化机器人被动缓解冲击的能力。将IMF与AMP结合后,所提方法使强化学习策略同时学习动物参考运动的显式轨迹和隐含的被动动态。实验表明,在AMP及人工设计奖励结构下,运输成本(CoT)均实现最高达32%的能效提升。

原文摘要 · Abstract (English)

Animals achieve energy-efficient locomotion by their implicit passive dynamics, a marvel that has captivated roboticists for decades.Recently, methods incorporated Adversarial Motion Prior (AMP) and Reinforcement learning (RL) shows promising progress to replicate Animals' naturalistic motion. However, such imitation learning approaches predominantly capture explicit kinematic patterns, so-called gaits, while overlooking the implicit passive dynamics. This work bridges this gap by incorporating a reward term guided by Impact Mitigation Factor (IMF), a physics-informed metric that quantifies a robot's ability to passively mitigate impacts. By integrating IMF with AMP, our approach enables RL policies to learn both explicit motion trajectories from animal reference motion and the implicit passive dynamic. We demonstrate energy efficiency improvements of up to 32%, as measured by the Cost of Transport (CoT), across both AMP and handcrafted reward structure.

强化学习仿生运动能效优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。