arXiv:2510.10759cs.RO2025-10

让机器人学习时自动调节奖励权重,零违规完成真实世界行走

Gain Tuning Is Not What You Need: Reward Gain Adaptation for Constrained Locomotion Learning

  • 根据惩罚反馈在线调整奖励权重比例,避免越界
  • 实机测试中无摔倒,主奖励比顶尖方法高50%
  • 适合需要持续学习的物理机器人研发人员

现有机器人步态学习方法严重依赖离线设置奖励权重,且无法保证训练过程中的约束满足。为此,本文提出基于具身调节的奖励导向增益机制(ROGER),通过在具身交互过程中实时接收惩罚信号,动态调整奖励权重。当学习接近约束阈值时,正奖励与负奖励的权重比自动降低以避免违反约束;在安全状态下则提高该比例以优先优化性能。在60公斤四足机器人上,ROGER在多次学习试验中实现近零约束违规,并使主奖励提升达50%。在MuJoCo连续行走基准测试中,包括单腿跳跃器,相比默认奖励函数,其性能相当或最高提升100%,扭矩使用减少60%,姿态偏差降低。最终,无需任何跌倒,在真实世界中仅用一小时即从零开始完成四足机器人步态学习。本工作推动了满足约束的真实世界持续机器人步态学习,简化了奖励权重调参,有助于物理机器人及现实学习系统的发展。

原文摘要 · Abstract (English)

Existing robot locomotion learning techniques rely heavily on the offline selection of proper reward weighting gains and cannot guarantee constraint satisfaction (i.e., constraint violation) during training. Thus, this work aims to address both issues by proposing Reward-Oriented Gains via Embodied Regulation (ROGER), which adapts reward-weighting gains online based on penalties received throughout the embodied interaction process. The ratio between the positive reward (primary reward) and negative reward (penalty) gains is automatically reduced as the learning approaches the constraint thresholds to avoid violation. Conversely, the ratio is increased when learning is in safe states to prioritize performance. With a 60-kg quadruped robot, ROGER achieved near-zero constraint violation throughout multiple learning trials. It also achieved up to 50% more primary reward than the equivalent state-of-the-art techniques. In MuJoCo continuous locomotion benchmarks, including a single-leg hopper, ROGER exhibited comparable or up to 100% higher performance and 60% less torque usage and orientation deviation compared to those trained with the default reward function. Finally, real-world locomotion learning of a physical quadruped robot was achieved from scratch within one hour without any falls. Therefore, this work contributes to constraint-satisfying real-world continual robot locomotion learning and simplifies reward weighting gain tuning, potentially facilitating the development of physical robots and those that learn in the real world.

机器人控制强化学习约束学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。