自动调整奖励的智能框架,让机器人学得更快更稳。
Reward Training Wheels: Adaptive Auxiliary Rewards for Robotics Reinforcement Learning
- 用师生机制动态调节辅助奖励权重,随机器人能力变化自适应。
- 仿真中导航成功率提升2.35%,越野性能提高122.62%,训练快3倍。
- 适合需要高效训练且避免人工调参的机器人强化学习场景。
机器人强化学习常依赖精心设计的辅助奖励来弥补稀疏主目标和真实世界试错数据不足的问题。然而,这些辅助奖励需大量工程投入,可能引入人为偏差,且无法随机器人能力演进自适应。本文提出奖励训练轮(Reward Training Wheels, RTW),一种教师-学生框架,实现辅助奖励的自动化动态调整。具体而言,RTW 教师根据学生当前能力动态调节辅助奖励权重,以决定哪些方面需加强或减弱以优化主目标。我们在两个挑战性任务上验证了RTW:在高度受限空间中的导航与在垂直地形上的非公路车辆机动。仿真结果表明,RTW相比专家设计奖励,导航成功率达2.35%提升,越野性能提升122.62%,训练效率分别提高35%和3倍。物理机器人实验进一步验证其有效性,实现5/5完美成功(对比专家设计为2/5),并使车辆姿态角最大降低47.4%。
原文摘要 · Abstract (English)
Robotics Reinforcement Learning (RL) often relies on carefully engineered auxiliary rewards to supplement sparse primary learning objectives to compensate for the lack of large-scale, real-world, trial-and-error data. While these auxiliary rewards accelerate learning, they require significant engineering effort, may introduce human biases, and cannot adapt to the robot's evolving capabilities during training. In this paper, we introduce Reward Training Wheels (RTW), a teacher-student framework that automates auxiliary reward adaptation for robotics RL. To be specific, the RTW teacher dynamically adjusts auxiliary reward weights based on the student's evolving capabilities to determine which auxiliary reward aspects require more or less emphasis to improve the primary objective. We demonstrate RTW on two challenging robot tasks: navigation in highly constrained spaces and off-road vehicle mobility on vertically challenging terrain. In simulation, RTW outperforms expert-designed rewards by 2.35% in navigation success rate and improves off-road mobility performance by 122.62%, while achieving 35% and 3X faster training efficiency, respectively. Physical robot experiments further validate RTW's effectiveness, achieving a perfect success rate (5/5 trials vs. 2/5 for expert-designed rewards) and improving vehicle stability with up to 47.4% reduction in orientation angles.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。