arXiv:2601.20668cs.RO2026-01被引 1

通过渐进式动作空间扩展,提升足式机器人控制的训练效率与性能。

GPO: Growing Policy Optimization for Legged Robot Locomotion and Whole-Body Control

  • 初期限制动作空间以高效收集数据,逐步放宽以增强探索。
  • 在四足和六足机器人上均实现更优性能,支持零样本硬件部署。
  • 无需环境特定调参,适用于各类足式机器人控制任务。

足式机器人强化学习策略的训练仍面临高维连续动作、硬件约束及探索受限等挑战。现有方法在基于位置控制且依赖环境特异性启发式(如奖励设计、课程学习、手动初始化)时表现良好,但在力矩控制下效果较差,因动作空间探索不足且梯度信号不充分。本文提出渐进式策略优化(GPO),采用随时间变化的动作变换,在训练初期限制有效动作空间,促进高效数据采集与策略学习,并逐步扩大以增强探索,提升预期回报。理论证明该变换保持PPO更新规则,仅引入有界且趋于消失的梯度失真,保障训练稳定。在四足与六足机器人上评估,包括仿真训练策略直接部署至硬件的零样本测试。使用GPO训练的策略始终表现更优,表明其为足式运动与全身控制提供通用、环境无关的优化框架。

原文摘要 · Abstract (English)

Training reinforcement learning (RL) policies for legged robots remains challenging due to high-dimensional continuous actions, hardware constraints, and limited exploration. Existing methods for locomotion and whole-body control work well for position-based control with environment-specific heuristics (e.g., reward shaping, curriculum design, and manual initialization), but are less effective for torque-based control, where sufficiently exploring the action space and obtaining informative gradient signals for training is significantly more difficult. We introduce Growing Policy Optimization (GPO), a training framework that applies a time-varying action transformation to restrict the effective action space in the early stage, thereby encouraging more effective data collection and policy learning, and then progressively expands it to enhance exploration and achieve higher expected return. We prove that this transformation preserves the PPO update rule and introduces only bounded, vanishing gradient distortion, thereby ensuring stable training. We evaluate GPO on both quadruped and hexapod robots, including zero-shot deployment of simulation-trained policies on hardware. Policies trained with GPO consistently achieve better performance. These results suggest that GPO provides a general, environment-agnostic optimization framework for learning legged locomotion.

强化学习足式机器人策略优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。