用残差策略提升四足机器人感知控制的稳定性和效率
Residual Policy Learning for Perceptive Quadruped Control Using Differentiable Simulation
- 在物理可微仿真中学习残差策略,优化基础政策
- 四足机器人在分钟级完成端到端步态与导航训练
- 适合需要快速收敛的复杂机器人控制场景
一阶策略梯度(FoPG)算法如通过时间反向传播和解析策略梯度,利用局部仿真物理加速策略搜索,相比标准无模型强化学习显著提升样本效率。然而,在接触丰富的任务(如行走)中,FoPG算法常表现出不良的学习动态。以往方法通过算法或仿真创新缓解接触动力学问题。本文提出通过在简单基线策略上学习残差策略来引导策略搜索。对于四足行走,我们发现基于FoPG的残差策略学习(FoPG RPL)主要提升最终奖励,而非提高样本效率。此外,我们展示了将FoPG应用于基于像素的局部导航,使质点机器人在数秒内训练收敛。最后,通过分钟级端到端训练,验证了FoPG RPL在四足机器人步态与感知导航中的通用性。
原文摘要 · Abstract (English)
First-order Policy Gradient (FoPG) algorithms such as Backpropagation through Time and Analytical Policy Gradients leverage local simulation physics to accelerate policy search, significantly improving sample efficiency in robot control compared to standard model-free reinforcement learning. However, FoPG algorithms can exhibit poor learning dynamics in contact-rich tasks like locomotion. Previous approaches address this issue by alleviating contact dynamics via algorithmic or simulation innovations. In contrast, we propose guiding the policy search by learning a residual over a simple baseline policy. For quadruped locomotion, we find that the role of residual policy learning in FoPG-based training (FoPG RPL) is primarily to improve asymptotic rewards, compared to improving sample efficiency for model-free RL. Additionally, we provide insights on applying FoPG's to pixel-based local navigation, training a point-mass robot to convergence within seconds. Finally, we showcase the versatility of FoPG RPL by using it to train locomotion and perceptive navigation end-to-end on a quadruped in minutes.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。