arXiv:2411.19732cs.ROcs.AI2024-11被引 1

用平滑度感知优化提升机器人行走策略的泛化能力

Improving generalization of robot locomotion policies via Sharpness-Aware Reinforcement Learning

  • 在梯度强化学习中引入平滑度感知优化,寻找损失曲面更平坦的解
  • 在接触密集环境中显著提升对动作噪声和环境变化的鲁棒性
  • 兼顾高效训练与真实世界迁移性能,适合机器人控制研究者

强化学习通常需要大量训练数据。仿真到现实的迁移为解决机器人领域的这一挑战提供了有前景的途径。尽管可微分仿真器可通过精确梯度提升样本效率,但在接触密集环境中可能不稳定,并导致泛化能力差。本文提出将平滑度感知优化集成到基于梯度的强化学习算法中。仿真结果表明,该方法在接触密集环境中显著增强了策略对环境变化和动作扰动的鲁棒性,同时保持了一阶方法的样本效率。具体而言,相比标准一阶方法,本方法提升了动作噪声容忍度,并实现了与零阶方法相当的泛化性能。这一改进源于在损失景观中找到更平坦的极小值,这与更好的泛化能力相关。本工作为平衡高效学习与鲁棒仿真到现实迁移提供了一个有前景的解决方案,有望缩小仿真与真实性能之间的差距。

原文摘要 · Abstract (English)

Reinforcement learning often requires extensive training data. Simulation-to-real transfer offers a promising approach to address this challenge in robotics. While differentiable simulators offer improved sample efficiency through exact gradients, they can be unstable in contact-rich environments and may lead to poor generalization. This paper introduces a novel approach integrating sharpness-aware optimization into gradient-based reinforcement learning algorithms. Our simulation results demonstrate that our method, tested on contact-rich environments, significantly enhances policy robustness to environmental variations and action perturbations while maintaining the sample efficiency of first-order methods. Specifically, our approach improves action noise tolerance compared to standard first-order methods and achieves generalization comparable to zeroth-order methods. This improvement stems from finding flatter minima in the loss landscape, associated with better generalization. Our work offers a promising solution to balance efficient learning and robust sim-to-real transfer in robotics, potentially bridging the gap between simulation and real-world performance.

强化学习机器人控制泛化能力平滑度感知

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。