arXiv:2412.04323cs.LGcs.AI2024-12中稿 · ICRA被引 5

提出统一处理分布内与分布外动态的强化学习泛化框架

GRAM: Generalization in Deep RL with a Robust Adaptation Module

  • 设计鲁棒自适应模块,自动识别并响应不同环境动态
  • 联合训练实现分布内适应与分布外鲁棒性的平衡
  • 在仿真与真实四足机器人上验证跨场景泛化能力

深度强化学习在现实场景中的可靠部署需要在多种条件下具备泛化能力,包括训练中见过的分布内情况以及未见过的分布外情况。本文提出一种统一处理两类泛化问题的框架,引入鲁棒自适应模块,用于识别并响应分布内与分布外的环境动态,并设计联合训练流程,兼顾分布内适应性与分布外鲁棒性。所提出的GRAM算法在部署后表现出强泛化性能,通过大量仿真及四足机器人硬件行走实验得到验证。

原文摘要 · Abstract (English)

The reliable deployment of deep reinforcement learning in real-world settings requires the ability to generalize across a variety of conditions, including both in-distribution scenarios seen during training as well as novel out-of-distribution scenarios. In this work, we present a framework for dynamics generalization in deep reinforcement learning that unifies these two distinct types of generalization within a single architecture. We introduce a robust adaptation module that provides a mechanism for identifying and reacting to both in-distribution and out-of-distribution environment dynamics, along with a joint training pipeline that combines the goals of in-distribution adaptation and out-of-distribution robustness. Our algorithm GRAM achieves strong generalization performance across in-distribution and out-of-distribution scenarios upon deployment, which we demonstrate through extensive simulation and hardware locomotion experiments on a quadruped robot.

强化学习泛化能力自适应

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。