arXiv:2410.16821cs.ROcs.SY2024-10中稿 · IROS 2024被引 3

用部分系统模型指导强化学习,提升采样效率与泛化能力

Guiding Reinforcement Learning with Incomplete System Dynamics

  • 分离已知与未知动态信息,构建嵌入式模型引导的强化学习框架
  • 在连续控制任务中样本效率显著优于标准强化学习方法
  • 适用于对数据效率和鲁棒性要求高的实际系统部署

无模型强化学习(RL)本质上是反应式方法,假设初始时对系统一无所知,完全依赖试错学习。这种方法面临样本效率低、泛化能力差以及需要精心设计奖励函数等挑战。而基于完整系统动力学的控制器则无需数据。本文针对介于两者之间的情况:既有足够信息表明完全无模型方法并非最优,又不足以进行完整控制器设计。通过精确解耦系统动力学中的已知与未知部分,我们构建了一个由部分模型引导的嵌入式控制器,从而提升增强型强化学习的学习效率。模块化设计使主流强化学习算法可被用于优化策略。仿真结果表明,该方法在连续控制任务上显著优于标准强化学习,性能也超过传统控制方法。真实地面车辆实验进一步验证了其在泛化性和鲁棒性方面的优势。

原文摘要 · Abstract (English)

Model-free reinforcement learning (RL) is inherently a reactive method, operating under the assumption that it starts with no prior knowledge of the system and entirely depends on trial-and-error for learning. This approach faces several challenges, such as poor sample efficiency, generalization, and the need for well-designed reward functions to guide learning effectively. On the other hand, controllers based on complete system dynamics do not require data. This paper addresses the intermediate situation where there is not enough model information for complete controller design, but there is enough to suggest that a model-free approach is not the best approach either. By carefully decoupling known and unknown information about the system dynamics, we obtain an embedded controller guided by our partial model and thus improve the learning efficiency of an RL-enhanced approach. A modular design allows us to deploy mainstream RL algorithms to refine the policy. Simulation results show that our method significantly improves sample efficiency compared with standard RL methods on continuous control tasks, and also offers enhanced performance over traditional control approaches. Experiments on a real ground vehicle also validate the performance of our method, including generalization and robustness.

强化学习系统建模样本效率

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。