用无量纲状态空间提升强化学习控制器泛化能力
Improving Controller Generalization with Dimensionless Markov Decision Processes
- 将状态与动作空间无量纲化,使策略对环境变化保持不变性
- 在单个环境中训练的策略可应对上下文分布变化,表现更鲁棒
- 适合需要跨环境泛化的机器人控制任务
基于强化学习的控制器通常高度依赖特定训练环境,导致泛化能力差。本文提出一种基于模型的方法,将世界模型和策略均在无量纲的状态-动作空间中训练。为此,引入维度无量纲马尔可夫决策过程(Π-MDP),通过鲍金汉姆-Π定理对状态与动作空间进行非维化处理,使策略对底层动力学上下文变化保持等变性。本文提供通用框架,并应用于基于高斯过程模型的模型基策略搜索算法。在模拟的驱动摆和小车-摆系统上验证了该方法的有效性:仅在一个环境训练的策略,即可稳健应对上下文分布的变化。
原文摘要 · Abstract (English)
Controllers trained with Reinforcement Learning tend to be very specialized and thus generalize poorly when their testing environment differs from their training one. We propose a Model-Based approach to increase generalization where both world model and policy are trained in a dimensionless state-action space. To do so, we introduce the Dimensionless Markov Decision Process ($Π$-MDP): an extension of Contextual-MDPs in which state and action spaces are non-dimensionalized with the Buckingham-$Π$ theorem. This procedure induces policies that are equivariant with respect to changes in the context of the underlying dynamics. We provide a generic framework for this approach and apply it to a model-based policy search algorithm using Gaussian Process models. We demonstrate the applicability of our method on simulated actuated pendulum and cartpole systems, where policies trained on a single environment are robust to shifts in the distribution of the context.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。