通过约束可访问状态提升跨动态强化学习性能
Policy Regularization on Globally Accessible States in Cross-Dynamics Reinforcement Learning
- 基于全局可访问状态设计正则化,避免不可达状态干扰
- 在多个基准上显著提升现有跨域策略迁移算法表现
- 可作为通用模块集成到离线与在线强化学习中
为从不同动态环境中学习,观察模仿(IfO)方法依赖专家状态轨迹,前提是恢复其他动态中的专家状态分布有助于当前环境的策略学习。然而,模仿学习天然存在性能上限。此外,当环境动态变化时,某些专家状态可能变得不可达,使其分布对模仿不再有用。为此,我们提出一种新框架,将奖励最大化与IfO结合,采用F-距离正则化策略优化。该框架对全局可访问状态——即在所有考虑动态中访问频率非零的状态——施加约束,缓解不可达状态带来的挑战。通过不同方式实例化F-距离,我们推导出两种理论分析,并开发出实用算法ASOR(可访问状态导向策略正则化)。ASOR可作为通用附加模块,集成至多种强化学习方法,包括离线和非策略强化学习。大量实验表明,其能有效提升最先进的跨域策略迁移算法性能。
原文摘要 · Abstract (English)
To learn from data collected in diverse dynamics, Imitation from Observation (IfO) methods leverage expert state trajectories based on the premise that recovering expert state distributions in other dynamics facilitates policy learning in the current one. However, Imitation Learning inherently imposes a performance upper bound of learned policies. Additionally, as the environment dynamics change, certain expert states may become inaccessible, rendering their distributions less valuable for imitation. To address this, we propose a novel framework that integrates reward maximization with IfO, employing F-distance regularized policy optimization. This framework enforces constraints on globally accessible states--those with nonzero visitation frequency across all considered dynamics--mitigating the challenge posed by inaccessible states. By instantiating F-distance in different ways, we derive two theoretical analysis and develop a practical algorithm called Accessible State Oriented Policy Regularization (ASOR). ASOR serves as a general add-on module that can be incorporated into various RL approaches, including offline RL and off-policy RL. Extensive experiments across multiple benchmarks demonstrate ASOR's effectiveness in enhancing state-of-the-art cross-domain policy transfer algorithms, significantly improving their performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。