通过局部动态规律发现可复用技能,提升离线层级强化学习效率
Exploiting Local Dynamics Regularity for Reusable Skills in Offline Hierarchical RL

- 基于局部动态一致性,识别不同情境下相似动作序列
- 在复杂人形环境实现有意义技能聚类,OGBench上性能提升
- 适用于需复用低层技能的层级强化学习算法
层级强化学习(HRL)通过发现并复用长时间跨度的技能,有望比非层级方法更高效地解决长时程任务。然而,如何获得真正可复用的技能仍是开放挑战。本文聚焦于局部动态的抽象:不同全局情境下的局部转移,往往需要相似的动作序列。通过将这些情境与其所需动作序列对齐,我们能够学习到哪些技能可复用以及在何处复用。该思想原则上可惠及多种HRL算法,其中高层策略需推理所使用的底层技能。提出的算法CARL(基于对比的动作表示用于可复用局部控制)在复杂人形环境中实现了有意义的技能聚类,并在集成至HIQL后,在OGBench基准上表现出显著的下游性能提升。
原文摘要 · Abstract (English)
Hierarchical Reinforcement Learning (HRL) promises to solve long-horizon Reinforcement Learning (RL) tasks more efficiently than non-hierarchical counterparts by discovering and reusing temporally-extended skills. However, obtaining skills that are actually reusable remains an open challenge. Towards this end, we focus on abstractions that exploit the intuition of local dynamics: local transitions in different global contexts require similar kinds of action sequences. By aligning these contexts with the action sequences they require, we are able to learn which skills to reuse and where to reuse them. In principle, this information should benefit many HRL algorithms, where high-level policies have to reason about the low-level skills they use. The resulting algorithm CARL (Contrastive Action-based Representations for Reusable Local Control) shows both qualitative clustering of meaningful skills in complex humanoid environments and improved downstream performance on the OGBench benchmark when integrated with HIQL.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。