通过区分不同技能的探索区域,无监督发现多样技能行为。
Unsupervised Skill Discovery through Skill Regions Differentiation
- 用状态密度差异最大化来鼓励技能间状态多样性。
- 在图像任务中实现比现有方法更优的下游任务性能。
- 适合需要自动发现通用技能的强化学习场景。
无监督强化学习旨在发现多样化的行为以加速下游任务学习。以往方法多依赖熵驱动探索或激励驱动技能学习,但熵方法在大规模状态空间(如图像)中表现不佳,而基于互信息(MI)的方法在状态探索上存在局限。为此,我们提出一种新技能发现目标:最大化某一技能的状态密度与其它技能已探索区域之间的差异,从而促进技能间的状态多样性,类似初始的MI目标。为估计状态密度,我们构建了一个具有软模块化的条件自编码器,适用于高维空间中的不同技能策略。同时,为激励技能内探索,我们基于学习到的自编码器设计了一种内在奖励,类似于紧凑潜在空间中的计数基探索。在多个具有挑战性的状态和图像任务中进行大量实验表明,该方法能学习到有意义的技能,并在多种下游任务中取得优越性能。
原文摘要 · Abstract (English)
Unsupervised Reinforcement Learning (RL) aims to discover diverse behaviors that can accelerate the learning of downstream tasks. Previous methods typically focus on entropy-based exploration or empowerment-driven skill learning. However, entropy-based exploration struggles in large-scale state spaces (e.g., images), and empowerment-based methods with Mutual Information (MI) estimations have limitations in state exploration. To address these challenges, we propose a novel skill discovery objective that maximizes the deviation of the state density of one skill from the explored regions of other skills, encouraging inter-skill state diversity similar to the initial MI objective. For state-density estimation, we construct a novel conditional autoencoder with soft modularization for different skill policies in high-dimensional space. Meanwhile, to incentivize intra-skill exploration, we formulate an intrinsic reward based on the learned autoencoder that resembles count-based exploration in a compact latent space. Through extensive experiments in challenging state and image-based tasks, we find our method learns meaningful skills and achieves superior performance in various downstream tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。