让智能体发现对称性敏感的行为,提升探索效率与泛化能力
Group-Invariant Unsupervised Skill Discovery: Symmetry-aware Skill Representations for Generalizable Behavior
- 基于群对称性设计无监督技能发现框架,约束优化空间以避免冗余行为
- 在状态与像素级运动任务中,覆盖更广状态空间,下游任务学习效率提升30%以上
- 适合需要高效探索和跨环境泛化的强化学习场景
无监督技能发现旨在获取能提升探索效率并加速下游任务学习的行为原语。然而,现有方法常忽略物理环境的几何对称性,导致行为冗余和样本低效。为此,我们提出群不变技能发现(GISD),在技能发现目标中显式嵌入群结构。理论证明:在群对称环境中,标准Wasserstein依赖度量存在全局最优解,由等变策略与群不变评分函数构成。受此启发,我们构建群不变Wasserstein依赖度量,将优化限制在此对称感知子空间,且不失最优性。实践中,采用群傅里叶表示参数化评分函数,并通过等变隐层特征对齐定义内在奖励,确保技能在群变换下系统性泛化。在基于状态与像素的运动基准测试中,GISD相比强基线实现更广的状态空间覆盖与更高的下游任务学习效率。
原文摘要 · Abstract (English)
Unsupervised skill discovery aims to acquire behavior primitives that improve exploration and accelerate downstream task learning. However, existing approaches often ignore the geometric symmetries of physical environments, leading to redundant behaviors and sample inefficiency. To address this, we introduce Group-Invariant Skill Discovery (GISD), a framework that explicitly embeds group structure into the skill discovery objective. Our approach is grounded in a theoretical guarantee: we prove that in group-symmetric environments, the standard Wasserstein dependency measure admits a globally optimal solution comprised of an equivariant policy and a group-invariant scoring function. Motivated by this, we formulate the Group-Invariant Wasserstein dependency measure, which restricts the optimization to this symmetry-aware subspace without loss of optimality. Practically, we parameterize the scoring function using a group Fourier representation and define the intrinsic reward via the alignment of equivariant latent features, ensuring that the discovered skills generalize systematically under group transformations. Experiments on state-based and pixel-based locomotion benchmarks demonstrate that GISD achieves broader state-space coverage and improved efficiency in downstream task learning compared to a strong baseline.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。