arXiv:2506.21044cs.LGcs.AI2025-06ICML被引 3

通过悔恨感知优化,高效发现多样技能,尤其在高维环境中表现更优。

Efficient Skill Discovery via Regret-Aware Optimization

  • 将技能发现建模为生成与策略学习的对抗博弈,用悔恨度衡量技能强度收敛程度。
  • 在高维环境中实现15%的零样本性能提升,且探索效率和多样性均优于基线方法。
  • 适合需要高效自适应技能学习的复杂强化学习任务,如机器人控制与多目标环境。

无监督技能发现旨在开放强化学习中学习多样且可区分的行为。现有方法主要通过纯探索、互信息优化和时序表征学习提升多样性,但在高维场景下仍存在效率瓶颈。本文将技能发现建模为技能生成与策略学习的最小-最大博弈,提出一种基于时序表征学习的悔恨感知方法,沿可升级策略强度方向扩展已发现的技能空间。核心思想是:技能发现与策略学习具有对抗性,强度弱的技能应持续探索,而强度已收敛的技能则减少探索。具体实现中,使用悔恨度评估强度收敛程度,并由可学习的技能生成器引导技能发现。为避免退化,技能生成来自一个可升级的生成器群体。在不同复杂度和维度规模的环境中进行实验,结果表明该方法在效率和多样性上均优于基线。此外,在高维环境中相比现有方法实现15%的零样本性能提升。

原文摘要 · Abstract (English)

Unsupervised skill discovery aims to learn diverse and distinguishable behaviors in open-ended reinforcement learning. For existing methods, they focus on improving diversity through pure exploration, mutual information optimization, and learning temporal representation. Despite that they perform well on exploration, they remain limited in terms of efficiency, especially for the high-dimensional situations. In this work, we frame skill discovery as a min-max game of skill generation and policy learning, proposing a regret-aware method on top of temporal representation learning that expands the discovered skill space along the direction of upgradable policy strength. The key insight behind the proposed method is that the skill discovery is adversarial to the policy learning, i.e., skills with weak strength should be further explored while less exploration for the skills with converged strength. As an implementation, we score the degree of strength convergence with regret, and guide the skill discovery with a learnable skill generator. To avoid degeneration, skill generation comes from an up-gradable population of skill generators. We conduct experiments on environments with varying complexities and dimension sizes. Empirical results show that our method outperforms baselines in both efficiency and diversity. Moreover, our method achieves a 15% zero shot improvement in high-dimensional environments, compared to existing methods.

强化学习技能发现高效探索时序表征

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。