arXiv:2601.23156cs.LGcs.FL2026-01中稿 · ICML被引 1

无监督发现强化学习中的层次化技能结构,无需标签或奖励

Unsupervised Hierarchical Skill Discovery

  • 用语法驱动方法自动分割无标签轨迹为可复用技能
  • 在高维像素环境(如Minecraft)中生成更结构化、语义合理的层次
  • 发现的技能层次能加速下游任务学习,适合复杂行为建模研究

我们研究强化学习中无监督的技能分割与层次结构发现。现有方法多依赖动作标签、奖励信号或人工标注,限制了适用性。本文提出一种基于语法的方法,从无标签轨迹中分割出技能,并构建其层次结构,捕捉低层行为及其组合形成的高层技能。我们在高维像素环境(包括Craftax和完整未修改版Minecraft)上评估,使用技能分割、复用性和层次质量等指标,结果表明本方法始终优于现有基线,生成更结构化且语义合理的内容。作为概念验证,我们进一步证明这些发现的层次结构能显著加速并稳定下游强化学习任务的学习过程。

原文摘要 · Abstract (English)

We consider the problem of unsupervised skill segmentation and hierarchical structure discovery in reinforcement learning. While recent approaches have sought to segment trajectories into reusable skills or options, most rely on action labels, rewards, or handcrafted annotations, limiting their applicability. We propose a method that segments unlabelled trajectories into skills and induces a hierarchical structure over them using a grammar-based approach. The resulting hierarchy captures both low-level behaviours and their composition into higher-level skills. We evaluate our approach in high-dimensional, pixel-based environments, including Craftax and the full, unmodified version of Minecraft. Using metrics for skill segmentation, reuse, and hierarchy quality, we find that our method consistently produces more structured and semantically meaningful hierarchies than existing baselines. Furthermore, as a proof of concept, we demonstrate that these discovered hierarchies accelerate and stabilise learning on downstream reinforcement learning tasks.

强化学习层次技能无监督学习技能发现

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。