arXiv:2410.18416cs.LGcs.RO2024-10NeurIPS被引 15

通过引导状态因子间交互,发现更有效的无监督技能

SkiLD: Unsupervised Skill Discovery Guided by Factor Interactions

  • 基于状态因子分解,鼓励技能产生多样交互
  • 在复杂环境中学习出语义清晰的技能,优于仅追求状态覆盖的方法
  • 适合需要长期规划的机器人任务,如家庭环境操作

无监督技能发现有望让智能体通过自主、无奖励的环境交互学习可复用技能。现有方法通过鼓励行为可区分且覆盖多样状态来学习技能,但在包含多个状态因子(如含多个物体的家庭环境)的复杂环境中,覆盖所有状态不可能实现,盲目追求状态多样性常导致简单技能,不利于下游任务。本文提出基于局部依赖的技能发现(Skild),利用状态因子分解作为自然归纳偏置,引导技能学习。核心思想是:能引发状态因子间多样化交互的技能,往往对解决下游任务更有价值。为此,Skild设计了一种新目标函数,显式鼓励掌握能有效诱发环境内不同交互的技能。我们在多个具有挑战性的长时程稀疏奖励任务领域评估了Skild,包括一个真实的模拟家庭机器人场景,结果表明,Skild成功学习到语义明确的技能,并在性能上优于仅最大化状态覆盖的现有无监督强化学习方法。

原文摘要 · Abstract (English)

Unsupervised skill discovery carries the promise that an intelligent agent can learn reusable skills through autonomous, reward-free environment interaction. Existing unsupervised skill discovery methods learn skills by encouraging distinguishable behaviors that cover diverse states. However, in complex environments with many state factors (e.g., household environments with many objects), learning skills that cover all possible states is impossible, and naively encouraging state diversity often leads to simple skills that are not ideal for solving downstream tasks. This work introduces Skill Discovery from Local Dependencies (Skild), which leverages state factorization as a natural inductive bias to guide the skill learning process. The key intuition guiding Skild is that skills that induce <b>diverse interactions</b> between state factors are often more valuable for solving downstream tasks. To this end, Skild develops a novel skill learning objective that explicitly encourages the mastering of skills that effectively induce different interactions within an environment. We evaluate Skild in several domains with challenging, long-horizon sparse reward tasks including a realistic simulated household robot domain, where Skild successfully learns skills with clear semantic meaning and shows superior performance compared to existing unsupervised reinforcement learning methods that only maximize state coverage.

无监督学习技能发现机器人强化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。