arXiv:2411.02998cs.AIcs.LG2024-11ICLR被引 2

通过分层技能结构加速强化学习任务泛化,表现超越现有方法。

Accelerating Task Generalisation with Multi-Level Skill Hierarchies

  • 构建多层级技能层次,基于行为模式的未来价值生成选项
  • 在复杂程序生成环境中实现更高分布内与分布外性能
  • 适合研究高效泛化和层次强化学习的学者与工程师

实现强化学习智能体在新任务上的有效泛化是人工智能研究的关键挑战。本文提出分形聚类选项(Fracture Cluster Options, FraCOs),一种多层级分层强化学习方法,在困难的泛化任务上达到当前最优表现。FraCOs识别智能体行为中的模式,并根据这些模式的预期未来效用形成选项,从而实现对新任务的快速适应。在表格设置中,FraCOs展现出良好的迁移能力,且随着层次深度增加性能持续提升。我们在多个复杂的程序生成环境中,将FraCOs与最先进深度强化学习算法进行对比,结果表明其在分布内和分布外任务上均优于竞争对手。

原文摘要 · Abstract (English)

Creating reinforcement learning agents that generalise effectively to new tasks is a key challenge in AI research. This paper introduces Fracture Cluster Options (FraCOs), a multi-level hierarchical reinforcement learning method that achieves state-of-the-art performance on difficult generalisation tasks. FraCOs identifies patterns in agent behaviour and forms options based on the expected future usefulness of those patterns, enabling rapid adaptation to new tasks. In tabular settings, FraCOs demonstrates effective transfer and improves performance as it grows in hierarchical depth. We evaluate FraCOs against state-of-the-art deep reinforcement learning algorithms in several complex procedurally generated environments. Our results show that FraCOs achieves higher in-distribution and out-of-distribution performance than competitors.

强化学习任务泛化分层结构

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。