arXiv:2410.15184cs.LGcs.AI2024-10ICLR被引 3

让智能体自动发现高层次动作,提升长程探索效率

Action abstractions for amortized sampling

  • 通过聚类高奖励轨迹中的动作序列,生成高层抽象动作
  • 在复杂环境中显著提升样本效率,更好发现多样高回报状态
  • 抽象动作具可解释性,适合需要高效探索的强化学习任务

随着强化学习和生成流网络(GFlowNets)所采样轨迹长度增加,信用分配与探索变得愈发困难,长规划时域阻碍了模式发现与泛化。这一挑战在熵驱动的强化学习方法中尤为突出,如生成流网络,其代理需学会从结构化分布中采样,并发现多个需多步到达的高奖励状态。为此,我们提出将动作抽象(即高层动作)的发现融入策略优化过程。该方法迭代提取多个高奖励轨迹中常见的动作子序列,并将其‘打包’为一个新动作加入动作空间。在合成与真实环境的实证评估中,该方法在发现多样高奖励物体方面表现出更优的样本效率,尤其在更难的探索问题上。我们还观察到,抽象后的高阶动作具有可解释性,捕捉了动作空间奖励景观的潜在结构。本工作提供了受认知启发的动作抽象方法,首次实现了在可摊销序列采样中的层次化规划。

原文摘要 · Abstract (English)

As trajectories sampled by policies used by reinforcement learning (RL) and generative flow networks (GFlowNets) grow longer, credit assignment and exploration become more challenging, and the long planning horizon hinders mode discovery and generalization. The challenge is particularly pronounced in entropy-seeking RL methods, such as generative flow networks, where the agent must learn to sample from a structured distribution and discover multiple high-reward states, each of which take many steps to reach. To tackle this challenge, we propose an approach to incorporate the discovery of action abstractions, or high-level actions, into the policy optimization process. Our approach involves iteratively extracting action subsequences commonly used across many high-reward trajectories and `chunking' them into a single action that is added to the action space. In empirical evaluation on synthetic and real-world environments, our approach demonstrates improved sample efficiency performance in discovering diverse high-reward objects, especially on harder exploration problems. We also observe that the abstracted high-order actions are interpretable, capturing the latent structure of the reward landscape of the action space. This work provides a cognitively motivated approach to action abstraction in RL and is the first demonstration of hierarchical planning in amortized sequential sampling.

强化学习动作抽象探索效率层次规划

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。