用奖励函数组合生成行为空间,提升策略表达力。
Hierarchical Behaviour Spaces
- 通过线性组合奖励函数构建行为空间,扩展策略表达能力
- 在NetHack环境表现优异,优于传统单奖励方法
- 突破常规认知:层次结构优势来自探索增强而非长程推理
近年来的分层强化学习在预定义选项奖励函数下已成功扩展至数十亿时间步。我们发现,与其为每个选项使用单一奖励函数,不如利用奖励函数共同构建行为空间——让控制器对奖励函数进行线性组合,从而更有效地表示多样化策略。我们称该方法为分层行为空间(Hierarchical Behaviour Spaces, HBS)。在NetHack学习环境中评估HBS,结果表现强劲。通过一系列实验我们发现,与传统观点相反,该方法中层次结构的优势主要来自探索能力的提升,而非长期规划能力。
原文摘要 · Abstract (English)
Recent work in hierarchical reinforcement learning has shown success in scaling to billions of timesteps when learning over a set of predefined option reward functions. We show that, instead of using a single reward function per option, the reward functions can be effectively used to induce a space of behaviours, by letting the controller specify linear combinations over reward functions, allowing a more expressive set of policies to be represented. We call this method Hierarchical Behaviour Spaces (HBS). We evaluate HBS on the NetHack Learning Environment, demonstrating strong performance. We conduct a series of experiments and determine that, perhaps going against conventional wisdom, the benefits of hierarchy in our method come from increased exploration rather than long term reasoning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。