arXiv:2506.14045cs.AI2025-06综述被引 27

探索智能体在复杂环境中的时间结构,提升决策与学习效率。

Discovering Temporal Structure: An Overview of Hierarchical Reinforcement Learning

  • 通过分层强化学习发现经验流中的时间层级结构
  • 不同方法可从在线数据、离线数据或大语言模型中挖掘结构
  • 适用于需要长期规划的复杂任务,如机器人控制与游戏

在复杂开放环境中实现探索、规划与学习是人工智能的一大挑战。分层强化学习(HRL)通过发现并利用经验流中的时间结构,为这一挑战提供了有前景的解决方案。尽管该框架吸引力强,相关研究丰富,但尚不明确何为有效的时间结构,以及在哪些问题中识别结构具有价值。本文从决策根本挑战出发,揭示了HRL的优势,并分析其对智能体性能权衡的影响。随后系统梳理了发现时间结构的方法家族,涵盖从在线经验、离线数据到大语言模型(LLMs)的应用。最后指出当前在时间结构发现上的挑战,以及特别适合此类研究的领域。

原文摘要 · Abstract (English)

Developing agents capable of exploring, planning and learning in complex open-ended environments is a grand challenge in artificial intelligence (AI). Hierarchical reinforcement learning (HRL) offers a promising solution to this challenge by discovering and exploiting the temporal structure within a stream of experience. The strong appeal of the HRL framework has led to a rich and diverse body of literature attempting to discover a useful structure. However, it is still not clear how one might define what constitutes good structure in the first place, or the kind of problems in which identifying it may be helpful. This work aims to identify the benefits of HRL from the perspective of the fundamental challenges in decision-making, as well as highlight its impact on the performance trade-offs of AI agents. Through these benefits, we then cover the families of methods that discover temporal structure in HRL, ranging from learning directly from online experience to offline datasets, to leveraging large language models (LLMs). Finally, we highlight the challenges of temporal structure discovery and the domains that are particularly well-suited for such endeavours.

强化学习分层结构时间建模智能体

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。