让强化学习模型像人一样自我反思,提升规划效率。
Efficient Planning in Reinforcement Learning via Model Introspection
- 将模型自省视为程序分析,挖掘内部模型信息
- 在关系型强化学习模型上实现高效目标规划
- 连接强化学习与经典规划,适合算法研究者
强化学习与经典规划通常被视为两种不同问题,因形式差异需采用不同解决方案。然而,当人类面对任务时,无论任务如何描述,往往能自主获取所需额外信息以高效求解。其关键在于自省:通过推理自身对问题的内部模型,直接生成相关任务信息。本文提出,这种自省可类比为程序分析。我们探讨了该方法在强化学习中各类模型上的应用实例,并提出一种算法,可在关系型强化学习模型上实现高效的面向目标规划,揭示了强化学习与经典规划间的新联系。
原文摘要 · Abstract (English)
Reinforcement learning and classical planning are typically seen as two distinct problems, with differing formulations necessitating different solutions. Yet, when humans are given a task, regardless of the way it is specified, they can often derive the additional information needed to solve the problem efficiently. The key to this ability is introspection: by reasoning about their internal models of the problem, humans directly synthesize additional task-relevant information. In this paper, we propose that this introspection can be thought of as program analysis. We discuss examples of how this approach can be applied to various kinds of models used in reinforcement learning. We then describe an algorithm that enables efficient goal-oriented planning over the class of models used in relational reinforcement learning, demonstrating a novel link between reinforcement learning and classical planning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。