用强化学习优化领域时序规划的启发式策略,提升搜索效率。
Exploiting Symbolic Heuristics for the Synthesis of Domain-Specific Temporal Planning Guidance using Reinforcement Learning
- 结合符号启发式设计奖励机制,缓解无限状态问题
- 学习符号启发式的残差修正,而非从零训练
- 多队列规划融合学习与符号启发式,适合复杂时序任务
近期工作探索了在固定领域和给定训练问题集(非计划)条件下,利用强化学习合成启发式引导以提升时序规划器性能。本文提出该框架的演进:首先,形式化多种奖励机制,利用符号启发式缓解因截断无限状态马尔可夫决策过程带来的问题;其次,提出学习现有符号启发式的残差修正值,而非从零开始学习整个启发式;最后,采用多队列规划方法,将学习得到的启发式与符号启发式结合,平衡系统性搜索与不完美学习信息。实验对比了各方法优劣,显著推进了该规划与学习范式的技术水平。
原文摘要 · Abstract (English)
Recent work investigated the use of Reinforcement Learning (RL) for the synthesis of heuristic guidance to improve the performance of temporal planners when a domain is fixed and a set of training problems (not plans) is given. The idea is to extract a heuristic from the value function of a particular (possibly infinite-state) MDP constructed over the training problems. In this paper, we propose an evolution of this learning and planning framework that focuses on exploiting the information provided by symbolic heuristics during both the RL and planning phases. First, we formalize different reward schemata for the synthesis and use symbolic heuristics to mitigate the problems caused by the truncation of episodes needed to deal with the potentially infinite MDP. Second, we propose learning a residual of an existing symbolic heuristic, which is a "correction" of the heuristic value, instead of eagerly learning the whole heuristic from scratch. Finally, we use the learned heuristic in combination with a symbolic heuristic using a multiple-queue planning approach to balance systematic search with imperfect learned information. We experimentally compare all the approaches, highlighting their strengths and weaknesses and significantly advancing the state of the art for this planning and learning schema.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。