arXiv:2505.03668cs.AI2025-05中稿 · 9th Conference on …被引 5

用逻辑规则自动学习长期有效的决策动作,提升不确定环境下的推理效率。

Learning Symbolic Persistent Macro-Actions for POMDP Solving Over Time

  • 通过事件演算的线性时序逻辑生成持续性宏观动作
  • 在Pocman和Rocksample任务中推理速度显著提升,性能稳定
  • 仅需少量执行轨迹即可学习,无需人工设计启发式规则

本文提出将时间逻辑推理与部分可观测马尔可夫决策过程(POMDP)结合,实现不确定性环境下可解释的决策。方法基于事件演算的线性时序逻辑(LTL)生成持久性(即恒定)的宏观动作,指导基于蒙特卡洛树搜索(MCTS)的POMDP求解器,在整个时间范围内大幅减少推理时间并保持稳健性能。这些宏观动作通过归纳逻辑编程(ILP)从少量执行轨迹(信念-动作对)中学习,无需人工设计启发式规则,仅需指定POMDP转移模型。在Pocman和Rocksample基准场景中,所学宏动作展现出更强的表达力与泛化能力,相比时间无关启发式,带来显著的计算效率提升。

原文摘要 · Abstract (English)

This paper proposes an integration of temporal logical reasoning and Partially Observable Markov Decision Processes (POMDPs) to achieve interpretable decision-making under uncertainty with macro-actions. Our method leverages a fragment of Linear Temporal Logic (LTL) based on Event Calculus (EC) to generate \emph{persistent} (i.e., constant) macro-actions, which guide Monte Carlo Tree Search (MCTS)-based POMDP solvers over a time horizon, significantly reducing inference time while ensuring robust performance. Such macro-actions are learnt via Inductive Logic Programming (ILP) from a few traces of execution (belief-action pairs), thus eliminating the need for manually designed heuristics and requiring only the specification of the POMDP transition model. In the Pocman and Rocksample benchmark scenarios, our learned macro-actions demonstrate increased expressiveness and generality when compared to time-independent heuristics, indeed offering substantial computational efficiency improvements.

POMDP逻辑推理宏动作强化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。