arXiv:2410.12175cs.LGcs.AI2024-10NeurIPS被引 16

将LTL与ω正则目标转化为平均奖励问题,实现可解释强化学习的渐近最优解。

Reinforcement Learning with LTL and $ω$-Regular Objectives via Optimality-Preserving Translation to Average Rewards

  • 通过有限记忆奖励机,将ω正则目标转为保优的平均奖励问题。
  • 证明可通过渐近求解一系列折扣和问题,获得最优策略。
  • 为LTL目标的强化学习提供理论保障,适合可解释性研究者。

线性时序逻辑(LTL)和更一般的ω正则目标是强化学习中传统折扣和与平均奖励目标的替代方案,具有更强的可读性和可解释性优势。本文研究了这些目标之间的关系。主要成果是:每个ω正则目标的强化学习问题均可通过有限记忆奖励机,以保优方式转化为极限平均奖励问题。此外,我们证明了通过渐近求解一系列近似折扣和问题,可找到极限平均问题的最优策略。因此,解决了长期未决问题:对于LTL和ω正则目标,最优策略可渐近学习得到。

原文摘要 · Abstract (English)

Linear temporal logic (LTL) and, more generally, $ω$-regular objectives are alternatives to the traditional discount sum and average reward objectives in reinforcement learning (RL), offering the advantage of greater comprehensibility and hence explainability. In this work, we study the relationship between these objectives. Our main result is that each RL problem for $ω$-regular objectives can be reduced to a limit-average reward problem in an optimality-preserving fashion, via (finite-memory) reward machines. Furthermore, we demonstrate the efficacy of this approach by showing that optimal policies for limit-average problems can be found asymptotically by solving a sequence of discount-sum problems approximately. Consequently, we resolve an open problem: optimal policies for LTL and $ω$-regular objectives can be learned asymptotically.

强化学习LTL可解释性平均奖励

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。