arXiv:2605.24740cs.LGcs.GT2026-05中稿 · ICML

通过迭代估计强化学习中的可达性最优策略,实现渐近最优。

Reinforcement Learning for Reachability: Guaranteeing Asymptotic Optimality

论文配图:Reinforcement Learning for Reachability: Guaranteeing Asymptotic Optimality
图 1 · 摘自论文原文
  • 基于PAC学习框架,逐步优化策略收敛条件
  • 理论证明在极限下可达到精确最优策略
  • 适合关注强化学习收敛机制的研究者

强化学习在满足可达性约束的序贯决策中至关重要,但理论保证仍不充分。近期工作实现了策略的渐近收敛,但对收敛动态缺乏深入理解。本文提出一种新方法,基于PAC学习框架,在有限时间内以高置信度获得近似最优策略,前提需已知马尔可夫决策过程(MDP)内部参数(如最小转移概率)。我们指出,尽管这些参数在强化学习中未知,但可通过迭代更新逐步精确估计。通过持续满足PAC条件,证明了在极限情况下可实现精确最优。标准基准测试的实证结果验证了该理论对收敛动态的洞察。

原文摘要 · Abstract (English)

Reinforcement learning (RL) for reachability specifications is fundamental in sequential decision-making, yet theoretical guarantees remain less explored. A recent work achieves asymptotic convergence to optimal policies. However, this approach provides limited insight into convergence dynamics. In this work, we present an alternative approach that provides deeper theoretical insights into convergence. Our approach builds on PAC learning with assumptions. PAC learning guarantees near-optimal policies with high confidence in finite time but requires knowing internal MDP parameters like minimum transition probability. We argue that while these parameters are unknown in RL, they can be iteratively refined and estimated with increasing accuracy. By iteratively satisfying PAC conditions, we show that exact optimality can be achieved in the limit. Empirical evaluations on standard benchmarks validate our theoretical insights into convergence dynamics.

强化学习可达性理论保证

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。