arXiv:2510.10544cs.LGcs.AI2025-10中稿 · ICML被引 1

提出新泛化界,让强化学习策略更可靠且可解释。

PAC-Bayesian Reinforcement Learning Trains Generalizable Policies

  • 基于马尔可夫链混合时间构建新泛化界,解决序列数据依赖问题。
  • 在连续控制任务中,算法提供有效置信度证书,性能不逊于SAC。
  • 适合关注模型可靠性与探索效率的研究者或工业应用。

我们推导出一种新的强化学习的PAC-Bayesian泛化界,通过马尔可夫链的混合时间显式建模数据中的依赖关系。该方法克服了传统泛化界在强化学习中因数据序列性而失效的问题。新界对现代离策略算法(如软演员-评论家,Soft Actor-Critic)提供了非平凡的泛化保证。我们通过PB-SAC算法验证其实际效用:在训练过程中优化该界以引导探索。在多个连续控制任务上的实验表明,该方法既能提供有意义的置信度证书,又能保持与SAC相当的性能。

原文摘要 · Abstract (English)

We derive a novel PAC-Bayesian generalization bound for reinforcement learning that explicitly accounts for Markov dependencies in the data, through the chain's mixing time. This contributes to overcoming challenges in obtaining generalization guarantees for reinforcement learning, where the sequential nature of data breaks the independence assumptions underlying classical bounds. The new bound provides non-vacuous certificates for modern off-policy algorithms such as Soft Actor-Critic. We demonstrate the practical utility of the bound through PB-SAC, a novel algorithm that optimizes the bound during training to guide exploration. Experiments across several continuous control tasks show that the proposed approach provides meaningful confidence certificates while maintaining competitive performance.

强化学习泛化界PAC-Bayes策略训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。