arXiv:2509.08660cs.LG2025-09被引 5

提出可复现的线性函数逼近强化学习方法,提升算法稳定性。

Replicable Reinforcement Learning with Linear Function Approximation

  • 设计高效可复现的随机设计回归与中心化协方差估计算法
  • 首次实现线性马尔可夫决策过程在生成模型与回合制设置下的可复现性保证
  • 实验验证可生成更一致的神经策略,适合追求稳定性的研究者

实验结果复现是机器学习等科学领域长期面临的挑战。近期研究将可复现性形式化为:算法在从同一分布中抽取的两个不同样本上执行时,应产生完全相同的结果。可证明可复现的算法对强化学习(RL)尤其重要,因现有算法在实践中常表现不稳定。尽管表格型强化学习已存在可复现算法,但将其扩展至更实用的函数逼近场景仍是开放问题。本文通过开发适用于线性函数逼近的可复现方法取得进展。首先提出两种高效的可复现随机设计回归与非中心协方差估计算法,具有独立价值;随后利用这些工具,首次构建出线性马尔可夫决策过程在生成模型与回合制设置下的可证明高效的可复现强化学习算法。最后通过实验评估,展示其能启发更一致的神经策略。

原文摘要 · Abstract (English)

Replication of experimental results has been a challenge faced by many scientific disciplines, including the field of machine learning. Recent work on the theory of machine learning has formalized replicability as the demand that an algorithm produce identical outcomes when executed twice on different samples from the same distribution. Provably replicable algorithms are especially interesting for reinforcement learning (RL), where algorithms are known to be unstable in practice. While replicable algorithms exist for tabular RL settings, extending these guarantees to more practical function approximation settings has remained an open problem. In this work, we make progress by developing replicable methods for linear function approximation in RL. We first introduce two efficient algorithms for replicable random design regression and uncentered covariance estimation, each of independent interest. We then leverage these tools to provide the first provably efficient replicable RL algorithms for linear Markov decision processes in both the generative model and episodic settings. Finally, we evaluate our algorithms experimentally and show how they can inspire more consistent neural policies.

强化学习可复现性线性函数逼近算法稳定性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。