arXiv:2603.26982stat.MLcs.AI2026-03被引 1

为样本平均Q-learning提供在线统计推断,可生成置信区间。

Online Statistical Inference of Constant Sample-averaged Q-Learning

  • 基于泛函中心极限定理构建在线推断框架。
  • 在网格世界和资源匹配任务中实现95%覆盖率的置信区间。
  • 适合需要可靠性评估的强化学习应用开发者。

强化学习算法广泛应用于各类决策任务,但其性能常受高方差和不稳定性影响,尤其在噪声环境或稀疏奖励场景下。本文提出一种针对样本平均Q-learning方法的在线统计推断框架。在一般条件下,我们适配了泛函中心极限定理(FCLT)于改进后的算法,并通过随机缩放构造Q值的置信区间。实验在两个问题上对比了改进方法与传统Q-learning:一个为简单玩具例题的网格世界,另一个为真实世界的动态资源匹配问题。结果报告了两种方法的覆盖率与置信区间宽度,验证了该框架的有效性。

原文摘要 · Abstract (English)

Reinforcement learning algorithms have been widely used for decision-making tasks in various domains. However, the performance of these algorithms can be impacted by high variance and instability, particularly in environments with noise or sparse rewards. In this paper, we propose a framework to perform statistical online inference for a sample-averaged Q-learning approach. We adapt the functional central limit theorem (FCLT) for the modified algorithm under some general conditions and then construct confidence intervals for the Q-values via random scaling. We conduct experiments to perform inference on both the modified approach and its traditional counterpart, Q-learning using random scaling and report their coverage rates and confidence interval widths on two problems: a grid world problem as a simple toy example and a dynamic resource-matching problem as a real-world example for comparison between the two solution approaches.

强化学习统计推断Q-learning

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。