提出一种新型强化学习方法,可更准确估计不确定性并构建置信区间。
Asymptotic Analysis of Sample-averaged Q-learning
- 用样本平均法改进Q-learning,融合多组奖励与状态数据
- 证明算法渐近正态性,支持无需调参的置信区间构造
- 在经典环境验证不同批量策略对学习效率的影响
强化学习在复杂不确定环境中已成为训练智能体的关键方法。将统计推断引入强化学习算法对于理解与管理模型性能的不确定性至关重要。本文提出一种广义的时间变化批平均Q-learning框架,称为样本平均Q-learning(SA-QL),通过聚合奖励和下一状态的多个样本,扩展了传统单样本Q-learning,以更好地应对数据变异性与不确定性。我们利用函数中心极限定理(FCLT)建立新理论框架,在较弱条件下揭示样本平均算法的渐近正态性。此外,我们设计了一种随机缩放方法用于区间估计,实现置信区间构造而无需额外超参数。在风速网格世界和滑溜冰面等经典随机OpenAI Gym环境中进行大量数值实验,展示了不同批量调度策略对学习效率、覆盖率及置信区间宽度的影响。本工作为样本平均Q-learning建立了统一的理论基础,为强化学习算法的有效批量调度与统计推断提供新见解。
原文摘要 · Abstract (English)
Reinforcement learning (RL) has emerged as a key approach for training agents in complex and uncertain environments. Incorporating statistical inference in RL algorithms is essential for understanding and managing uncertainty in model performance. This paper introduces a generalized framework for time-varying batch-averaged Q-learning, termed sample-averaged Q-learning (SA-QL), which extends traditional single-sample Q-learning by aggregating samples of rewards and next states to better account for data variability and uncertainty. We leverage the functional central limit theorem (FCLT) to establish a novel framework that provides insights into the asymptotic normality of the sample-averaged algorithm under mild conditions. Additionally, we develop a random scaling method for interval estimation, enabling the construction of confidence intervals without requiring extra hyperparameters. Extensive numerical experiments across classic stochastic OpenAI Gym environments, including windy gridworld and slippery frozenlake, demonstrate how different batch scheduling strategies affect learning efficiency, coverage rates, and confidence interval widths. This work establishes a unified theoretical foundation for sample-averaged Q-learning, providing insights into effective batch scheduling and statistical inference for RL algorithms.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。