研究神经网络宽度缩放对强化学习算法不确定性的影响。
Scaling Effects and Uncertainty Quantification in Neural Actor Critic Algorithms
- 提出通用反多项式缩放机制,调节网络宽度与训练步数关系。
- 发现方差随宽度提升以幂律衰减,参数越接近1越稳定。
- 给出学习率与探索率的可证明优良配置方法,适合高可靠性场景。
本文研究使用浅层神经网络作为演员和评论家模型的神经演员-评论家算法。研究聚焦两点:其一,在网络宽度和训练步数趋于无穷时,不同缩放方案下网络输出的收敛性质;其二,精确控制各缩放模式下的近似误差。已有研究表明,在网络宽度的反平方根缩放下,系统收敛至具有随机初值的常微分方程。本文将关注点从收敛速度拓展至算法输出的全面统计表征,旨在量化神经演员-评论家方法中的不确定性。具体地,我们研究一种广义的反多项式缩放,指数为介于0.5与1之间的可调超参数。通过推导网络输出的渐近展开(视为统计估计量),揭示其结构。在主导项中,我们证明方差随网络宽度以幂律衰减,衰减速率等于0.5减去缩放参数,表明当参数趋近1时统计鲁棒性显著提升。数值实验支持该行为,并进一步提示此缩放选择可带来更快收敛。最终分析为学习率、探索率等算法超参数提供了明确配置建议,这些建议依赖于网络宽度与缩放参数,确保可证明的优良统计行为。
原文摘要 · Abstract (English)
We investigate the neural Actor Critic algorithm using shallow neural networks for both the Actor and Critic models. The focus of this work is twofold: first, to compare the convergence properties of the network outputs under various scaling schemes as the network width and the number of training steps tend to infinity; and second, to provide precise control of the approximation error associated with each scaling regime. Previous work has shown convergence to ordinary differential equations with random initial conditions under inverse square root scaling in the network width. In this work, we shift the focus from convergence speed alone to a more comprehensive statistical characterization of the algorithm's output, with the goal of quantifying uncertainty in neural Actor Critic methods. Specifically, we study a general inverse polynomial scaling in the network width, with an exponent treated as a tunable hyperparameter taking values strictly between one half and one. We derive an asymptotic expansion of the network outputs, interpreted as statistical estimators, in order to clarify their structure. To leading order, we show that the variance decays as a power of the network width, with an exponent equal to one half minus the scaling parameter, implying improved statistical robustness as the scaling parameter approaches one. Numerical experiments support this behavior and further suggest faster convergence for this choice of scaling. Finally, our analysis yields concrete guidelines for selecting algorithmic hyperparameters, including learning rates and exploration rates, as functions of the network width and the scaling parameter, ensuring provably favorable statistical behavior.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。