证明深度Q网络可高概率逼近最优价值函数。
Universal Approximation Theorem of Deep Q-Networks
- 用随机控制与前向后向随机微分方程建模连续时间DQN
- 在紧集上以任意精度逼近最优Q函数,概率趋近1
- 揭示层数、离散化与粘性解对性能的影响,适合强化学习研究者
我们通过随机控制和前向后向随机微分方程(FBSDEs)建立了一个分析深度Q网络(DQN)的连续时间框架。针对由平方可积鞅驱动的连续时间马尔可夫决策过程(MDP),我们分析了DQN的逼近性质。结果表明,在紧集上,DQN可高概率地以任意精度逼近最优Q函数,这依赖于残差网络逼近定理及状态-动作过程的大偏差界。随后,我们分析了该设定下一般Q-learning算法的收敛性,采用随机逼近理论进行推导。研究强调了DQN层数、时间离散化与粘性解(主要针对值函数$V^*$)在处理最优Q函数可能非光滑性中的作用。本工作连接了深度强化学习与随机控制,为物理系统或高频数据应用中的连续时间DQN提供了理论支持。
原文摘要 · Abstract (English)
We establish a continuous-time framework for analyzing Deep Q-Networks (DQNs) via stochastic control and Forward-Backward Stochastic Differential Equations (FBSDEs). Considering a continuous-time Markov Decision Process (MDP) driven by a square-integrable martingale, we analyze DQN approximation properties. We show that DQNs can approximate the optimal Q-function on compact sets with arbitrary accuracy and high probability, leveraging residual network approximation theorems and large deviation bounds for the state-action process. We then analyze the convergence of a general Q-learning algorithm for training DQNs in this setting, adapting stochastic approximation theorems. Our analysis emphasizes the interplay between DQN layer count, time discretization, and the role of viscosity solutions (primarily for the value function $V^*$) in addressing potential non-smoothness of the optimal Q-function. This work bridges deep reinforcement learning and stochastic control, offering insights into DQNs in continuous-time settings, relevant for applications with physical systems or high-frequency data.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。