用随机微分方程理论证明深度Q网络可逼近贝尔曼算子迭代过程。
Universal Approximation Theorem for Deep Q-Learning via FBSDE System
- 基于后向随机微分方程,构建与贝尔曼更新结构一致的神经网络架构。
- 证明深度残差网络层能以可控误差逼近贝尔曼算子作用,且深度对应迭代次数。
- 揭示网络在值函数空间中的动态系统行为,适用于强化学习理论研究者。
深度Q网络(DQN)的近似能力通常依赖于不利用最优Q函数内在结构的通用近似定理(UAT)。本文建立了针对一类特定结构DQN的UAT,其网络设计模拟贝尔曼更新的迭代精化过程。分析核心为正则性传播:单次贝尔曼算子作用具有的正则性,可通过后向随机微分方程(BSDE)理论处理;而整个值迭代序列的统一正则性——在标准利普希茨假设下,紧凑域上的统一利普希茨连续性——由有限时域动态规划原理导出。我们证明,深度残差网络的层可视为作用于函数空间的神经算子,逼近贝尔曼算子的作用。由此得到的近似定理与控制问题结构紧密关联,提供一种网络深度直接对应值函数精化迭代次数、误差可控传播的证明方法。这一视角揭示了网络在值函数空间中的动态系统行为。
原文摘要 · Abstract (English)
The approximation capabilities of Deep Q-Networks (DQNs) are commonly justified by general Universal Approximation Theorems (UATs) that do not leverage the intrinsic structural properties of the optimal Q-function, the solution to a Bellman equation. This paper establishes a UAT for a class of DQNs whose architecture is designed to emulate the iterative refinement process inherent in Bellman updates. A central element of our analysis is the propagation of regularity: while the transformation induced by a single Bellman operator application exhibits regularity, for which Backward Stochastic Differential Equations (BSDEs) theory provides analytical tools, the uniform regularity of the entire sequence of value iteration iterates--specifically, their uniform Lipschitz continuity on compact domains under standard Lipschitz assumptions on the problem data--is derived from finite-horizon dynamic programming principles. We demonstrate that layers of a deep residual network, conceived as neural operators acting on function spaces, can approximate the action of the Bellman operator. The resulting approximation theorem is thus intrinsically linked to the control problem's structure, offering a proof technique wherein network depth directly corresponds to iterations of value function refinement, accompanied by controlled error propagation. This perspective reveals a dynamic systems view of the network's operation on a space of value functions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。