arXiv:2505.00382cs.LGmath.PR2025-05

用随机延迟微分方程解析DQN的稳定机制,揭示经验回放与目标网络的连续系统本质。

Approximation to Deep Q-Network by Stochastic Delay Differential Equations

  • 构建基于DQN的随机延迟微分方程,从连续视角建模算法动态。
  • 证明两者的Wasserstein-1距离随步长减小趋近于零,实现理论逼近。
  • 揭示目标网络的延迟项提升系统稳定性,适合研究强化学习理论者阅读。

尽管深度Q网络(DQN)在强化学习中取得了显著突破,但其理论分析仍有限。本文基于DQN算法构造了一个随机延迟微分方程(SDDE),并估计两者之间的Wasserstein-1距离。我们给出了该距离的上界,并证明当步长趋于零时,两者距离收敛至零。该结果使我们能从连续系统角度理解DQN的两大关键技术:经验回放与目标网络。具体而言,方程中的延迟项对应目标网络,有助于系统稳定性。我们的方法结合了改进的Lindeberg原则与算子比较,建立了上述结论。

原文摘要 · Abstract (English)

Despite the significant breakthroughs that the Deep Q-Network (DQN) has brought to reinforcement learning, its theoretical analysis remains limited. In this paper, we construct a stochastic differential delay equation (SDDE) based on the DQN algorithm and estimate the Wasserstein-1 distance between them. We provide an upper bound for the distance and prove that the distance between the two converges to zero as the step size approaches zero. This result allows us to understand DQN's two key techniques, the experience replay and the target network, from the perspective of continuous systems. Specifically, the delay term in the equation, corresponding to the target network, contributes to the stability of the system. Our approach leverages a refined Lindeberg principle and an operator comparison to establish these results.

强化学习DQN随机微分方程理论分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。