分析连续控制中Q学习的光滑性,揭示其对状态-动作空间的混合正则性。
Deep Q-Learning on Hölder Spaces
- 从贝尔曼目标出发,研究连续时间下的正则性传播机制
- 证明更新算子使状态变量光滑,动作变量仅保持Lipschitz依赖
- 提出适配混合正则性的张量型DeepONet架构,可调参数平衡精度与复杂度
我们研究连续时间随机控制中基于值函数强化学习的算子理论核心,针对连续状态与动作空间。在值函数型强化学习中,每个Q-learning或DQN更新均基于贝尔曼最优性目标;本文在扩散设定下隔离该目标,分析其正则性与近似复杂度。在一致椭圆性和Hölder-连续系数条件下,证明贝尔曼更新将有界输入映射至各向异性正则类:状态变量被平滑,动作变量仅保留Lipschitz依赖。这导致贝尔曼迭代族具有紧性,从而启发一种适配问题混合正则性的张量积DeepONet架构。我们进一步推导出显式的近似误差与资源消耗边界,并揭示当时间步δ→0时的刚度-复杂度权衡关系。该理论直接贡献于连续随机控制中贝尔曼目标正则性与近似性的理解。但未给出含探索、回放与随机梯度更新的实用采样Q-learning的完整收敛定理。
原文摘要 · Abstract (English)
We study the operator-theoretic core of Q-learning in continuous-time stochastic control with continuous states and actions. In value-based reinforcement learning, each Q-learning or DQN update is built from a Bellman optimality target; our analysis isolates this target in a diffusion setting and studies its regularity and approximation complexity. Under uniform ellipticity and Hölder-regular coefficients, we show that a Bellman update maps bounded inputs into an anisotropic regularity class, smoothing the state variable while leaving only Lipschitz dependence on the action variable. This yields a compact family of Bellman iterates and motivates a tensor-product DeepONet architecture adapted to the mixed regularity of the problem. We then derive explicit approximation and resource bounds, together with a stiffness--complexity trade-off as the time step $δ\to 0$. The resulting theory makes a direct contribution to Q-learning theory at the level of Bellman target regularity and approximation in continuous stochastic control. At the same time, we do not claim a full convergence theorem for practical sampled Q-learning with exploration, replay, and stochastic gradient updates.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。