arXiv:2605.08417cs.LGmath.OC2026-05

提出一种新型强化学习算法,解决分布鲁棒性学习的计算难题。

Central Limit Theorem for Two-Time-Scale Approximate Distributionally Robust RL

论文配图:Central Limit Theorem for Two-Time-Scale Approximate Distributionally Robust RL
图 1 · 摘自论文原文
  • 基于一阶近似构建无对抗优化的近似贝尔曼方程
  • 算法在小歧义范围内实现 $n^{-1/2}$ 收敛率与中心极限定理
  • 适合追求理论保证的强化学习研究者

设计无模型的分布鲁棒强化学习(DRRL)算法面临根本挑战:鲁棒贝尔曼算子对转移核呈非线性,导致单样本更新有偏;而鲁棒性背后的对抗优化使评估计算昂贵。针对此,本文考虑在KL散度歧义集下的小歧义情形,提出基于相关鲁棒泛函一阶展开的近似DRRL框架。该框架得到一个移除对抗优化但仍保持歧义半径一阶精度的近似鲁棒贝尔曼方程。为求解该方程的不动点,提出均值-方差随机逼近(MVSA)算法,仅需单样本更新,通过升维随机逼近动态与双时标设计实现。进一步证明了MVSA的收敛性及中心极限定理:主迭代在标准 $n^{-1/2}$ 标度下满足中心极限定理,且渐近协方差可显式刻画。最后通过数值实验验证了理论结果。

原文摘要 · Abstract (English)

Designing model-free algorithms for distributionally robust reinforcement learning (DRRL) poses fundamental challenges. The robust Bellman operator is nonlinear in the transition kernel, which makes one-sample Bellman updates biased, while the adversarial optimization underlying robustness makes robust evaluation computationally demanding. To address these difficulties, we consider the natural small-ambiguity regime under Kullback--Leibler ambiguity sets and propose an approximate DRRL framework based on a first-order expansion of the relevant robust functional. This yields an approximate robust Bellman equation that removes the adversarial optimization while remaining first-order accurate in the ambiguity radius. To learn the fixed point of this approximate equation, we propose Mean-Variance Stochastic Approximation (MVSA), a model-free algorithm that uses only one-sample updates. This is achieved via a lifted stochastic approximation dynamics and a two-time-scale design. We then prove convergence and a central limit theorem for MVSA: its main iterate satisfies a central limit theorem at the canonical $n^{-1/2}$ scale, with explicitly characterized asymptotic covariances. Finally, we validate our theoretical findings with a numerical experiment.

强化学习分布鲁棒随机逼近中心极限定理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。