证明非均匀更新会导致强化学习算法无法收敛。
Scalar-Stepsize Nonuniform Monte Carlo Optimistic Policy Iteration: A Certified Counterexample
- 采用固定非均匀采样概率的标量步长异步更新
- 在三状态两动作环境中出现周期性吸引轨道,无法收敛
- 揭示非均匀采样会破坏残差收缩机制,适合研究收敛性边界
Tsitsiklis 证明了在均匀更新结构下蒙特卡洛乐观策略迭代的收敛性,并指出非均匀更新频率是一个关键障碍。本文针对具有固定非均匀状态选择概率的标量步长、未归一化的异步状态值递归,给出了一个经认证的反例。在一个三状态、两动作的折扣MDP中,非均匀更新频率导致对角缩放的贪婪策略均值场出现已认证的非恒定吸引混合周期轨道。即使使用有界无偏几何时域估计器和Robbins–Monro步长,原始随机递归仍以正概率陷入该周期轨道,因而不收敛。该例子揭示了一个几何障碍:均匀采样产生径向残差收缩,而标量非均匀采样则各向异性地扭曲残差动态,可能生成切换吸引周期。
原文摘要 · Abstract (English)
Tsitsiklis proved convergence of Monte Carlo optimistic policy iteration under a uniform update structure and identified nonuniform update frequencies as a delicate obstruction. We give a certified negative answer for the natural scalar-stepsize, unnormalized asynchronous state-value recursion with fixed nonuniform state-selection probabilities. In a three-state, two-action discounted MDP, the nonuniform update frequencies induce a diagonally scaled greedy-policy mean field with a certified nonconstant attracting hybrid periodic orbit. With a bounded unbiased geometric-horizon estimator and Robbins--Monro stepsizes, the original stochastic recursion remains trapped near the cycle with positive probability and therefore fails to converge. The example pinpoints a geometric obstruction: uniform sampling gives radial residual contraction, whereas scalar nonuniform sampling anisotropically distorts the residual dynamics and can generate switched attracting cycles.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。