提出几何化非平稳MDP的稳定分析框架,提升动态环境下的策略跟踪能力。
Geometry of Drifting MDPs with Path-Integral Stability Certificates
- 将环境建模为可微同伦路径,追踪最优贝尔曼不动点的运动轨迹
- 证明路径积分稳定性界,给出远离突变区的稳定可行区域
- 设计轻量级自适应算法,根据漂移强度在线调节学习或规划强度
现实中的强化学习常面临非平稳性:奖励与动态随时间漂移、加速、振荡甚至发生最优动作的突变。现有理论多用粗粒度模型衡量环境变化的幅度,却忽略了局部变化方式——而加速度和近似平局正是导致跟踪误差和策略抖动的关键。本文从几何视角出发,将非平稳折扣马尔可夫决策过程(MDP)建模为可微同伦路径,追踪由此引发的最优贝尔曼不动点运动,揭示出内在复杂性的长度-曲率-拐点特征:累积漂移、加速度/振荡以及动作间隙引起的非光滑性。我们证明了与求解器无关的路径积分稳定性界,并推导出间隙安全的可行区域,可认证远离突变区的局部稳定性。基于此,提出轻量级封装方法:同伦追踪强化学习(HT-RL)与HT-MCTS,能在线估计路径长度、曲率及近似平局接近度的重放缓冲代理,并据此自适应调整学习或规划强度。实验表明,相比静态基线,该方法在动态环境中的跟踪性能与动态遗憾均有提升,尤其在振荡和突变频繁场景中收益显著。
原文摘要 · Abstract (English)
Real-world reinforcement learning is often \emph{nonstationary}: rewards and dynamics drift, accelerate, oscillate, and trigger abrupt switches in the optimal action. Existing theory often represents nonstationarity with coarse-scale models that measure \emph{how much} the environment changes, not \emph{how} it changes locally -- even though acceleration and near-ties drive tracking error and policy chattering. We take a geometric view of nonstationary discounted Markov Decision Processes (MDPs) by modeling the environment as a differentiable homotopy path and tracking the induced motion of the optimal Bellman fixed point. This yields a length-curvature-kink signature of intrinsic complexity: cumulative drift, acceleration/oscillation, and action-gap-induced nonsmoothness. We prove a solver-agnostic path-integral stability bound and derive gap-safe feasible regions that certify local stability away from switch regimes. Building on these results, we introduce \textit{Homotopy-Tracking RL (HT-RL)} and \textit{HT-MCTS}, lightweight wrappers that estimate replay-based proxies of length, curvature, and near-tie proximity online and adapt learning or planning intensity accordingly. Experiments show improved tracking and dynamic regret over matched static baselines, with the largest gains in oscillatory and switch-prone regimes.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。