arXiv:2609.06467cs.LG2026-09

提出混合策略稳定性新框架,突破传统收敛限制。

Local and Global Stability in Performative Reinforcement Learning

  • 用策略混合重构稳定性定义,摆脱对环境敏感性的依赖
  • 局部稳定可实现$O(1/\\/sqrt{T})$收敛率,无需连续性假设
  • 全局稳定需弱于Lipschitz的过渡范围假设,适用于多人博弈

在执行性强化学习中,部署的策略会改变生成未来数据的环境,自然的解概念是能诱导出自身最优的执行性稳定策略。现有收敛保证依赖于环境映射$\pi \mapsto (P_\pi, r_\pi)$的Lipschitz敏感性假设,这类假设难以验证且在多智能体最优响应动态等场景中失效。本文改用策略混合来研究稳定性,发现其与执行性预测有本质不同:随机化可消除对敏感性假设的需求。我们区分了局部混合稳定性(基于占用加权的一阶松弛,等价于驻定性)与全局混合稳定性(抵御任意偏离策略)。第一个结果是:任意、可能不连续的环境映射下,加权状态级Hedge动态可使局部稳定性差距以$O(1/\sqrt{T})$速率趋于零,无论使用精确反馈或轨迹反馈。两种稳定性本质上不同:存在实例使局部稳定精确达成,但所有混合策略的全局稳定性差距始终远离零。针对全局稳定性,我们引入弱于Lipschitz敏感性的有界转移范围假设,使得无权重状态级Hedge收敛至$O(\gamma\epsilon_P/(1-\gamma)^3)$的下限,并在轨迹反馈下证明了关于$\epsilon_P$的匹配下界$\Omega(\gamma\epsilon_P/(1-\gamma))$,说明该下限不可避免。最后,我们将两种稳定性推广至$n$人执行性马尔可夫博弈,无需联合环境映射或博弈结构假设即可获得局部稳定性;对于执行性马尔可夫势博弈,可实现全局稳定性。

原文摘要 · Abstract (English)

In performative reinforcement learning the deployed policy shapes the environment that generates the learner's future data, and the natural solution concept is a performatively stable policy that is optimal in the environment it induces. Existing convergence guarantees rely on Lipschitz sensitivity assumptions on the environment map $\pi \mapsto (P_\pi, r_\pi)$, which are hard to verify and fail in settings such as multi-agent best-response dynamics. We instead study stability for mixtures of policies, and show that the resulting picture is fundamentally different from performative prediction, where randomization removes the need for any sensitivity assumption. We distinguish local mixed stability, an occupancy-weighted first-order relaxation that we show is equivalent to stationarity, from global mixed stability, which certifies against arbitrary deviating policies. Our first result is that a weighted per-state Hedge dynamic drives the local stability gap to zero at an $O(1/\sqrt{T})$ rate for an arbitrary, possibly discontinuous, environment map, both with exact and with trajectory feedback. The two notions genuinely differ: we exhibit an instance where local stability is achieved exactly but every mixture has global stability gap bounded away from zero. For global stability we introduce a bounded transition range assumption, strictly weaker than Lipschitz sensitivity, under which unweighted per-state Hedge converges up to a floor of $O(\gamma\epsilon_P/(1-\gamma)^3)$, and we prove a matching-in-$\epsilon_P$ lower bound of $\Omega(\gamma\epsilon_P/(1-\gamma))$ under trajectory feedback, so this floor is unavoidable. Finally, we extend both notions to $n$-player performative Markov games, obtaining local stability with no assumption on the joint environment map or game structure, and global stability for performative Markov potential games.

强化学习博弈论稳定性分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。