用奖励函数当测试工具,新方法揭示强化学习与系统稳定性的深层联系。
A Test-Function Approach to Incremental Stability
- 以奖励函数为测试工具,重构增量输入-状态稳定性分析框架。
- 证明了在霍尔德连续奖励下,值函数的光滑性与系统稳定性等价。
- 为强化学习与控制理论融合提供新视角,适合研究稳定性的学者。
本文提出一种分析增量输入-状态稳定性(δISS)的新框架,核心思想是将奖励函数作为“测试函数”。传统控制理论依赖满足时间递减条件的李雅普诺夫函数,而强化学习中的值函数则通过指数衰减构造,其奖励函数可能非光滑且无界。因此,这类值函数无法直接视为李雅普诺夫证书。本文建立了闭环系统在给定策略下的δISS变体与对抗性选择霍尔德连续奖励函数下值函数正则性之间的新等价关系。该结果表明,值函数的正则性及其与增量稳定性的关联,可脱离传统李雅普诺夫方法独立理解,为控制与强化学习的交叉研究提供了新路径。
原文摘要 · Abstract (English)
This paper presents a novel framework for analyzing Incremental-Input-to-State Stability ($δ$ISS) based on the idea of using rewards as "test functions." Whereas control theory traditionally deals with Lyapunov functions that satisfy a time-decrease condition, reinforcement learning (RL) value functions are constructed by exponentially decaying a Lipschitz reward function that may be non-smooth and unbounded on both sides. Thus, these RL-style value functions cannot be directly understood as Lyapunov certificates. We develop a new equivalence between a variant of incremental input-to-state stability of a closed-loop system under given a policy, and the regularity of RL-style value functions under adversarial selection of a Hölder-continuous reward function. This result highlights that the regularity of value functions, and their connection to incremental stability, can be understood in a way that is distinct from the traditional Lyapunov-based approach to certifying stability in control theory.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。