用因果监控生成实时奖励,让强化学习更稳定地控制复杂系统
Control Synthesis of Cyber-Physical Systems for Real-Time Specifications through Causation-Guided Reinforcement Learning
- 通过在线监控系统行为与时序逻辑规范的因果关系,生成即时奖励
- 在多个连续控制基准上,训练稳定性与收敛速度优于现有方法
- 适合需要高可靠性的实时控制系统设计者使用
在实时且安全关键的网络物理系统(CPS)中,控制合成需在不确定和动态环境下保证策略满足严格的时序与正确性要求。信号时序逻辑(STL)作为表达实时约束的强大形式化工具,其语义可对系统行为进行量化评估。同时,强化学习(RL)已成为解决未知环境中控制合成问题的重要方法。近期研究将基于STL的奖励函数引入RL,以自动合成控制策略。然而,这些方法获得的奖励反映的是整段或部分路径的全局评估,无法准确累积局部变化的奖励,导致稀疏的全局奖励可能引发训练不收敛和性能不稳定。本文提出一种基于在线因果监控的实时奖励生成方法。该方法在每个控制步骤持续监控系统行为与STL规范的符合程度,计算向满足或违反目标的定量距离,从而生成反映瞬时状态动态的奖励。此外,我们提供了因果语义的平滑近似,克服其不连续性,使其对深度强化学习方法可微。我们实现了一个原型工具,并在Gym环境的多种连续控制基准上进行了评估。实验结果表明,所提出的基于因果语义的STL引导强化学习方法优于现有相关方法,为深度强化学习提供了更鲁棒、高效的奖励生成框架。
原文摘要 · Abstract (English)
In real-time and safety-critical cyber-physical systems (CPSs), control synthesis must guarantee that generated policies meet stringent timing and correctness requirements under uncertain and dynamic conditions. Signal temporal logic (STL) has emerged as a powerful formalism of expressing real-time constraints, with its semantics enabling quantitative assessment of system behavior. Meanwhile, reinforcement learning (RL) has become an important method for solving control synthesis problems in unknown environments. Recent studies incorporate STL-based reward functions into RL to automatically synthesize control policies. However, the automatically inferred rewards obtained by these methods represent the global assessment of a whole or partial path but do not accumulate the rewards of local changes accurately, so the sparse global rewards may lead to non-convergence and unstable training performances. In this paper, we propose an online reward generation method guided by the online causation monitoring of STL. Our approach continuously monitors system behavior against an STL specification at each control step, computing the quantitative distance toward satisfaction or violation and thereby producing rewards that reflect instantaneous state dynamics. Additionally, we provide a smooth approximation of the causation semantics to overcome the discontinuity of the causation semantics and make it differentiable for using deep-RL methods. We have implemented a prototype tool and evaluated it in the Gym environment on a variety of continuously controlled benchmarks. Experimental results show that our proposed STL-guided RL method with online causation semantics outperforms existing relevant STL-guided RL methods, providing a more robust and efficient reward generation framework for deep-RL.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。