arXiv:2605.12561cs.LGcs.RO2026-05

让智能体学会何时行动,实现更安全的低通信强化学习。

Learning When to Act: Communication-Efficient Reinforcement Learning via Run-Time Assurance

论文配图:Learning When to Act: Communication-Efficient Reinforcement Learning via Run-Time Assurance
图 1 · 摘自论文原文
  • 用时序保障层动态决定行动时机,结合李雅普诺夫屏障与LQR备份。
  • 在三类系统上使平均采样间隔提升1.45至3.51倍,且保持稳定。
  • 适应不同环境无需重训,适合高维、强鲁棒性需求的控制场景。

安全强化学习通常关注智能体应执行什么动作,本文则关注其应在何时行动。通过基于点态李雅普诺夫安全屏障的运行时保障(RTA)机制,单一策略可联合学习控制输入与通信高效的行动时机。研究聚焦于已知平衡点附近的稳定化问题,此时基于CARE的LQR备份、李雅普诺夫证书及经典李雅普诺夫-STC均有明确定义,便于与解析基准对比。RTA层通过一步前预测李雅普诺夫值和预计算的LQR备份覆盖策略,提供比仅期望安全的约束MDP方法更强的保证。在倒立摆、小车-摆杆和平面四旋翼系统上,所学策略的平均采样间隔(MSI)分别达到李雅普诺夫触发基线的1.91倍、1.45倍和3.51倍;相同平均速率下固定LQR控制器在三者上均不稳定,表明稀疏性安全源于自适应时机而非更低频率。基于CARE的李雅普诺夫奖励可跨环境迁移,仅一个权重参数$w_c$调节稳定性与通信权衡;消融实验表明RTA至关重要,移除后MSI下降1.27至1.84倍,状态范数恶化。偏好条件扩展仅需原训练算力的$\tfrac{2}{11}$即可恢复完整权衡前沿;SAC实验显示结果在离散与连续域中均具算法无关性。12状态3D四旋翼案例将框架扩展至高维系统,经典STC在此失效,且对±30%质量变化和扰动表现出良好鲁棒性,系统性能随干扰平滑退化,由RTA吸收无法处理的部分。

原文摘要 · Abstract (English)

Safe reinforcement learning (RL) typically asks $\textit{what}$ an agent should do. We ask $\textit{when}$ it needs to act, and show that a single policy can jointly learn control inputs and communication-efficient timing decisions under a pointwise Lyapunov safety shield. We focus on stabilization around a known equilibrium, where CARE-based LQR backups, Lyapunov certificates, and classical Lyapunov-STC are well defined, enabling clean comparison against analytical baselines. A run-time assurance (RTA) layer overrides the policy via a one-step-ahead Lyapunov prediction and a precomputed LQR backup, providing a strictly stronger guarantee than constrained MDP methods that enforce safety only in expectation. On an inverted pendulum, cart--pole, and planar quadrotor, the learned policy achieves $1.91\times$, $1.45\times$, and $3.51\times$ higher mean inter-sample interval (MSI) than a Lyapunov-triggered baseline; a fixed LQR controller at the same average rate is unstable on all three plants, showing that adaptive timing, not a lower average rate, makes sparsity safe. A CARE-derived Lyapunov reward transfers across environments without redesign, with a single weight $w_c$ controlling the stability--communication tradeoff; ablations confirm the RTA shield is essential, with its removal reducing MSI by $1.27$--$1.84\times$ and degrading state norms. A preference-conditioned extension recovers the full tradeoff frontier from one model at $\tfrac{2}{11}$ of training compute, and SAC experiments show the results are algorithm-agnostic across discrete and continuous domains. A 12-state 3D quadrotor case study extends the framework to higher-dimensional systems where classical STC is intractable, and robustness to $\pm30\%$ mass variation and disturbances shows graceful degradation, with the RTA absorbing what the learned policy cannot.

强化学习安全控制通信效率李雅普诺夫

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。