用核Stein散度检测罕见高成本事件,实现更安全的强化学习。
SteinGate: Tail-Sensitive Safe Reinforcement Learning via Stein Discrepancy

- 基于核Stein散度构建非参数安全证书,动态验证策略表现
- 实验显示约束违规频率与严重程度显著降低,性能媲美顶尖方法
- 适合对安全性要求高的连续控制任务,如机器人操控
安全强化学习通常通过限制期望累积成本来保证安全,但这一标准难以发现罕见但灾难性的尾部事件。本文提出SteinGate,一种边界感知的分布式安全证书,以核化Stein散度的稳健一致性检验替代脆弱的尾部拟合,并考虑了截断成本带来的边界原子。SteinGate评估策略回放成本是否与安全参考分布保持一致,从而提供非参数化安全验证。该证书用于动态调整学习机制:当回放成本与安全分布一致时,优先进行收益提升的策略更新;一旦成本尾部偏离,则切换至恢复行为。在连续控制基准测试中,SteinGate显著降低了训练期间约束违反的频率和严重程度,同时保持与当前最优基线相当的回报水平。
原文摘要 · Abstract (English)
Safe reinforcement learning typically enforces safety by bounding expected cumulative costs, a criterion that often fails to detect rare but catastrophic tail events. To overcome these limitations, this paper introduces SteinGate, a boundary-aware distributional safety certificate that replaces fragile tail fitting with a robust consistency check using Kernelized Stein Discrepancy while accounting for boundary atoms induced by clipped costs. SteinGate evaluates whether observed policy rollout costs remain consistent with a safe reference distribution, providing a non-parametric safety certificate. This certificate is used to dynamically adapt the learning regime: favoring reward-improving policy updates when rollouts remain consistent with the safe reference and switching to recovery behavior when the cost tail deviates. Experiments on continuous-control benchmarks demonstrate that SteinGate significantly reduces both the frequency and severity of constraint violations during training while maintaining competitive returns relative to state-of-the-art baselines.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。