解决强化学习中峰值成本失控的安全问题,提升真实场景下的鲁棒性。
Robust Peak-cost Constrained Reinforcement Learning

- 采用代理优化与积分概率度量,处理动态不确定性
- 在扰动环境下仍能保持安全约束,误差不超过ε
- 适合高风险场景如自动驾驶、医疗机器人
我们研究鲁棒峰值成本约束强化学习(RP-CRL),目标是在最大化期望奖励的同时控制轨迹中遇到的最大成本。该设定源于安全关键应用,单次严重违规可能造成灾难性后果,无法用基于期望累积成本的标准约束马尔可夫决策过程(CMDP)充分建模。现有可达性约束强化学习方法采用基于拉格朗日的方案,但峰值成本约束MDP的对偶性质尚不明确。我们证明,与标准CMDP不同,峰值成本约束MDP可能不存在零对偶间隙。进一步考虑了应对模拟器到现实世界动态差异的鲁棒形式。为此,我们提出一种代理优化框架和基于积分概率度量的鲁棒值估计方法。证明在适当超参数选择下,代理解能达到原问题的鲁棒奖励值,且约束违反不超过ε。实验表明,该方法在动态扰动下有效保障安全,同时保持强奖励性能。
原文摘要 · Abstract (English)
We study robust peak-cost constrained reinforcement learning (RP-CRL), where the objective is to maximize expected reward while controlling the maximum cost encountered along a trajectory. This setting is motivated by safety-critical applications in which a single large violation can be catastrophic and therefore cannot be adequately captured by the standard CMDP framework based on expected cumulative cost. Existing reachability-constrained RL methods adopt Lagrangian-based approaches, yet the underlying duality properties of peak-cost constrained MDPs remain unclear. We show that, unlike standard CMDPs, peak-cost constrained MDPs may not admit zero duality gap. We further consider a robust formulation to address simulator-to-real-world mismatch in the transition dynamics. To solve this problem, we develop a surrogate optimization framework and a robust value estimation method based on integral probability metrics. We prove that, with appropriate hyperparameter choices, the surrogate solution attains the same robust reward value as the original problem while violating the constraint by at most epsilon. Experiments show that the proposed method effectively enforces safety under dynamics perturbations while retaining strong reward performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。