arXiv:2505.21841cs.LGcs.AI2025-05ICML被引 4

首个应对任意对抗性约束的在线安全强化学习算法。

An Optimistic Algorithm for online CMDPS with Anytime Adversarial Constraints

  • 基于乐观镜面下降的对偶优化框架,动态适应未知变化的约束。
  • 实现最优后悔度 O(sqrt(K)) 与强约束违反度 O(sqrt(K))。
  • 无需严格安全策略或Slater条件,适合高风险动态场景应用。

在线安全强化学习在自动驾驶、机器人和网络安全等动态环境中至关重要,目标是学习在满足安全约束条件下最大化奖励的策略,此类问题可建模为受限马尔可夫决策过程(CMDPs)。现有方法在随机约束下可实现次线性后悔,但在对抗性设置中表现不佳——约束未知、时变且可能被恶意设计。本文提出首个针对任意对抗性约束的在线CMDPs求解算法:乐观镜面下降对偶算法(OMDPD),在不依赖Slater条件或已知严格安全策略的前提下,实现了最优后悔度O(sqrt(K))和强约束违反度O(sqrt(K))。进一步表明,若能获取准确的奖励与转移函数估计,还可进一步优化这些界。结果为对抗性环境中安全决策提供了实用保障。

原文摘要 · Abstract (English)

Online safe reinforcement learning (RL) plays a key role in dynamic environments, with applications in autonomous driving, robotics, and cybersecurity. The objective is to learn optimal policies that maximize rewards while satisfying safety constraints modeled by constrained Markov decision processes (CMDPs). Existing methods achieve sublinear regret under stochastic constraints but often fail in adversarial settings, where constraints are unknown, time-varying, and potentially adversarially designed. In this paper, we propose the Optimistic Mirror Descent Primal-Dual (OMDPD) algorithm, the first to address online CMDPs with anytime adversarial constraints. OMDPD achieves optimal regret O(sqrt(K)) and strong constraint violation O(sqrt(K)) without relying on Slater's condition or the existence of a strictly known safe policy. We further show that access to accurate estimates of rewards and transitions can further improve these bounds. Our results offer practical guarantees for safe decision-making in adversarial environments.

强化学习在线学习安全决策对抗约束

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。