arXiv:2511.07486cs.LGcs.SY2025-11被引 1

提出首个鲁棒约束MDP的采样复杂度保证,解决安全强化学习中的不确定性挑战。

Provably Efficient Sample Complexity for Robust CMDP

  • 构建带剩余安全预算的状态空间,确保策略在最坏动态下仍满足约束
  • 新算法RCVI在生成模型下仅需约|S||A|H⁵/ε²次采样,误差不超过ε
  • 适用于高安全要求场景,如自动驾驶、医疗决策等关键系统

我们研究在真实环境与模拟器或基准模型存在差异时,如何学习最大化累积奖励且满足安全约束的策略。聚焦于鲁棒约束马尔可夫决策过程(RCMDP),要求在不确定性集内的最坏动态下,累计效用超过阈值的同时最大化奖励。尽管已有工作建立了基于策略优化的有限时间迭代复杂度,但采样复杂度尚未被充分探讨。本文首次证明,在矩形不确定性集下,马尔可夫策略可能并非最优,不同于无约束鲁棒MDP的情况。为此,我们引入一个扩展状态空间,将剩余效用预算纳入状态表示。在此基础上,提出新型鲁棒约束值迭代算法(RCVI),在生成模型下实现采样复杂度$ ilde{O}(|S||A|H^5/ε^2)$,保证最大违反度不超过$ε$,其中$|S|$和$|A|$分别为状态和动作空间大小,$H$为每回合长度。据我们所知,这是首个针对RCMDP的采样复杂度保证。实验结果进一步验证了方法的有效性。

原文摘要 · Abstract (English)

We study the problem of learning policies that maximize cumulative reward while satisfying safety constraints, even when the real environment differs from a simulator or nominal model. We focus on robust constrained Markov decision processes (RCMDPs), where the agent must maximize reward while ensuring cumulative utility exceeds a threshold under the worst-case dynamics within an uncertainty set. While recent works have established finite-time iteration complexity guarantees for RCMDPs using policy optimization, their sample complexity guarantees remain largely unexplored. In this paper, we first show that Markovian policies may fail to be optimal even under rectangular uncertainty sets unlike the {\em unconstrained} robust MDP. To address this, we introduce an augmented state space that incorporates the remaining utility budget into the state representation. Building on this formulation, we propose a novel Robust constrained Value iteration (RCVI) algorithm with a sample complexity of $\mathcal{\tilde{O}}(|S||A|H^5/ε^2)$ achieving at most $ε$ violation using a generative model where $|S|$ and $|A|$ denote the sizes of the state and action spaces, respectively, and $H$ is the episode length. To the best of our knowledge, this is the {\em first sample complexity guarantee} for RCMDP. Empirical results further validate the effectiveness of our approach.

强化学习安全控制鲁棒优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。