arXiv:2602.17312cs.LGcs.SY2026-02

用分层安全优先级实现离线强化学习中的零安全违规

LexiSafe: Offline Safe Reinforcement Learning with Lexicographic Safety-Reward Hierarchy

  • 通过分层优先级结构强制安全第一,避免安全漂移
  • 在多个基准上安全违规减少,任务性能更优
  • 适合对安全要求严苛的物理系统决策场景

离线安全强化学习在信息物理系统(CPS)中日益重要,训练期间不可容忍安全违规,且仅能使用预采集数据。现有方法通常通过约束松弛或联合优化平衡奖励与安全,但缺乏防止安全漂移的结构性机制。本文提出LexiSafe,一种用于保持安全对齐行为的分层离线强化学习框架。首先构建LexiSafe-SC,一种针对标准离线安全强化学习的单成本形式化,推导出安全违规和性能次优性的边界,共同提供样本复杂度保证。随后扩展至多层级安全需求的LexiSafe-MC,支持多个安全成本,并具备自身的样本复杂度分析。实验表明,相较于约束式离线基线,LexiSafe显著降低安全违规并提升任务性能。通过将分层优先级与结构偏差结合,LexiSafe为关键安全系统决策提供了实用且理论完备的方法。

原文摘要 · Abstract (English)

Offline safe reinforcement learning (RL) is increasingly important for cyber-physical systems (CPS), where safety violations during training are unacceptable and only pre-collected data are available. Existing offline safe RL methods typically balance reward-safety tradeoffs through constraint relaxation or joint optimization, but they often lack structural mechanisms to prevent safety drift. We propose LexiSafe, a lexicographic offline RL framework designed to preserve safety-aligned behavior. We first develop LexiSafe-SC, a single-cost formulation for standard offline safe RL, and derive safety-violation and performance-suboptimality bounds that together yield sample-complexity guarantees. We then extend the framework to hierarchical safety requirements with LexiSafe-MC, which supports multiple safety costs and admits its own sample-complexity analysis. Empirically, LexiSafe demonstrates reduced safety violations and improved task performance compared to constrained offline baselines. By unifying lexicographic prioritization with structural bias, LexiSafe offers a practical and theoretically grounded approach for safety-critical CPS decision-making.

离线RL安全强化学习分层优先级

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。