arXiv:2412.08794cs.LGstat.ML2024-12ICLR被引 9

用隐变量建模安全约束,让离线强化学习更安全高效

Latent Safety-Constrained Policy Approach for Safe Offline Reinforcement Learning

  • 用条件变分自编码器学习隐式安全约束
  • 在安全约束下优化累积奖励,显著提升性能
  • 适合自动驾驶等高风险场景的离线安全决策

在安全离线强化学习中,目标是仅利用离线数据训练出在严格遵守安全约束前提下最大化累积奖励的策略。传统方法常难以平衡安全性与性能,导致表现下降或安全风险上升。本文提出新方法:首先通过条件变分自编码器(Conditional Variational Autoencoders)学习一个保守的安全策略,以建模潜在的安全约束;随后将问题转化为受约束的奖励-回报最大化问题,在隐空间中通过奖励优势加权回归训练编码器,使策略在满足推断出的隐式安全约束的同时优化奖励。理论分析提供了策略性能与样本复杂度的边界。在多个基准数据集(包括挑战性的自动驾驶场景)上的大量实验表明,该方法不仅确保安全合规,且在累积奖励优化上优于现有方法。可视化结果进一步揭示了方法的有效性及内在机制。

原文摘要 · Abstract (English)

In safe offline reinforcement learning (RL), the objective is to develop a policy that maximizes cumulative rewards while strictly adhering to safety constraints, utilizing only offline data. Traditional methods often face difficulties in balancing these constraints, leading to either diminished performance or increased safety risks. We address these issues with a novel approach that begins by learning a conservatively safe policy through the use of Conditional Variational Autoencoders, which model the latent safety constraints. Subsequently, we frame this as a Constrained Reward-Return Maximization problem, wherein the policy aims to optimize rewards while complying with the inferred latent safety constraints. This is achieved by training an encoder with a reward-Advantage Weighted Regression objective within the latent constraint space. Our methodology is supported by theoretical analysis, including bounds on policy performance and sample complexity. Extensive empirical evaluation on benchmark datasets, including challenging autonomous driving scenarios, demonstrates that our approach not only maintains safety compliance but also excels in cumulative reward optimization, surpassing existing methods. Additional visualizations provide further insights into the effectiveness and underlying mechanisms of our approach.

离线RL安全强化学习隐变量建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。