arXiv:2602.11018cs.LGcs.AI2026-02中稿 · AAMAS 2026

从不安全轨迹中学习安全策略,无需事先标注安全成本。

OSIL: Learning Offline Safe Imitation Policies with Safety Inferred from Non-preferred Trajectories

  • 通过非偏好轨迹推断安全约束,构建安全策略学习框架。
  • 在不降低奖励性能前提下,显著提升策略安全性与约束满足率。
  • 适合需离线学习且难以定义安全成本的现实场景应用。

本文针对离线安全模仿学习问题,目标是从缺乏每步安全成本或奖励信息的示范轨迹中学习安全且高回报的策略。在许多真实场景中,在环境中在线学习存在风险,而准确设定安全成本又十分困难。然而,收集反映不良或不安全行为的轨迹往往可行,这些轨迹隐含了应避免的行为模式,我们称之为非偏好轨迹。本文提出一种名为OSIL的新算法,通过非偏好轨迹推断安全信息。将安全策略学习建模为约束马尔可夫决策过程(CMDP),不依赖显式安全成本与奖励标注,而是通过推导奖励最大化目标的下界,并学习一个估计非偏好行为概率的成本模型。该方法使智能体仅凭离线示范即可学习到既安全又高回报的策略。实验表明,所提方法能在不降低奖励表现的前提下,有效满足成本约束,优于多个基线方法。

原文摘要 · Abstract (English)

This work addresses the problem of offline safe imitation learning (IL), where the goal is to learn safe and reward-maximizing policies from demonstrations that do not have per-timestep safety cost or reward information. In many real-world domains, online learning in the environment can be risky, and specifying accurate safety costs can be difficult. However, it is often feasible to collect trajectories that reflect undesirable or unsafe behavior, implicitly conveying what the agent should avoid. We refer to these as non-preferred trajectories. We propose a novel offline safe IL algorithm, OSIL, that infers safety from non-preferred demonstrations. We formulate safe policy learning as a Constrained Markov Decision Process (CMDP). Instead of relying on explicit safety cost and reward annotations, OSIL reformulates the CMDP problem by deriving a lower bound on reward maximizing objective and learning a cost model that estimates the likelihood of non-preferred behavior. Our approach allows agents to learn safe and reward-maximizing behavior entirely from offline demonstrations. We empirically demonstrate that our approach can learn safer policies that satisfy cost constraints without degrading the reward performance, thus outperforming several baselines.

安全强化学习离线学习模仿学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。