arXiv:2602.23816cs.LGcs.AI2026-02中稿 · publication at AAM…

在未知约束下学习安全策略,通过专家示范提升高回报轨迹概率。

Learning to maintain safety through expert demonstrations in settings with unknown constraints: A Q-learning perspective

  • 用Q值融合奖励与安全评估,定义状态动作对的潜在价值。
  • 在多个基准任务中优于现有逆向约束强化学习算法。
  • 适合需要平衡安全与效率的高风险决策场景。

在可观测奖励但未知约束且不可观测代价的受限马尔可夫决策过程(constrained MDP)中,给定一组安全执行任务的轨迹,目标是找到一种策略,最大化示范轨迹的出现概率,同时在保守性与高回报轨迹(可能含不安全步骤)的概率之间取得平衡。为此,我们基于演示数据,以“承诺”为指标优化最有望成功的轨迹。通过将每个状态-动作对的“承诺”建模为依赖于任务奖励和安全评估的Q值,提出了一种安全Q值逆向约束强化学习(SafeQIL)算法。该方法结合了奖励与安全性的期望,形成一种新的安全Q学习视角。在一系列具有挑战性的基准任务上,SafeQIL与当前最先进的逆向约束强化学习算法进行了对比,验证了其有效性。

原文摘要 · Abstract (English)

Given a set of trajectories demonstrating the execution of a task safely in a constrained MDP with observable rewards but with unknown constraints and non-observable costs, we aim to find a policy that maximizes the likelihood of demonstrated trajectories trading the balance between being conservative and increasing significantly the likelihood of high-rewarding trajectories but with potentially unsafe steps. Having these objectives, we aim towards learning a policy that maximizes the probability of the most $promising$ trajectories with respect to the demonstrations. In so doing, we formulate the ``promise" of individual state-action pairs in terms of $Q$ values, which depend on task-specific rewards as well as on the assessment of states' safety, mixing expectations in terms of rewards and safety. This entails a safe Q-learning perspective of the inverse learning problem under constraints: The devised Safe $Q$ Inverse Constrained Reinforcement Learning (SafeQIL) algorithm is compared to state-of-the art inverse constraint reinforcement learning algorithms to a set of challenging benchmark tasks, showing its merits.

强化学习安全策略逆向学习Q学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。