arXiv:2607.13938cs.RO2026-07

从专家观察中学习安全控制策略,无需预先设计安全约束。

Discriminative Barrier Functions for Safe Adversarial Imitation Learning from Observation

论文配图:Discriminative Barrier Functions for Safe Adversarial Imitation Learning from Observation
图 1 · 摘自论文原文
  • 用控制屏障函数约束奖励函数空间,实现安全探索
  • 仅靠未标注的专家数据就能恢复出鲁棒的屏障函数
  • 适合需要安全保证的机器人导航与真实场景部署

逆强化学习(IRL)虽能从专家示范中学习并泛化,但常依赖无约束探索,难以保障现实应用安全。控制屏障函数(CBFs)可确保控制系统安全,但其解析设计耗时且复杂。本文提出将奖励函数候选限制在CBF空间内,使IRL兼具安全在线控制与持续经验优化能力。关键在于,该框架可直接从无标签的专家观测中数据驱动地恢复屏障函数。实验表明,恢复的屏障函数对专家数据中完全未出现的危险状态也具备鲁棒性。在仿真导航环境中,本方法优于标准IRL基线,显著提升安全性。进一步对比了基于规划与基于策略的IRL方法在仿真和真实障碍避让任务中的权衡。

原文摘要 · Abstract (English)

Inverse Reinforcement Learning (IRL) algorithms are powerful tools for learning from and generalizing expert demonstrations, but they often rely on unconstrained exploration, rendering them unsafe for real-world deployment. Meanwhile, Control Barrier Functions (CBFs) can guarantee the safety of control systems, but the analytical design of CBFs can be time-consuming and esoteric. In this work, we address these limitations jointly by constraining reward function candidacy during IRL to the space of CBFs, yielding a formulation that exhibits safe online control with continuous experiential improvement. Crucially, this framework enables the data-driven recovery of barrier functions directly from unlabeled expert observations. We demonstrate that the recovered barrier function is robust to unsafe states entirely absent from the expert data. Furthermore, we benchmark our method against standard IRL baselines in a simulated navigation environment, demonstrating improved safety performance. Finally, we investigate the trade-offs of planning-based versus policy-based IRL methods across both simulation and a real world obstacle avoidance task.

安全强化学习逆强化学习屏障函数

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。