提出新方法,在交互环境中安全规划并保持安全保证。
Safe Planning in Interactive Environments via Iterative Policy Updates and Adversarially Robust Conformal Prediction
- 通过敏感性分析量化策略更新对环境的影响,动态调整安全边界。
- 在二维汽车行人和四旋翼飞行器场景中验证了安全与收敛性。
- 适合自动驾驶、机器人控制等需实时安全决策的场景。
在交互式环境(如自动驾驶车辆与行人共存)中实现自主代理的安全规划面临重大挑战,因环境行为未知且会随代理策略变化而改变。这种耦合导致交互驱动的分布偏移,使现有方法的安全保障失效。尽管近期工作利用可 conformal prediction (CP) 在无分布假设下生成安全保证,但其数据可交换性假设在交互设置中被破坏,因策略更新与环境行为存在循环依赖。为此,本文提出一种迭代框架,通过量化策略更新对环境行为的潜在影响,持续维护安全保证。方法结合对抗鲁棒 CP:每轮使用当前策略下的观测数据进行常规 CP,再通过解析调整将安全保证跨策略更新传递,调整基于策略到轨迹的敏感性分析,实现安全的周期性开环规划。进一步通过系统收缩性分析,给出 CP 结果与策略更新收敛的条件。在二维汽车-行人及高维四旋翼飞行器案例中,实证验证了安全性和收敛性。据我们所知,这是首个在交互环境中提供有效安全保证的结果。
原文摘要 · Abstract (English)
Safe planning of an autonomous agent in interactive environments -- such as the control of a self-driving vehicle among pedestrians -- poses a major challenge as the behavior of the environment is unknown and reactive to the behavior of the autonomous agent. This coupling gives rise to interaction-driven distribution shifts where the autonomous agent's control policy may change the environment's behavior, thereby invalidating safety guarantees in existing work. Indeed, recent works have used conformal prediction (CP) to generate distribution-free safety guarantees using observed data of the environment. However, CP's assumption on data exchangeability is violated in interactive settings due to a circular dependency where a control policy update changes the environment's behavior, and vice versa. To address this gap, we propose an iterative framework that robustly maintains safety guarantees across policy updates by quantifying the potential impact of a planned policy update on the environment's behavior. We realize this via adversarially robust CP where we perform a regular CP step in each episode using observed data under the current policy, but then transfer safety guarantees across policy updates by analytically adjusting the CP result to account for distribution shifts. This adjustment is performed based on a policy-to-trajectory sensitivity analysis, resulting in a safe, episodic open-loop planner. We further conduct a contraction analysis of the system providing conditions under which both the CP results and the policy updates are guaranteed to converge. We empirically demonstrate these safety and convergence guarantees on a two-dimensional car-pedestrian and a high-dimensional quadcopter case study. To the best of our knowledge, these are the first results that provide valid safety guarantees in such interactive settings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。