让视觉控制更安全平滑,提升任务成功率一倍
How to Train Your Latent Control Barrier Function: Smooth Safety Filtering Under Hard-to-Model Constraints
- 用梯度惩罚让潜在空间的约束函数更平滑
- 混合正常与安全策略数据训练,提升预测精度
- 在仿真和真实机械臂上实现连续安全过滤
潜在国内安全滤波器将哈密顿-雅可比可达性扩展至从高维观测中直接学习的潜在状态表示,实现了在难以建模的约束下安全的视觉运动控制。然而,现有方法采用‘最宽松’的切换策略,在正常与安全策略间离散切换,可能损害现代视觉运动策略的任务性能。尽管可达性值函数理论上可转化为控制屏障函数(CBF)以实现基于优化的平滑过滤,我们从理论和实证上表明,当前潜在空间学习方法产生的值函数本质上不兼容。主要存在两方面不兼容:其一,在哈密顿-雅可比可达性中,失败通过潜在空间中的“裕度函数”编码,其符号决定潜在状态是否属于约束集;但将裕度函数表示为分类器会导致饱和值函数并出现不连续跳变。我们证明,值函数的利普希茨常数与裕度函数的利普希茨常数呈线性关系,揭示了平滑的CBF要求平滑的裕度。其二,仅在安全策略数据上训练的强化学习近似,对正常策略动作的值估计不准确,而这正是CBF过滤所需的关键区域。为此,我们提出LatentCBF,通过梯度惩罚使裕度函数平滑且无需额外标注,同时采用混合正常与安全策略数据的值函数训练流程。在模拟基准和硬件平台上使用基于视觉的抓取策略的实验表明,LatentCBF实现了平滑安全过滤,并使任务完成率相比之前切换方法提升一倍。
原文摘要 · Abstract (English)
Latent safety filters extend Hamilton-Jacobi (HJ) reachability to operate on latent state representations and dynamics learned directly from high-dimensional observations, enabling safe visuomotor control under hard-to-model constraints. However, existing methods implement "least-restrictive" filtering that discretely switch between nominal and safety policies, potentially undermining the task performance that makes modern visuomotor policies valuable. While reachability value functions can, in principle, be adapted to be control barrier functions (CBFs) for smooth optimization-based filtering, we theoretically and empirically show that current latent-space learning methods produce fundamentally incompatible value functions. We identify two sources of incompatibility: First, in HJ reachability, failures are encoded via a "margin function" in latent space, whose sign indicates whether or not a latent is in the constraint set. However, representing the margin function as a classifier yields saturated value functions that exhibit discontinuous jumps. We prove that the value function's Lipschitz constant scales linearly with the margin function's Lipschitz constant, revealing that smooth CBFs require smooth margins. Second, reinforcement learning (RL) approximations trained solely on safety policy data yield inaccurate value estimates for nominal policy actions, precisely where CBF filtering needs them. We propose the LatentCBF, which addresses both challenges through gradient penalties that lead to smooth margin functions without additional labeling, and a value-training procedure that mixes data from both nominal and safety policy distributions. Experiments on simulated benchmarks and hardware with a vision-based manipulation policy demonstrate that LatentCBF enables smooth safety filtering while doubling the task-completion rate over prior switching methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。