arXiv:2410.22585cs.CVcs.LG2024-10被引 2

用预训练视觉模型做自动驾驶安全过滤器,效果接近已知真实状态时的表现。

Pre-Trained Vision Models as Perception Backbones for Safety Filters in Autonomous Driving

  • 用冻结的预训练视觉模型作为感知骨干设计安全过滤器。
  • 在DeepAccident数据集上,性能接近已知真实状态的过滤器。
  • 适合关注高维视觉控制安全性的研究者和工程师。

端到端视觉自动驾驶已取得显著进展,但安全性仍是主要挑战。在低维设置中,安全过滤器(如基于控制屏障函数的方法)已用于解决安全控制问题。在自动驾驶的高维视觉设置中设计此类过滤器同样可缓解安全问题,但更具挑战性。本文受预训练视觉模型在机器人控制策略中成功应用的启发,采用冻结的预训练视觉表征模型作为感知骨干,构建视觉安全过滤器。我们对四种常见预训练视觉模型在此场景下的离线性能进行了评估,并尝试了三种针对黑箱动态系统的安全过滤器训练方法,因表征空间的动力学未知。实验使用DeepAccident数据集,该数据集包含多个摄像头在CARLA中模拟真实事故场景的动作标注视频。结果表明,本方法所得过滤器性能与已知车辆及环境真实状态的过滤器相当。

原文摘要 · Abstract (English)

End-to-end vision-based autonomous driving has achieved impressive success, but safety remains a major concern. The safe control problem has been addressed in low-dimensional settings using safety filters, e.g., those based on control barrier functions. Designing safety filters for vision-based controllers in the high-dimensional settings of autonomous driving can similarly alleviate the safety problem, but is significantly more challenging. In this paper, we address this challenge by using frozen pre-trained vision representation models as perception backbones to design vision-based safety filters, inspired by these models' success as backbones of robotic control policies. We empirically evaluate the offline performance of four common pre-trained vision models in this context. We try three existing methods for training safety filters for black-box dynamics, as the dynamics over representation spaces are not known. We use the DeepAccident dataset that consists of action-annotated videos from multiple cameras on vehicles in CARLA simulating real accident scenarios. Our results show that the filters resulting from our approach are competitive with the ones that are given the ground truth state of the ego vehicle and its environment.

自动驾驶安全过滤视觉模型预训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。