针对鸟瞰图目标检测模型,提出黑盒鲁棒性评估框架,可生成致命语义扰动。
A Black-Box Evaluation Framework for Semantic Robustness in Bird's Eye View Detection
- 设计平滑距离代理函数与确定性优化算法,精准生成三类语义扰动
- 实测10个主流模型,PolarFormer最鲁棒,BEVDet精度降为零
- 首次系统评估鸟瞰图检测模型在对抗扰动下的脆弱性,适合自动驾驶安全研究者
基于摄像头的鸟瞰图(BEV)感知模型在自动驾驶中日益重要,但深度学习模型的鲁棒性问题备受关注。现有工作仅考察随机语义扰动对多视角BEV检测的影响,本文首次提出一种黑盒鲁棒性评估框架,通过对抗优化三种常见语义扰动——几何变换、色彩偏移和运动模糊——来欺骗BEV模型。为解决语义扰动优化难题,设计了基于距离的平滑代理函数替代mAP,并引入SimpleDIRECT算法,利用观测梯度方向指导优化。实验对比随机扰动与两种优化基线,验证框架有效性。同时构建了十个近期BEV模型的语义鲁棒性基准。结果表明,强调多视图几何信息的PolarFormer鲁棒性最强,而BEVDet则完全失效,精度降至零。
原文摘要 · Abstract (English)
Camera-based Bird's Eye View (BEV) perception models receive increasing attention for their crucial role in autonomous driving, a domain where concerns about the robustness and reliability of deep learning have been raised. While only a few works have investigated the effects of randomly generated semantic perturbations, aka natural corruptions, on the multi-view BEV detection task, we develop a black-box robustness evaluation framework that adversarially optimises three common semantic perturbations: geometric transformation, colour shifting, and motion blur, to deceive BEV models, serving as the first approach in this emerging field. To address the challenge posed by optimising the semantic perturbation, we design a smoothed, distance-based surrogate function to replace the mAP metric and introduce SimpleDIRECT, a deterministic optimisation algorithm that utilises observed slopes to guide the optimisation process. By comparing with randomised perturbation and two optimisation baselines, we demonstrate the effectiveness of the proposed framework. Additionally, we provide a benchmark on the semantic robustness of ten recent BEV models. The results reveal that PolarFormer, which emphasises geometric information from multi-view images, exhibits the highest robustness, whereas BEVDet is fully compromised, with its precision reduced to zero.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。