提升自动驾驶交通灯识别鲁棒性,应对光照雨雾等自然干扰。
Sequence-Preserving Dual-FoV Defense for Traffic Sign and Light Recognition in Autonomous Vehicles
- 设计双视角时序保持防御框架,融合特征压缩与熵检测。
- 在真实场景下实现79.8mAP、攻击成功率降至18.2%。
- 适合关注自动驾驶感知安全的工程师与研究者。
交通灯与交通标志识别对自动驾驶车辆至关重要,感知错误直接影响导航与安全。现有模型不仅易受数字对抗攻击,还面临眩光、雨天、污渍或涂鸦等自然退化影响,导致危险误判。当前研究缺乏对时序连续性、多静态视场(FoV)感知及数字与自然退化的联合鲁棒性考量。本研究基于aiMotive、Udacity、Waymo及德克萨斯州自录视频构建多源数据集,针对高速公路、夜间、雨天和城市四种运行设计域(ODD),对中长序列RGB图像进行时序对齐。通过一系列真实应用中的异常检测实验,提出统一三层防御栈:特征挤压、防御性蒸馏与基于熵的异常检测,并引入序列级时序投票进一步增强性能。评估指标包括准确率、攻击成功率(ASR)、风险加权误分类严重度及置信度稳定性。物理可迁移性经探针复现验证。结果表明,统一防御栈达到79.8mAP,ASR降低至18.2%,优于YOLOv8、YOLOv9与BEVFormer,高风险误分类减少至32%。
原文摘要 · Abstract (English)
Traffic light and sign recognition are key for Autonomous Vehicles (AVs) because perception mistakes directly influence navigation and safety. In addition to digital adversarial attacks, models are vulnerable to existing perturbations (glare, rain, dirt, or graffiti), which could lead to dangerous misclassifications. The current work lacks consideration of temporal continuity, multistatic field-of-view (FoV) sensing, and robustness to both digital and natural degradation. This study proposes a dual FoV, sequence-preserving robustness framework for traffic lights and signs in the USA based on a multi-source dataset built on aiMotive, Udacity, Waymo, and self-recorded videos from the region of Texas. Mid and long-term sequences of RGB images are temporally aligned for four operational design domains (ODDs): highway, night, rainy, and urban. Over a series of experiments on a real-life application of anomaly detection, this study outlines a unified three-layer defense stack framework that incorporates feature squeezing, defensive distillation, and entropy-based anomaly detection, as well as sequence-wise temporal voting for further enhancement. The evaluation measures included accuracy, attack success rate (ASR), risk-weighted misclassification severity, and confidence stability. Physical transferability was confirmed using probes for recapture. The results showed that the Unified Defense Stack achieved 79.8mAP and reduced the ASR to 18.2%, which is superior to YOLOv8, YOLOv9, and BEVFormer, while reducing the high-risk misclassification to 32%.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。