用生成模型模拟罕见事故场景,发现自动驾驶感知漏洞。
AUTHENTICATION: Identifying Rare Failure Modes in Autonomous Vehicle Perception Systems using Adversarially Guided Diffusion Models
- 用扩散模型逆向生成环境掩码,结合文本提示构造对抗样本。
- 在真实数据集上成功诱导检测模型失效,暴露隐藏缺陷。
- 输出可解释的自然语言报告,适合安全评估与政策制定。
自动驾驶汽车依赖人工智能精准识别物体并理解环境,但即便基于数百万英里的真实数据训练,仍难以察觉罕见失败模式(RFMs)。这类问题被称为“长尾挑战”,源于数据中存在大量极少见的实例。本文提出一种新方法,结合先进生成式与可解释AI技术,以揭示RFMs。我们提取目标物体(如车辆)的分割掩码,并反向生成环境掩码;结合精心设计的文本提示,输入定制扩散模型。利用受对抗噪声优化引导的Stable Diffusion修复模型,生成能规避目标检测模型的多样化环境图像,暴露出AI系统的潜在漏洞。最终生成自然语言描述,帮助开发者与政策制定者提升自动驾驶系统的安全性和可靠性。
原文摘要 · Abstract (English)
Autonomous Vehicles (AVs) rely on artificial intelligence (AI) to accurately detect objects and interpret their surroundings. However, even when trained using millions of miles of real-world data, AVs are often unable to detect rare failure modes (RFMs). The problem of RFMs is commonly referred to as the "long-tail challenge", due to the distribution of data including many instances that are very rarely seen. In this paper, we present a novel approach that utilizes advanced generative and explainable AI techniques to aid in understanding RFMs. Our methods can be used to enhance the robustness and reliability of AVs when combined with both downstream model training and testing. We extract segmentation masks for objects of interest (e.g., cars) and invert them to create environmental masks. These masks, combined with carefully crafted text prompts, are fed into a custom diffusion model. We leverage the Stable Diffusion inpainting model guided by adversarial noise optimization to generate images containing diverse environments designed to evade object detection models and expose vulnerabilities in AI systems. Finally, we produce natural language descriptions of the generated RFMs that can guide developers and policymakers to improve the safety and reliability of AV systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。