用可控图像编辑实现车辆隐身攻击,骗过检测器又不被人类察觉
In-the-Wild Camouflage Attack on Vehicle Detectors through Controllable Image Editing
- 将车辆隐身攻击建模为条件图像编辑,用ControlNet直接生成真实场景中的伪装车
- 在COCO和LINZ数据集上使检测准确率下降超38%,结构保持更好、人眼更难察觉
- 可攻击未知检测模型,且具备物理世界适用潜力,适合研究对抗防御的学者
深度神经网络在计算机视觉中取得显著成功,但仍极易受到对抗攻击。其中,隐身攻击通过改变物体的可见外观来欺骗检测器,同时对人类保持隐蔽。本文提出一种新框架,将车辆隐身攻击建模为条件图像编辑问题。具体地,探索了图像级与场景级的隐身生成策略,并微调ControlNet直接在真实图像上合成伪装车辆。设计统一目标函数,联合约束车辆结构保真度、风格一致性与对抗有效性。在COCO和LINZ数据集上的大量实验表明,该方法攻击效果显著更强,使AP50下降超过38%,同时更好地保留车辆结构并提升人眼感知的隐蔽性。此外,该框架对未见过的黑盒检测器具有良好泛化能力,并展现出向物理世界迁移的潜力。项目页面见https://humansensinglab.github.io/CtrlCamo。
原文摘要 · Abstract (English)
Deep neural networks (DNNs) have achieved remarkable success in computer vision but remain highly vulnerable to adversarial attacks. Among them, camouflage attacks manipulate an object's visible appearance to deceive detectors while remaining stealthy to humans. In this paper, we propose a new framework that formulates vehicle camouflage attacks as a conditional image-editing problem. Specifically, we explore both image-level and scene-level camouflage generation strategies, and fine-tune a ControlNet to synthesize camouflaged vehicles directly on real images. We design a unified objective that jointly enforces vehicle structural fidelity, style consistency, and adversarial effectiveness. Extensive experiments on the COCO and LINZ datasets show that our method achieves significantly stronger attack effectiveness, leading to more than 38% AP50 decrease, while better preserving vehicle structure and improving human-perceived stealthiness compared to existing approaches. Furthermore, our framework generalizes effectively to unseen black-box detectors and exhibits promising transferability to the physical world. Project page is available at https://humansensinglab.github.io/CtrlCamo
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。