利用自然磨损生成更逼真的物理对抗样本,提升攻击效果与隐蔽性。
Adversarial Wear and Tear: Exploiting Natural Damage for Generating Physical-World Adversarial Examples
- 基于生成对抗网络建模自然磨损风格,通过可学习的损伤代码实现可控扰动。
- 在两个交通标志数据集上攻击成功率超90%,且视觉真实感更强。
- 适用于研究模型鲁棒性或安全验证的工程师与研究人员。
物理世界中的对抗样本对自动驾驶等安全关键场景的深度神经网络部署构成重大挑战。现有方法多依赖临时修改(如阴影、激光或贴纸),且针对特定场景。本文提出一种新类物理对抗样本AdvWT,灵感源自物体自然磨损这一固有现象。不同于人工构造的扰动,磨损随时间自然产生,如户外路牌的逐渐退化。AdvWT采用两步策略:首先,使用基于GAN的无监督图像到图像转换网络,建模路牌等物体的自然损伤特征,将损伤特性编码为潜在的‘损伤风格码’;其次,在该风格码中引入对抗扰动,优化其变换过程,使生成的损伤外观仍具现实感,同时有效误导神经网络。在两个交通标志数据集上的实验表明,AdvWT在数字与物理域均能高效误导深度神经网络,攻击成功率高于现有方法,且具有更强鲁棒性与自然外观。此外,将AdvWT融入训练可增强模型对真实损坏标志的泛化能力。
原文摘要 · Abstract (English)
The presence of adversarial examples in the physical world poses significant challenges to the deployment of Deep Neural Networks in safety-critical applications such as autonomous driving. Most existing methods for crafting physical-world adversarial examples are ad-hoc, relying on temporary modifications like shadows, laser beams, or stickers that are tailored to specific scenarios. In this paper, we introduce a new class of physical-world adversarial examples, AdvWT, which draws inspiration from the naturally occurring phenomenon of `wear and tear', an inherent property of physical objects. Unlike manually crafted perturbations, `wear and tear' emerges organically over time due to environmental degradation, as seen in the gradual deterioration of outdoor signboards. To achieve this, AdvWT follows a two-step approach. First, a GAN-based, unsupervised image-to-image translation network is employed to model these naturally occurring damages, particularly in the context of outdoor signboards. The translation network encodes the characteristics of damaged signs into a latent `damage style code'. In the second step, we introduce adversarial perturbations into the style code, strategically optimizing its transformation process. This manipulation subtly alters the damage style representation, guiding the network to generate adversarial images where the appearance of damages remains perceptually realistic, while simultaneously ensuring their effectiveness in misleading neural networks. Through comprehensive experiments on two traffic sign datasets, we show that AdvWT effectively misleads DNNs in both digital and physical domains. AdvWT achieves an effective attack success rate, greater robustness, and a more natural appearance compared to existing physical-world adversarial examples. Additionally, integrating AdvWT into training enhances a model's generalizability to real-world damaged signs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。