arXiv:2411.16622cs.CVcs.AI2024-11被引 3

用可微渲染技术实现物理世界中人眼不可察觉的对抗样本

Imperceptible Adversarial Examples in the Physical World

  • 通过直通估计器处理视觉传感非可微问题,实现物理对抗样本生成
  • 在真实打印图像和CARLA模拟中实现ℓ∞有界对抗攻击,准确率降至0%
  • 首次展示物理世界中高隐蔽性对抗样本,适合安全与防御研究者关注

数字域中的对抗样本对基于深度学习的计算机视觉模型具有人眼不可察觉的扰动。但在物理世界中,由于视觉传感系统的非可微失真函数,生成类似对抗样本仍具挑战。现有方法常放宽定义,允许无界扰动,导致明显或异常的视觉模式。本文采用直通估计器(STE,又称BPDA)克服非可微性:前向传播使用精确的非可微失真,反向传播则用恒等函数。我们进一步扩展了可微渲染到STE,实现了物理世界中人眼不可察觉的对抗补丁。通过打印照片和CARLA模拟实验,证明STE能在非可微失真下快速生成ℓ∞有界对抗样本。据我们所知,这是首个在物理世界中实现小ℓ∞范数有界、全局扰动威胁模型下分类准确率为0%、补丁扰动下AP50仅为4.22%的不可察觉对抗样本的工作。我们呼吁学术界重新评估物理世界中对抗样本的威胁。

原文摘要 · Abstract (English)

Adversarial examples in the digital domain against deep learning-based computer vision models allow for perturbations that are imperceptible to human eyes. However, producing similar adversarial examples in the physical world has been difficult due to the non-differentiable image distortion functions in visual sensing systems. The existing algorithms for generating physically realizable adversarial examples often loosen their definition of adversarial examples by allowing unbounded perturbations, resulting in obvious or even strange visual patterns. In this work, we make adversarial examples imperceptible in the physical world using a straight-through estimator (STE, a.k.a. BPDA). We employ STE to overcome the non-differentiability -- applying exact, non-differentiable distortions in the forward pass of the backpropagation step, and using the identity function in the backward pass. Our differentiable rendering extension to STE also enables imperceptible adversarial patches in the physical world. Using printout photos, and experiments in the CARLA simulator, we show that STE enables fast generation of $\ell_\infty$ bounded adversarial examples despite the non-differentiable distortions. To the best of our knowledge, this is the first work demonstrating imperceptible adversarial examples bounded by small $\ell_\infty$ norms in the physical world that force zero classification accuracy in the global perturbation threat model and cause near-zero ($4.22\%$) AP50 in object detection in the patch perturbation threat model. We urge the community to re-evaluate the threat of adversarial examples in the physical world.

对抗样本物理世界可微渲染视觉安全

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。