arXiv:2409.17874cs.AI2024-09NeurIPS被引 43

DarkSAM用通用干扰让SAM无法分割任何物体

DarkSAM: Fooling Segment Anything Model to Segment Nothing

  • 不依赖提示的攻击,从空间和频域破坏关键特征
  • 单个扰动即可让SAM在多数据集失效
  • 揭示SAM在语义与纹理上的脆弱性,适合安全研究者

分割一切模型(SAM)因其出色的泛化能力备受关注。然而,其对通用对抗扰动(UAP)的脆弱性尚未被充分研究。本文提出DarkSAM,首个无需提示的通用攻击框架,包含基于语义解耦的空间攻击与基于纹理扭曲的频域攻击。首先将SAM输出分为前景与背景,设计阴影目标策略获取图像语义蓝图作为攻击目标。DarkSAM致力于在空间与频域中提取并破坏图像的关键对象特征。空间域中扰乱前景与背景语义以混淆SAM;频域中通过扭曲高频成分(即纹理信息)进一步增强攻击效果。结果表明,仅需单一UAP,DarkSAM即可使SAM在多种图像上、不同提示下均无法完成分割。在四个数据集及两个变体模型上的实验验证了DarkSAM的强大攻击能力与迁移性。

原文摘要 · Abstract (English)

Segment Anything Model (SAM) has recently gained much attention for its outstanding generalization to unseen data and tasks. Despite its promising prospect, the vulnerabilities of SAM, especially to universal adversarial perturbation (UAP) have not been thoroughly investigated yet. In this paper, we propose DarkSAM, the first prompt-free universal attack framework against SAM, including a semantic decoupling-based spatial attack and a texture distortion-based frequency attack. We first divide the output of SAM into foreground and background. Then, we design a shadow target strategy to obtain the semantic blueprint of the image as the attack target. DarkSAM is dedicated to fooling SAM by extracting and destroying crucial object features from images in both spatial and frequency domains. In the spatial domain, we disrupt the semantics of both the foreground and background in the image to confuse SAM. In the frequency domain, we further enhance the attack effectiveness by distorting the high-frequency components (i.e., texture information) of the image. Consequently, with a single UAP, DarkSAM renders SAM incapable of segmenting objects across diverse images with varying prompts. Experimental results on four datasets for SAM and its two variant models demonstrate the powerful attack capability and transferability of DarkSAM.

视觉安全对抗攻击SAM

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。