arXiv:2507.21985cs.CVcs.CR2025-07ICCV

零样本攻击未学习模型,精准生成符合意图的违规内容

ZIUM: Zero-Shot Intent-Aware Adversarial Attack on Unlearned Models

  • 无需优化即可攻击已移除概念,实现零样本攻击
  • 攻击成功率高于现有方法,可按用户意图定制输出内容
  • 显著降低攻击耗时,适合快速验证模型隐私漏洞

机器去学习(MU)通过从深度学习模型中移除特定数据或概念来增强隐私保护并防止敏感内容生成。对抗性提示可利用已去学习模型生成包含被移除概念的内容,构成重大安全风险。然而,现有对抗攻击方法在生成符合攻击者意图的内容方面仍面临挑战,且需高昂计算成本以识别有效提示。为此,我们提出ZIUM——一种零样本意图感知的未学习模型对抗攻击方法,支持灵活定制目标攻击图像以反映攻击者意图。此外,ZIUM可在无需对已攻击过的去学习概念进行额外优化的情况下实现零样本攻击。在多种去学习场景下的评估表明,ZIUM能有效根据用户意图提示生成内容,攻击成功率优于现有方法;同时,其零样本攻击显著减少了对已攻击概念的攻击时间。

原文摘要 · Abstract (English)

Machine unlearning (MU) removes specific data points or concepts from deep learning models to enhance privacy and prevent sensitive content generation. Adversarial prompts can exploit unlearned models to generate content containing removed concepts, posing a significant security risk. However, existing adversarial attack methods still face challenges in generating content that aligns with an attacker's intent while incurring high computational costs to identify successful prompts. To address these challenges, we propose ZIUM, a Zero-shot Intent-aware adversarial attack on Unlearned Models, which enables the flexible customization of target attack images to reflect an attacker's intent. Additionally, ZIUM supports zero-shot adversarial attacks without requiring further optimization for previously attacked unlearned concepts. The evaluation across various MU scenarios demonstrated ZIUM's effectiveness in successfully customizing content based on user-intent prompts while achieving a superior attack success rate compared to existing methods. Moreover, its zero-shot adversarial attack significantly reduces the attack time for previously attacked unlearned concepts.

对抗攻击隐私安全去学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。