arXiv:2511.10382cs.CV2025-11

现有隐私保护方法在个性化生成中易被简单滤镜破解,效果虚假可靠。

Fragile by Design: On the Limits of Adversarial Defenses in Personalized Generation

  • 用对抗扰动伪装用户图像以防止身份泄露
  • 多数方法经简单滤镜处理后即失效
  • 适合关注生成式模型隐私安全的研究者

个性化AI应用如DreamBooth可基于用户图像生成定制内容,但存在面部身份泄露风险。近期防御方法如Anti-DreamBooth通过向用户照片注入对抗扰动来缓解此风险。然而我们发现两大被忽视的局限:其一,对抗样本常含明显伪影(如条纹或图案),易被识别为篡改内容;其二,扰动极脆弱,即使使用简单的非学习型滤镜也能有效去除,使模型恢复对用户身份的记忆与复现能力。为此,我们提出新型评估框架AntiDB_Purify,系统评估现有防御在真实净化威胁下的表现,涵盖传统图像滤镜与对抗净化。结果表明,当前所有方法在该威胁下均失去防护效力。这揭示了现有防御仅提供虚假安全感,亟需更不可见且鲁棒的保护机制以保障个性化生成中的用户身份安全。

原文摘要 · Abstract (English)

Personalized AI applications such as DreamBooth enable the generation of customized content from user images, but also raise significant privacy concerns, particularly the risk of facial identity leakage. Recent defense mechanisms like Anti-DreamBooth attempt to mitigate this risk by injecting adversarial perturbations into user photos to prevent successful personalization. However, we identify two critical yet overlooked limitations of these methods. First, the adversarial examples often exhibit perceptible artifacts such as conspicuous patterns or stripes, making them easily detectable as manipulated content. Second, the perturbations are highly fragile, as even a simple, non-learned filter can effectively remove them, thereby restoring the model's ability to memorize and reproduce user identity. To investigate this vulnerability, we propose a novel evaluation framework, AntiDB_Purify, to systematically evaluate existing defenses under realistic purification threats, including both traditional image filters and adversarial purification. Results reveal that none of the current methods maintains their protective effectiveness under such threats. These findings highlight that current defenses offer a false sense of security and underscore the urgent need for more imperceptible and robust protections to safeguard user identity in personalized generation.

隐私保护对抗攻击生成模型身份泄露

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。