arXiv:2507.17554cs.CV2025-07中稿 · ACM Multimedia 202…被引 2

利用隐空间对抗攻击,防止扩散模型用少量样本个性化生成隐私内容。

An h-space Based Adversarial Attack for Protection Against Few-shot Personalization

论文配图:An h-space Based Adversarial Attack for Protection Against Few-shot Personalization
图 1 · 摘自论文原文
  • 在语义隐空间(h-space)中设计对抗扰动,干扰模型生成
  • 新方法在少样本个性化场景下显著降低生成质量,优于现有技术
  • 只需修改键值参数,效率高,适合部署于隐私保护系统

扩散模型能仅用少量样本生成定制化图像,带来严重的隐私风险,尤其是未经授权的内容篡改。为此,研究者致力于开发基于对抗攻击的保护机制,通过生成有效扰动来污染扩散模型。本文观察到,这类模型在语义隐空间(h-space)中具有高度抽象性,该空间编码了生成连贯内容的关键高层特征。为此提出新型反个性化方法HAAD(基于h-space的扩散模型对抗攻击),利用h-space中的对抗扰动有效破坏图像生成过程。在此基础上,进一步提出更高效的HAAD-KV版本,仅基于h-space的键值(KV)参数构造扰动,计算开销更低但防护更强。尽管结构简单,本方法在性能上超越现有先进攻击手段,证明其有效性。

原文摘要 · Abstract (English)

The versatility of diffusion models in generating customized images from few samples raises significant privacy concerns, particularly regarding unauthorized modifications of private content. This concerning issue has renewed the efforts in developing protection mechanisms based on adversarial attacks, which generate effective perturbations to poison diffusion models. Our work is motivated by the observation that these models exhibit a high degree of abstraction within their semantic latent space (`h-space'), which encodes critical high-level features for generating coherent and meaningful content. In this paper, we propose a novel anti-customization approach, called HAAD (h-space based Adversarial Attack for Diffusion models), that leverages adversarial attacks to craft perturbations based on the h-space that can efficiently degrade the image generation process. Building upon HAAD, we further introduce a more efficient variant, HAAD-KV, that constructs perturbations solely based on the KV parameters of the h-space. This strategy offers a stronger protection, that is computationally less expensive. Despite their simplicity, our methods outperform state-of-the-art adversarial attacks, highlighting their effectiveness.

扩散模型隐私保护对抗攻击少样本生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。