arXiv:2412.18302cs.CVcs.CR2024-12

攻击者可操纵文本嵌入,让文生图模型生成特定名人图像。

FameBias: Embedding Manipulation Bias Attack in Text-to-Image Models

  • 通过修改输入文本的嵌入向量实现攻击,无需重新训练模型。
  • 在多个触发词与目标人物组合下,攻击成功率高且保持语义一致。
  • 适合研究安全风险或对抗样本防御的研究者关注。

文生图扩散模型虽能生成高质量图像,但其可能被滥用于传播虚假信息。近期研究发现,攻击者可通过简单微调在模型中植入偏见,使特定关键词触发特定图像生成。本文提出FameBias,一种仅操作输入嵌入向量的偏见攻击方法,无需额外训练即可生成包含特定公众人物的图像。我们在Stable Diffusion V2上全面评估该方法,基于多种触发词与目标人物生成大量图像。实验表明,FameBias在多组触发-目标对中均实现高攻击成功率,同时保持原始提示的语义上下文完整性。

原文摘要 · Abstract (English)

Text-to-Image (T2I) diffusion models have rapidly advanced, enabling the generation of high-quality images that align closely with textual descriptions. However, this progress has also raised concerns about their misuse for propaganda and other malicious activities. Recent studies reveal that attackers can embed biases into these models through simple fine-tuning, causing them to generate targeted imagery when triggered by specific phrases. This underscores the potential for T2I models to act as tools for disseminating propaganda, producing images aligned with an attacker's objective for end-users. Building on this concept, we introduce FameBias, a T2I biasing attack that manipulates the embeddings of input prompts to generate images featuring specific public figures. Unlike prior methods, Famebias operates solely on the input embedding vectors without requiring additional model training. We evaluate FameBias comprehensively using Stable Diffusion V2, generating a large corpus of images based on various trigger nouns and target public figures. Our experiments demonstrate that FameBias achieves a high attack success rate while preserving the semantic context of the original prompts across multiple trigger-target pairs.

文生图对抗攻击偏见注入

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。