无需文字提示,悄悄在图像中植入品牌标识
Silent Branding Attack: Trigger-free Data Poisoning Attack on Text-to-Image Diffusion Models
- 利用重复视觉模式让模型无意识生成指定标志
- 攻击后图像质量不降,标志隐蔽且能被检测到
- 适合研究模型安全与隐秘水印的学者使用
文本生成图像扩散模型在高质量内容生成方面取得显著进展。然而,其依赖公开数据并日益普遍的数据共享趋势使其易受数据投毒攻击。本文提出一种名为“静默品牌攻击”的新型投毒方法,可使模型在无任何文本触发的情况下生成包含特定品牌标识或符号的图像。我们发现,当某些视觉模式在训练数据中反复出现时,模型会自然地在输出中重现这些模式,即使未在提示中提及。基于此,我们开发了一种自动化投毒算法,将标志以自然方式嵌入原始图像,实现无缝融合且难以察觉。经训练的模型在生成图像时会自动包含标志,且不降低图像质量或破坏文本对齐。我们在大规模高质量图像数据集和风格个性化数据集上验证了该攻击在两种现实场景下的有效性,成功率高,且无需特定文本触发。人工评估及量化指标(包括标志检测)表明,该方法可隐蔽地嵌入标志。
原文摘要 · Abstract (English)
Text-to-image diffusion models have achieved remarkable success in generating high-quality contents from text prompts. However, their reliance on publicly available data and the growing trend of data sharing for fine-tuning make these models particularly vulnerable to data poisoning attacks. In this work, we introduce the Silent Branding Attack, a novel data poisoning method that manipulates text-to-image diffusion models to generate images containing specific brand logos or symbols without any text triggers. We find that when certain visual patterns are repeatedly in the training data, the model learns to reproduce them naturally in its outputs, even without prompt mentions. Leveraging this, we develop an automated data poisoning algorithm that unobtrusively injects logos into original images, ensuring they blend naturally and remain undetected. Models trained on this poisoned dataset generate images containing logos without degrading image quality or text alignment. We experimentally validate our silent branding attack across two realistic settings on large-scale high-quality image datasets and style personalization datasets, achieving high success rates even without a specific text trigger. Human evaluation and quantitative metrics including logo detection show that our method can stealthily embed logos.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。