arXiv:2506.21874cs.CRcs.AI2025-06被引 4

用对抗扰动骗VLM误标图像,低成本污染文生图模型训练数据

On the Feasibility of Poisoning Text-to-Image AI Models via Adversarial Mislabeling

  • 通过对抗扰动让VLM错误生成图像描述,制造脏标签样本
  • 仅用少量样本即可成功改变文生图模型行为,攻击成功率超73%
  • 黑盒攻击对商用VLM有效,现有防御易被绕过

当前文生图模型依赖互联网上数百万张图像及其由视觉语言模型(VLM)生成的详细描述进行训练。然而,近期研究表明VLM容易受到隐蔽对抗攻击:向图像添加微小扰动可误导VLM生成错误描述。本文探索利用此类对抗误标攻击,作为污染文生图模型训练流程的手段。实验表明,VLM对对抗扰动高度敏感,攻击者可生成外观正常的图像,使其被持续误标。这导致大量“脏标签”样本注入训练数据,仅需少量中毒样本即能显著改变文生图模型的行为。尽管存在潜在防御措施,但可被自适应攻击者针对性绕过。该攻防博弈可能降低训练数据质量,推高模型开发成本。最终实验证明,该攻击在黑盒场景下对商业VLM(Google Vertex AI、Microsoft Azure)仍具高成功率(>73%)。

原文摘要 · Abstract (English)

Today's text-to-image generative models are trained on millions of images sourced from the Internet, each paired with a detailed caption produced by Vision-Language Models (VLMs). This part of the training pipeline is critical for supplying the models with large volumes of high-quality image-caption pairs during training. However, recent work suggests that VLMs are vulnerable to stealthy adversarial attacks, where adversarial perturbations are added to images to mislead the VLMs into producing incorrect captions. In this paper, we explore the feasibility of adversarial mislabeling attacks on VLMs as a mechanism to poisoning training pipelines for text-to-image models. Our experiments demonstrate that VLMs are highly vulnerable to adversarial perturbations, allowing attackers to produce benign-looking images that are consistently miscaptioned by the VLM models. This has the effect of injecting strong "dirty-label" poison samples into the training pipeline for text-to-image models, successfully altering their behavior with a small number of poisoned samples. We find that while potential defenses can be effective, they can be targeted and circumvented by adaptive attackers. This suggests a cat-and-mouse game that is likely to reduce the quality of training data and increase the cost of text-to-image model development. Finally, we demonstrate the real-world effectiveness of these attacks, achieving high attack success (over 73%) even in black-box scenarios against commercial VLMs (Google Vertex AI and Microsoft Azure).

对抗攻击文生图数据污染VLM

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。