arXiv:2510.26105cs.CVcs.AI2025-10

攻击者可仅改图生成违规内容,暴露多模态模型安全漏洞

Security Risk of Misalignment between Text and Image in Multi-modal Model

  • 通过修改输入图生成违规内容,不改动提示词
  • 在多种模型上验证成功,图像修复与风格迁移任务均有效
  • 首个仅靠恶意图像攻击的多模态攻击,威胁图像编辑应用

尽管多模态扩散模型(如文生图模型)取得了显著进展,其对对抗性输入的脆弱性仍研究不足。我们的研究发现,现有扩散模型中文本与图像模态的对齐能力不足,导致生成不当或不适宜工作场合(NSFW)内容的风险。为此,我们提出一种新型攻击方法——提示受限多模态攻击(PReMA),通过修改输入图像并结合任意指定提示词,在不改变提示词本身的情况下操控生成结果。PReMA是首个仅通过生成对抗性图像来操纵模型输出的攻击,区别于以往主要依赖对抗性提示生成NSFW内容的方法。在多个模型上的图像修复和风格迁移任务中,全面评估证实了PReMA的强大有效性。

原文摘要 · Abstract (English)

Despite the notable advancements and versatility of multi-modal diffusion models, such as text-to-image models, their susceptibility to adversarial inputs remains underexplored. Contrary to expectations, our investigations reveal that the alignment between textual and Image modalities in existing diffusion models is inadequate. This misalignment presents significant risks, especially in the generation of inappropriate or Not-Safe-For-Work (NSFW) content. To this end, we propose a novel attack called Prompt-Restricted Multi-modal Attack (PReMA) to manipulate the generated content by modifying the input image in conjunction with any specified prompt, without altering the prompt itself. PReMA is the first attack that manipulates model outputs by solely creating adversarial images, distinguishing itself from prior methods that primarily generate adversarial prompts to produce NSFW content. Consequently, PReMA poses a novel threat to the integrity of multi-modal diffusion models, particularly in image-editing applications that operate with fixed prompts. Comprehensive evaluations conducted on image inpainting and style transfer tasks across various models confirm the potent efficacy of PReMA.

多模态安全对抗攻击扩散模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。