arXiv:2410.13010cs.LGcs.AI2024-10

让目标物体在图像中‘消失’,却让模型以为它本就不存在。

Hiding-in-Plain-Sight (HiPS) Attack on CLIP for Targetted Object Removal from Images

  • 通过隐藏目标物体实现隐秘攻击,不改变原图外观。
  • 成功让下游图像描述模型完全忽略指定物体。
  • 适合研究模型隐私与视觉安全的学者参考。

机器学习模型易受对抗攻击,但传统攻击多针对单一模态。随着像CLIP这样的大型多模态模型兴起,新漏洞随之出现。以往多模态攻击旨在彻底改变模型输出,但在真实场景中,攻击者可能只需细微改动,使变化难以被下游模型或人类察觉。本文提出一种新型对抗攻击——‘显眼处隐藏’(Hiding-in-Plain-Sight, HiPS),通过选择性地隐藏目标物体,使模型误以为该物体不存在。提出两种变体:HiPS-cls与HiPS-cap,实验证明其能有效迁移至下游图像描述模型(如CLIP-Cap),实现目标物体在图像描述中的精准移除。

原文摘要 · Abstract (English)

Machine learning models are known to be vulnerable to adversarial attacks, but traditional attacks have mostly focused on single-modalities. With the rise of large multi-modal models (LMMs) like CLIP, which combine vision and language capabilities, new vulnerabilities have emerged. However, prior work in multimodal targeted attacks aim to completely change the model's output to what the adversary wants. In many realistic scenarios, an adversary might seek to make only subtle modifications to the output, so that the changes go unnoticed by downstream models or even by humans. We introduce Hiding-in-Plain-Sight (HiPS) attacks, a novel class of adversarial attacks that subtly modifies model predictions by selectively concealing target object(s), as if the target object was absent from the scene. We propose two HiPS attack variants, HiPS-cls and HiPS-cap, and demonstrate their effectiveness in transferring to downstream image captioning models, such as CLIP-Cap, for targeted object removal from image captions.

对抗攻击多模态隐匿攻击

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。