arXiv:2502.19697cs.CV2025-02ICCV被引 2

通过破坏属性文本嵌入,提升行人重识别的可迁移攻击效果。

Prompt-driven Transferable Adversarial Attack on Person Re-Identification with Attribute-aware Textual Inversion

  • 利用视觉语言模型对齐能力,针对性破坏行人属性的文本表示。
  • 在跨模型与跨数据集攻击中,平均降率提升22.9%,达到领先水平。
  • 适合研究对抗攻击、安全评估与细粒度特征敏感性的研究人员。

行人重识别(re-id)模型在安防系统中至关重要,需通过可迁移对抗攻击探测其漏洞。近期基于视觉-语言模型(VLM)的攻击方法虽展现优越可迁移性,但因过度强调整体表征中的判别语义,导致特征破坏不充分。本文提出属性感知提示攻击(AP-Attack),利用VLM的图文对齐能力,通过破坏特定属性的文本嵌入,显式干扰行人图像的细粒度语义特征。设计文本反演网络,将行人图像映射为表示语义嵌入的伪标记,以对比学习方式在预定义属性提示模板下训练。良性与对抗性细粒度文本语义的反演,使攻击者能实现更彻底的破坏,显著提升对抗样本的可迁移性。大量实验表明,AP-Attack在跨模型&数据集攻击场景中,平均降率(mean Drop Rate)相比先前方法提升22.9%,达到当前最优性能。

原文摘要 · Abstract (English)

Person re-identification (re-id) models are vital in security surveillance systems, requiring transferable adversarial attacks to explore the vulnerabilities of them. Recently, vision-language models (VLM) based attacks have shown superior transferability by attacking generalized image and textual features of VLM, but they lack comprehensive feature disruption due to the overemphasis on discriminative semantics in integral representation. In this paper, we introduce the Attribute-aware Prompt Attack (AP-Attack), a novel method that leverages VLM's image-text alignment capability to explicitly disrupt fine-grained semantic features of pedestrian images by destroying attribute-specific textual embeddings. To obtain personalized textual descriptions for individual attributes, textual inversion networks are designed to map pedestrian images to pseudo tokens that represent semantic embeddings, trained in the contrastive learning manner with images and a predefined prompt template that explicitly describes the pedestrian attributes. Inverted benign and adversarial fine-grained textual semantics facilitate attacker in effectively conducting thorough disruptions, enhancing the transferability of adversarial examples. Extensive experiments show that AP-Attack achieves state-of-the-art transferability, significantly outperforming previous methods by 22.9% on mean Drop Rate in cross-model&dataset attack scenarios.

对抗攻击行人重识别视觉语言模型文本反演

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。