用文本引导扩散模型攻击图像哈希检索,更精准且更难被发现。
DiffHash: Text-Guided Targeted Attack via Diffusion Models against Deep Hashing Image Retrieval
- 通过大语言模型生成文本提示,指导图像隐空间优化攻击。
- 在多个数据集上实现更高攻击成功率和更强的黑盒迁移能力。
- 适合研究模型安全、对抗样本防御的学者与工程师。
深度哈希模型广泛用于大规模图像检索,但易受对抗样本攻击。现有目标攻击方法仍存在多模态引导不足、依赖标签信息及像素级操作等问题。为此,我们提出 DiffHash,一种基于扩散模型的目标攻击方法。不同于传统像素级修改,该方法通过大语言模型(LLM)生成的目标图像文本提示,引导优化图像隐表示,并设计多空间哈希对齐网络,将高维图像与文本空间映射至低维二值哈希空间。重建时引入文本引导注意力机制,确保对抗样本在保持视觉合理性的同时准确匹配目标语义。大量实验表明,该方法优于当前最先进(SOTA)的靶向攻击方法,在跨数据集测试中展现出更优的黑盒迁移性与稳定性。
原文摘要 · Abstract (English)
Deep hashing models have been widely adopted to tackle the challenges of large-scale image retrieval. However, these approaches face serious security risks due to their vulnerability to adversarial examples. Despite the increasing exploration of targeted attacks on deep hashing models, existing approaches still suffer from a lack of multimodal guidance, reliance on labeling information and dependence on pixel-level operations for attacks. To address these limitations, we proposed DiffHash, a novel diffusion-based targeted attack for deep hashing. Unlike traditional pixel-based attacks that directly modify specific pixels and lack multimodal guidance, our approach focuses on optimizing the latent representations of images, guided by text information generated by a Large Language Model (LLM) for the target image. Furthermore, we designed a multi-space hash alignment network to align the high-dimension image space and text space to the low-dimension binary hash space. During reconstruction, we also incorporated text-guided attention mechanisms to refine adversarial examples, ensuring them aligned with the target semantics while maintaining visual plausibility. Extensive experiments have demonstrated that our method outperforms state-of-the-art (SOTA) targeted attack methods, achieving better black-box transferability and offering more excellent stability across datasets.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。