无需训练数据即可植入后门,同时保持图像检索精度。
DarkHash: A Data-Free Backdoor Attack Against Deep Hashing
- 设计双语义引导的影子攻击框架,仅微调特定层。
- 在四个数据集上均实现高攻击成功率,超越现有方法。
- 适合研究模型安全与后门防御的学者参考。
得益于出色的特征学习能力和高效性,深度哈希在大规模图像检索中取得了显著成功。近期研究揭示了深度哈希模型易受后门攻击的影响。尽管已有研究取得良好攻击效果,但均依赖于访问训练数据以植入后门。现实中,出于隐私保护和知识产权考虑,获取此类数据(如身份信息)通常被禁止。在无训练数据的前提下,向深度哈希模型嵌入后门并维持原始任务检索精度,是一个新颖且具有挑战性的问题。本文提出 DarkHash,首个针对深度哈希的数据无关后门攻击方法。具体而言,设计了一种新型影子后门攻击框架,采用双语义引导机制,通过使用替代数据集仅微调目标模型的特定层,实现后门功能植入与原始检索精度的保持。通过利用样本与其邻域间的关系,设计拓扑对齐损失,优化单个及邻近污染样本朝向目标样本,进一步提升攻击能力。在四个图像数据集、五种模型架构和两种哈希方法上的实验表明,DarkHash 具有极高有效性,优于现有最先进后门攻击方法。防御实验显示,DarkHash 能抵御现有主流后门防御手段。
原文摘要 · Abstract (English)
Benefiting from its superior feature learning capabilities and efficiency, deep hashing has achieved remarkable success in large-scale image retrieval. Recent studies have demonstrated the vulnerability of deep hashing models to backdoor attacks. Although these studies have shown promising attack results, they rely on access to the training dataset to implant the backdoor. In the real world, obtaining such data (e.g., identity information) is often prohibited due to privacy protection and intellectual property concerns. Embedding backdoors into deep hashing models without access to the training data, while maintaining retrieval accuracy for the original task, presents a novel and challenging problem. In this paper, we propose DarkHash, the first data-free backdoor attack against deep hashing. Specifically, we design a novel shadow backdoor attack framework with dual-semantic guidance. It embeds backdoor functionality and maintains original retrieval accuracy by fine-tuning only specific layers of the victim model using a surrogate dataset. We consider leveraging the relationship between individual samples and their neighbors to enhance backdoor attacks during training. By designing a topological alignment loss, we optimize both individual and neighboring poisoned samples toward the target sample, further enhancing the attack capability. Experimental results on four image datasets, five model architectures, and two hashing methods demonstrate the high effectiveness of DarkHash, outperforming existing state-of-the-art backdoor attack methods. Defense experiments show that DarkHash can withstand existing mainstream backdoor defense methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。