arXiv:2505.14531cs.LG2025-05被引 2

无需模型细节,可清除各类视觉模型中的恶意触发器。

SifterNet: A Generalized and Model-Agnostic Trigger Purification Approach

  • 基于霍普菲尔德网络的关联记忆机制净化输入触发器。
  • 在多个数据集上实现高精度且有效清除后门触发器。
  • 黑盒适用,不依赖模型结构或训练权限,适合实际部署。

针对卷积神经网络和基于视觉变换器的大模型中的后门攻击,本文提出一种通用且模型无关的触发器净化方法SifterNet,借助经典的伊辛模型思想。现有触发器检测与清除方法通常需预先掌握目标模型细节、大量干净样本甚至重训练授权,极大限制了实际应用,尤其在无法访问目标模型时难以实施。理想的防御手段应不依赖具体模型即可消除植入的触发器。为此,SifterNet通过利用霍普菲尔德网络的记忆-关联功能,在无需了解目标模型的前提下,实现对输入样本中触发器的有效净化。其核心创新在于引入伊辛模型的思想。大量实验验证了该方法在触发器净化效果和精度保持方面的有效性,相较于当前主流基线,在多个常用数据集上均表现出显著优势。

原文摘要 · Abstract (English)

Aiming at resisting backdoor attacks in convolution neural networks and vision Transformer-based large model, this paper proposes a generalized and model-agnostic trigger-purification approach resorting to the classic Ising model. To date, existing trigger detection/removal studies usually require to know the detailed knowledge of target model in advance, access to a large number of clean samples or even model-retraining authorization, which brings the huge inconvenience for practical applications, especially inaccessible to target model. An ideal countermeasure ought to eliminate the implanted trigger without regarding whatever the target models are. To this end, a lightweight and black-box defense approach SifterNet is proposed through leveraging the memorization-association functionality of Hopfield network, by which the triggers of input samples can be effectively purified in a proper manner. The main novelty of our proposed approach lies in the introduction of ideology of Ising model. Extensive experiments also validate the effectiveness of our approach in terms of proper trigger purification and high accuracy achievement, and compared to the state-of-the-art baselines under several commonly-used datasets, our SiferNet has a significant superior performance.

后门防御触发器净化黑盒防御伊辛模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。