arXiv:2501.17151cs.LG2025-01中稿 · the Thirty-Eighth …被引 6

通过异常分布样本检测模型后门,无需预设攻击方式。

Scanning Trojaned Models Using Out-of-Distribution Samples

  • 利用分布外样本的对抗扰动识别后门盲区。
  • 对对抗训练的后门模型仍保持高检测率。
  • 无需训练数据,适用于多种场景。

深度神经网络的后门扫描至关重要,因其在现实应用中广泛使用。现有方法多依赖对攻击方式的先验假设,且难以检测对抗训练生成的后门模型。为此,本文提出TRODO(基于分布外样本对抗偏移的后门检测)方法,利用“盲区”机制——即后门模型将分布外样本误判为分布内样本的现象。通过对抗性地将分布外样本向分布内方向扰动,若扰动后样本被分类为分布内,则表明存在后门。该方法不依赖具体后门类型或标签映射,对对抗训练的后门模型同样有效。即使在无训练数据条件下,也能在多个数据集和场景下实现高精度检测,具备强适应性与鲁棒性。

原文摘要 · Abstract (English)

Scanning for trojan (backdoor) in deep neural networks is crucial due to their significant real-world applications. There has been an increasing focus on developing effective general trojan scanning methods across various trojan attacks. Despite advancements, there remains a shortage of methods that perform effectively without preconceived assumptions about the backdoor attack method. Additionally, we have observed that current methods struggle to identify classifiers trojaned using adversarial training. Motivated by these challenges, our study introduces a novel scanning method named TRODO (TROjan scanning by Detection of adversarial shifts in Out-of-distribution samples). TRODO leverages the concept of "blind spots"--regions where trojaned classifiers erroneously identify out-of-distribution (OOD) samples as in-distribution (ID). We scan for these blind spots by adversarially shifting OOD samples towards in-distribution. The increased likelihood of perturbed OOD samples being classified as ID serves as a signature for trojan detection. TRODO is both trojan and label mapping agnostic, effective even against adversarially trained trojaned classifiers. It is applicable even in scenarios where training data is absent, demonstrating high accuracy and adaptability across various scenarios and datasets, highlighting its potential as a robust trojan scanning strategy.

后门检测对抗样本模型安全

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。