arXiv:2411.12914cs.LGcs.CR2024-11被引 4

利用神经坍缩现象检测并清除模型中的后门攻击

Trojan Cleansing with Neural Collapse

  • 通过神经坍缩特性识别后门攻击导致的特征异常
  • 在多种架构与数据集上验证了清除效果
  • 轻量级方法适用于广泛模型,适合安全审查场景

后门攻击是一种在训练阶段植入的复杂攻击,通过特定触发器使神经网络在任何含触发器的输入上产生指定输出。随着深度网络规模越来越大,且训练数据难以全面审计,这类攻击风险日益突出。本文将后门攻击与神经坍缩(Neural Collapse)现象关联,发现后门攻击会破坏多种数据集和模型架构下的神经坍缩收敛性。基于此,我们设计了一种轻量级、广适应性的清洗机制,可有效清除多种不同架构中的后门攻击,并在实验中验证其有效性。

原文摘要 · Abstract (English)

Trojan attacks are sophisticated training-time attacks on neural networks that embed backdoor triggers which force the network to produce a specific output on any input which includes the trigger. With the increasing relevance of deep networks which are too large to train with personal resources and which are trained on data too large to thoroughly audit, these training-time attacks pose a significant risk. In this work, we connect trojan attacks to Neural Collapse, a phenomenon wherein the final feature representations of over-parameterized neural networks converge to a simple geometric structure. We provide experimental evidence that trojan attacks disrupt this convergence for a variety of datasets and architectures. We then use this disruption to design a lightweight, broadly generalizable mechanism for cleansing trojan attacks from a wide variety of different network architectures and experimentally demonstrate its efficacy.

后门攻击神经坍缩模型安全清洗机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。