通过神经网络中的活跃路径检测并清除后门,提升入侵检测模型安全性
Detecting and Eliminating Neural Network Backdoors Through Active Paths with Application to Intrusion Detection
- 基于神经网络中激活的路径识别后门触发特征
- 在入侵检测模型上成功检测并消除注入的后门
- 方法可解释性强,适合安全敏感场景使用
机器学习后门具有在正常输入下表现正常,但当输入包含特定触发器时则按攻击者意图行为的特点。检测此类触发器已被证明极为困难。本文提出一种新颖且可解释的方法,通过神经网络中的活跃路径检测并消除后门触发器。我们在用于入侵检测的机器学习模型上注入后门,实验结果表明该方法具有显著有效性。
原文摘要 · Abstract (English)
Machine learning backdoors have the property that the machine learning model should work as expected on normal inputs, but when the input contains a specific $\textit{trigger}$, it behaves as the attacker desires. Detecting such triggers has been proven to be extremely difficult. In this paper, we present a novel and explainable approach to detect and eliminate such backdoor triggers based on active paths found in neural networks. We present promising experimental evidence of our approach, which involves injecting backdoors into a machine learning model used for intrusion detection.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。