提出新指标NCCR,评估神经网络抗攻击能力与输入稳定性。
NCCR: to Evaluate the Robustness of Neural Networks and Adversarial Examples
- 通过监测特定神经元输出变化率衡量模型鲁棒性。
- 实验显示该指标能有效区分正常输入与对抗样本。
- 适用于图像识别和语音识别模型的鲁棒性检测。
近年来,神经网络受到广泛关注,随之而来的是安全问题。研究表明,神经网络对经过微小扰动的对抗样本极为敏感,这种扰动人眼难以察觉。尽管已有多种攻击与防御方法,但针对神经网络及其输入鲁棒性的评估研究仍不足。本文提出一种名为神经元覆盖变化率(NCCR)的新指标,用于衡量深度学习模型抵抗攻击的能力及对抗样本的稳定性。NCCR通过监控输入扰动时特定神经元输出的变化程度,变化越小表示模型越鲁棒。在图像识别与说话人识别模型上的实验表明,该指标能有效评估模型或输入的鲁棒性,并可用于检测输入是否为对抗样本,因为对抗样本通常表现出更低的鲁棒性。
原文摘要 · Abstract (English)
Neural networks have received a lot of attention recently, and related security issues have come with it. Many studies have shown that neural networks are vulnerable to adversarial examples that have been artificially perturbed with modification, which is too small to be distinguishable by human perception. Different attacks and defenses have been proposed to solve these problems, but there is little research on evaluating the robustness of neural networks and their inputs. In this work, we propose a metric called the neuron cover change rate (NCCR) to measure the ability of deep learning models to resist attacks and the stability of adversarial examples. NCCR monitors alterations in the output of specifically chosen neurons when the input is perturbed, and networks with a smaller degree of variation are considered to be more robust. The results of the experiment on image recognition and the speaker recognition model show that our metrics can provide a good assessment of the robustness of neural networks or their inputs. It can also be used to detect whether an input is adversarial or not, as adversarial examples are always less robust.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。