基于物理约束的攻击方法,可量化神经网络在粒子物理中的分类鲁棒性。
MiniFool -- Physics-Constraint-Aware Minimizer-Based Adversarial Attacks in Deep Neural Networks
- 通过最小化包含χ²统计量和目标分数偏差的代价函数生成对抗样本。
- 在MNIST和CMS数据上验证,错误分类事件的扰动概率更高。
- 适用于评估未标注实验数据的分类可靠性,尤其适合高能物理场景。
本文提出一种名为MiniFool的新算法,用于在粒子与天体物理领域的神经网络分类任务中实施受物理启发的对抗攻击。该算法最初针对冰立方中微子天文台探测天体物理τ中微子的任务开发,但也可扩展至其他科学领域。我们在经典MNIST数据集及大型强子对撞机CMS实验的公开数据上应用该方法。算法基于最小化一个代价函数,该函数结合了基于χ²的检验统计量与目标分数的偏离度。检验统计量根据实验不确定性量化施加于数据的扰动概率。在研究案例中,我们发现被正确和错误分类的事件在分类翻转的概率上存在差异。通过测试分类变化随缩放实验不确定性的攻击参数的变化,可量化网络决策的鲁棒性,并进一步评估未标注实验数据的分类可靠性。
原文摘要 · Abstract (English)
In this paper, we present a new algorithm, MiniFool, that implements physics-inspired adversarial attacks for testing neural network-based classification tasks in particle and astroparticle physics. While we initially developed the algorithm for the search for astrophysical tau neutrinos with the IceCube Neutrino Observatory, we apply it to further data from other science domains, thus demonstrating its general applicability. Here, we apply the algorithm to the well-known MNIST data set and furthermore, to Open Data data from the CMS experiment at the Large Hadron Collider. The algorithm is based on minimizing a cost function that combines a $χ^2$ based test-statistic with the deviation from the desired target score. The test statistic quantifies the probability of the perturbations applied to the data based on the experimental uncertainties. For our studied use cases, we find that the likelihood of a flipped classification differs for both the initially correctly and incorrectly classified events. When testing changes of the classifications as a function of an attack parameter that scales the experimental uncertainties, the robustness of the network decision can be quantified. Furthermore, this allows testing the robustness of the classification of unlabeled experimental data.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。