参数空间分类器比传统模型更抗对抗攻击,无需额外训练。
Adversarial Attacks in Weight-Space Classifiers
- 在参数空间直接进行分类,利用隐式神经表示的特性提升鲁棒性。
- 实验表明参数空间模型对白盒攻击的鲁棒性显著优于传统分类器。
- 揭示梯度混淆是原因,提出新攻击方法验证其局限性,适合安全研究者。
隐式神经表示(INRs)近年来在多个研究领域受到关注,因其能以紧凑、连续的方式表征大规模复杂数据。已有研究表明,许多主流下游任务可直接在INR参数空间中完成,大幅降低处理原始数据域时的计算资源消耗。然而,现代机器学习方法普遍易受对抗攻击影响,严重限制了其在各类场景中的可靠性与应用。本文深入分析了参数空间分类器在对抗攻击下的行为。结果表明,直接在参数空间训练的分类模型相比运行于原始信号空间的标准分类器,对标准白盒攻击具有更强的鲁棒性,且无需任何鲁棒性训练。这种鲁棒性源于INR优化过程中产生的梯度混淆现象。同时,我们发现该鲁棒性在某些替代攻击方法下存在局限。为支持结论,我们构建了一套新型针对参数空间分类器的对抗攻击方法,并进一步分析了实际攻击中的关键考量。
原文摘要 · Abstract (English)
Implicit Neural Representations (INRs) have been recently garnering increasing interest in various research fields, mainly due to their ability to represent large, complex data in a compact, continuous manner. Past work further showed that numerous popular downstream tasks can be performed directly in the INR parameter-space. Doing so can substantially reduce the computational resources required to process the represented data in their native domain. A major difficulty in using modern machine-learning approaches, is their high susceptibility to adversarial attacks, which have been shown to greatly limit the reliability and applicability of such methods in a wide range of settings. In this work, we perform an in-depth security analysis of the behavior of weight-space classifiers under adversarial attacks. Our study reveals that parameter-space models trained for classification exhibit increased robustness to standard white-box adversarial attacks compared to standard classifiers that operate in the original signal space. This is achieved without the need of any robust training. We source this robust behavior to the phenomenon of gradient-obfuscation promoted during the INR optimization process, and pinpoint the limitations of this robustness under alternative adversarial approaches. To support our claims, we develop a novel suite of adversarial attacks targeting parameter-space classifiers, and furthermore analyze practical considerations of such attacks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。