arXiv:2502.15017cs.LGcs.CR2025-02中稿 · AAAI被引 1

用可解释网络分析对抗攻击,发现鲁棒模型特征更分散且不重叠。

Interpreting Adversarial Attacks and Defences using Architectures with Enhanced Interpretability

  • 采用可解释的线性门控网络,仅通过特征网络分析对抗攻击机制。
  • 对抗训练模型的超平面与主成分相似,且远离数据点,更稳定。
  • 可视化显示鲁棒模型跨类别激活路径差异大,不易被攻击破坏。

深度学习中的对抗攻击对模型的完整性与可靠性构成重大威胁。对抗训练是常见的防御方法。本文利用具有更强可解释性的深度线性门控网络(DLGN),分析经过PGD对抗训练的鲁棒模型与标准训练模型的差异。在DLGN中,特征网络作为唯一攻击入口,我们研究其超平面方向、与PCA的关系以及类间子网络重叠情况。结果表明,PGD-AT模型的超平面更接近主成分方向,且离数据点更远。通过路径活动分析,发现PGD-AT模型在不同类别间形成多样且非重叠的活跃子网络,避免攻击引发的门控重叠。可视化揭示了两种模型所学表示的本质差异。

原文摘要 · Abstract (English)

Adversarial attacks in deep learning represent a significant threat to the integrity and reliability of machine learning models. Adversarial training has been a popular defence technique against these adversarial attacks. In this work, we capitalize on a network architecture, namely Deep Linearly Gated Networks (DLGN), which has better interpretation capabilities than regular deep network architectures. Using this architecture, we interpret robust models trained using PGD adversarial training and compare them with standard training. Feature networks in DLGN act as feature extractors, making them the only medium through which an adversary can attack the model. We analyze the feature network of DLGN with fully connected layers with respect to properties like alignment of the hyperplanes, hyperplane relation with PCA, and sub-network overlap among classes and compare these properties between robust and standard models. We also consider this architecture having CNN layers wherein we qualitatively (using visualizations) and quantitatively contrast gating patterns between robust and standard models. We uncover insights into hyperplanes resembling principal components in PGD-AT and STD-TR models, with PGD-AT hyperplanes aligned farther from the data points. We use path activity analysis to show that PGD-AT models create diverse, non-overlapping active subnetworks across classes, preventing attack-induced gating overlaps. Our visualization ideas show the nature of representations learnt by PGD-AT and STD-TR models.

对抗攻击可解释性特征分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。