arXiv:2503.12683cs.LGmath.GR2025-03被引 2

用代数方法生成可追溯的对抗样本,突破传统攻击不可解释的局限。

Algebraic Adversarial Attacks on Explainability Models

  • 基于神经网络对称性群构造对抗样本,实现数学可追踪的攻击。
  • 在三个数据集上验证有效,包括一个真实世界数据集。
  • 适合研究对抗鲁棒性与可解释性交叉问题的研究者。

经典对抗攻击被表述为约束优化问题,虽有效但难以追溯对抗样本的生成过程。本文提出一种代数方法构建对抗攻击,研究后验可解释模型生成对抗样本的条件。通过几何深度学习框架分析神经网络的对称性群,构造出数学可解析的对抗样本。该方法在两个知名数据集和一个真实世界数据集上得到验证,展示了其有效性与可追溯性。

原文摘要 · Abstract (English)

Classical adversarial attacks are phrased as a constrained optimisation problem. Despite the efficacy of a constrained optimisation approach to adversarial attacks, one cannot trace how an adversarial point was generated. In this work, we propose an algebraic approach to adversarial attacks and study the conditions under which one can generate adversarial examples for post-hoc explainability models. Phrasing neural networks in the framework of geometric deep learning, algebraic adversarial attacks are constructed through analysis of the symmetry groups of neural networks. Algebraic adversarial examples provide a mathematically tractable approach to adversarial examples. We validate our approach of algebraic adversarial examples on two well-known and one real-world dataset.

对抗攻击可解释性代数方法

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。