对抗训练通过放弃非鲁棒特征,减少了特征数量,从而降低超叠加现象。
Why Does Robustness Reduce Superposition?

- 基于特征分类框架,揭示对抗训练如何逐步淘汰非鲁棒特征。
- 实验表明,对抗训练后特征总数减少,导致超叠加程度下降。
- 适合关注模型可解释性与鲁棒性关系的研究者阅读。
对抗样本及其成因的研究仍处于开放状态。机制可解释性,尤其是超叠加现象,为该问题提供了新视角。Gorton & Lewis (2025) 表明对抗样本源于超叠加,并实证发现对抗训练能减少超叠加,但未提供机制解释。本文受 Ilyas 等人 (2019) 特征分类框架启发,提出因果链:对抗训练主动放弃非鲁棒特征,导致需表征的总特征数减少,从而引发超叠加减弱。
原文摘要 · Abstract (English)
The study of adversarial examples and their origins remains an open area of research. Mechanistic interpretability, and superposition in particular, offers new avenues for approaching this problem. Gorton & Lewis (2025) demonstrate that adversarial examples arise from superposition and show empirically that adversarial training reduces superposition, yet provide no mechanistic account of why this occurs. We present an empirical explanation inspired by the feature taxonomy of Ilyas et al. (2019), tracing the following chain of causalities: adversarial training abandons non-robust features, leading to fewer total features to represent, resulting in less superposition.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。