平坦的损失曲面只带来局部鲁棒性,不保证全局抗干扰能力。
When Flatness Does (Not) Guarantee Adversarial Robustness
- 通过推导最后一层前的平坦度公式,建立损失变化与输入空间的关系
- 实验验证:平坦区域易出现高置信度错误的对抗样本
- 揭示了对抗脆弱性的几何本质,适合研究模型鲁棒性的人阅读
尽管神经网络在实践中表现良好,但仍易受微小对抗扰动影响。长期以来有观点认为,损失曲面中平坦的极小值区域能提升鲁棒性。虽然直观合理,但该关系一直缺乏严谨论证。本文通过严格形式化这一关系,发现该直觉仅部分成立:平坦性仅意味着局部鲁棒性,而非全局鲁棒性。我们首先推导出最后一层前相对平坦度的闭式表达,并据此约束输入空间中损失的变化。由此可对整个网络的对抗鲁棒性进行形式化分析。结果表明,要实现超出局部邻域的鲁棒性,损失必须在数据流形外急剧弯曲。我们在多种架构和数据集上验证了理论预测,揭示了决定对抗脆弱性的几何结构,并发现平坦区域常对应高置信度错误的对抗样本。研究挑战了对平坦度的简化理解,提供了其在鲁棒性中作用的精细刻画。
原文摘要 · Abstract (English)
Despite their empirical success, neural networks remain vulnerable to small, adversarial perturbations. A longstanding hypothesis suggests that flat minima, regions of low curvature in the loss landscape, offer increased robustness. While intuitive, this connection has remained largely informal and incomplete. By rigorously formalizing the relationship, we show this intuition is only partially correct: flatness implies local but not global adversarial robustness. To arrive at this result, we first derive a closed-form expression for relative flatness in the penultimate layer, and then show we can use this to constrain the variation of the loss in input space. This allows us to formally analyze the adversarial robustness of the entire network. We then show that to maintain robustness beyond a local neighborhood, the loss needs to curve sharply away from the data manifold. We validate our theoretical predictions empirically across architectures and datasets, uncovering the geometric structure that governs adversarial vulnerability, and linking flatness to model confidence: adversarial examples often lie in large, flat regions where the model is confidently wrong. Our results challenge simplified views of flatness and provide a nuanced understanding of its role in robustness.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。