发现深度网络输入空间中存在低损失连接路径,揭示对抗样本新特性。
Input Space Mode Connectivity in Deep Neural Networks
- 在输入空间构建插值路径,验证不同图像间存在低损失连通性
- 路径偏离直线度小于0.15,表明连接路径高度线性
- 可解释对抗样本本质,适合研究模型鲁棒性与可解释性
我们将损失景观中的模式连通性概念拓展至深度神经网络的输入空间。该现象最初在参数空间中被研究,描述了通过梯度下降获得的不同解(损失极小点)之间存在低损失路径。本文提供了理论与实证证据,证明其在深层网络输入空间中同样存在,凸显该现象的普遍性。我们观察到具有相似预测结果的不同输入图像通常可被连接,且对于训练好的模型,路径接近线性,仅偏离直线0.15以内。方法采用真实、插值及通过输入优化技术生成的合成输入。我们推测,高维空间中的输入空间模式连通性是一种几何效应,甚至在未训练模型中也存在,可能由渗流理论解释。利用该连通性,我们获得了关于对抗样本的新见解,并展示了其在对抗检测中的潜力。此外,还讨论了其在深度网络可解释性中的应用。
原文摘要 · Abstract (English)
We extend the concept of loss landscape mode connectivity to the input space of deep neural networks. Mode connectivity was originally studied within parameter space, where it describes the existence of low-loss paths between different solutions (loss minimizers) obtained through gradient descent. We present theoretical and empirical evidence of its presence in the input space of deep networks, thereby highlighting the broader nature of the phenomenon. We observe that different input images with similar predictions are generally connected, and for trained models, the path tends to be simple, with only a small deviation from being a linear path. Our methodology utilizes real, interpolated, and synthetic inputs created using the input optimization technique for feature visualization. We conjecture that input space mode connectivity in high-dimensional spaces is a geometric effect that takes place even in untrained models and can be explained through percolation theory. We exploit mode connectivity to obtain new insights about adversarial examples and demonstrate its potential for adversarial detection. Additionally, we discuss applications for the interpretability of deep networks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。