通过插入像素揭示了模型准确率与鲁棒性之间的几何矛盾
A Geometric Probe of the Accuracy-Robustness Trade-off: Sharp Boundaries in Symmetry-Breaking Dimensional Expansion
- 用插入常数像素打破对称性,提升准确率
- 准确率从90.47%升至95.63%,但对抗鲁棒性下降
- 修复插入维度可恢复鲁棒性,揭示脆弱边界机制
深度学习中清洁准确率与对抗鲁棒性之间的权衡现象普遍存在,但其几何根源仍不明确。本文利用对称性破缺维度扩展(SBDE)作为可控探针,研究该权衡的机制。SBDE通过插入常数像素扩展输入图像,破坏平移对称性,显著提升清洁准确率(如在CIFAR-10上,ResNet-18从90.47%提升至95.63%),其原理是减少参数退化。然而,这一准确率提升以牺牲对抗鲁棒性为代价。通过测试时的掩码投影(将插入的辅助像素重置为训练值),我们发现攻击脆弱性几乎完全源自插入维度。投影能有效消除攻击并恢复鲁棒性,表明模型通过在辅助轴上构建陡峭的损失梯度(尖锐边界)实现高准确率。研究揭示:优化景观加深了吸引盆以提高准确率,却不可避免地在辅助自由度上形成陡峭屏障,导致对流形外扰动极度敏感。
原文摘要 · Abstract (English)
The trade-off between clean accuracy and adversarial robustness is a pervasive phenomenon in deep learning, yet its geometric origin remains elusive. In this work, we utilize Symmetry-Breaking Dimensional Expansion (SBDE) as a controlled probe to investigate the mechanism underlying this trade-off. SBDE expands input images by inserting constant-valued pixels, which breaks translational symmetry and consistently improves clean accuracy (e.g., from $90.47\%$ to $95.63\%$ on CIFAR-10 with ResNet-18) by reducing parameter degeneracy. However, this accuracy gain comes at the cost of reduced robustness against iterative white-box attacks. By employing a test-time \emph{mask projection} that resets the inserted auxiliary pixels to their training values, we demonstrate that the vulnerability stems almost entirely from the inserted dimensions. The projection effectively neutralizes the attacks and restores robustness, revealing that the model achieves high accuracy by creating \emph{sharp boundaries} (steep loss gradients) specifically along the auxiliary axes. Our findings provide a concrete geometric explanation for the accuracy-robustness paradox: the optimization landscape deepens the basin of attraction to improve accuracy but inevitably erects steep walls along the auxiliary degrees of freedom, creating a fragile sensitivity to off-manifold perturbations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。