PGD攻击轨迹分析揭示:梯度对齐等指标不能独立衡量鲁棒性。
Geometry Is Not Robustness: A Trajectory-Level Study of PGD Evaluation

- 通过记录3000个样本的完整攻击轨迹,分析损失变化、梯度对齐与失败步数。
- 不同鲁棒性模型的平均损失和梯度对齐曲线相似,但失败步数分布差异明显。
- 失败步数是更可靠的鲁棒性指示器,需结合上下文解读轨迹诊断结果。
投影梯度下降(PGD)常用于评估对抗鲁棒性,通常仅依赖最终对抗准确率,忽略了攻击过程中的模型行为。近期研究提出轨迹级诊断方法,如损失演化、梯度对齐和失败步数,以深入理解对抗优化动态。然而,这些指标是否可靠反映鲁棒性仍不明确。本文在Fashion-MNIST上对干净训练与对抗训练的卷积神经网络进行轨迹级研究,采用20步严格PGD攻击(随机初始化+多重启)测量鲁棒性,单初始值记录轨迹。每模型记录3000个正确分类样本的完整攻击轨迹,分析各迭代下的损失演化、梯度对齐及失败时间。结果显示,模型间存在清晰鲁棒性层级;但轨迹指标贡献不均:对抗训练模型间平均损失曲线与梯度对齐模式高度相似,却具有显著不同的鲁棒准确率。相反,失败步数分布能更好区分鲁棒性等级,直接体现对对抗扰动的抵抗能力。这表明轨迹诊断描述优化几何,但不能独立衡量鲁棒性。其可解释性取决于鲁棒性区间、攻击强度及多指标联合评估。轨迹分析应作为补充诊断工具,结合上下文理解,而非替代标准鲁棒性测试。
原文摘要 · Abstract (English)
Projected Gradient Descent (PGD) is widely used to evaluate adversarial robustness, typically via final adversarial accuracy, which does not capture model behaviour throughout the attack. Recent work proposes trajectory-level diagnostics, such as loss evolution, gradient alignment, and steps-to-failure, for deeper insight into adversarial optimisation dynamics. However, whether these diagnostics reliably indicate robustness strength remains unclear. We conduct a trajectory-level investigation of PGD attacks on convolutional neural networks trained on Fashion-MNIST. We compare clean-trained and adversarially-trained models across multiple robustness regimes, using rigorous 20-step PGD evaluations with random initialisation and multiple restarts for robustness measurement, and single-initialisation trajectory recording for diagnostics. We record full PGD trajectories across 3000 clean-correct samples per model and analyse loss evolution, gradient alignment, and failure timing across attack iterations. Our results reveal a clear robustness hierarchy across models; however, trajectory metrics do not contribute equally to its identification. Mean loss trajectories and gradient alignment patterns appear quantitatively similar across adversarially-trained models with substantially different robust accuracies. In contrast, steps-to-failure distributions provide a clearer separation of robustness regimes, directly reflecting functional resistance to adversarial perturbation. These findings indicate that trajectory-level diagnostics describe optimisation geometry but do not independently measure adversarial robustness. Their interpretability depends on robustness regime, attack strength, and multi-metric evaluation. Trajectory-level analysis should be a complementary diagnostic tool, interpreted in context, rather than a replacement for standard robustness measurements.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。