arXiv:2510.19427cs.LG2025-10

发现鲁棒模型并非普遍相似,分类层是差异根源。

Revisiting the Relation Between Robustness and Universality

  • 测试不同数据集下鲁棒模型的表征相似性
  • 预测行为随鲁棒性提升不收敛,无严格普适性
  • 简单重训分类器可提升预测一致性,适合模型分析者

Jones 等人(2022)提出的修正普遍性假说认为,针对特定任务训练的对抗鲁棒模型高度相似。本文重新审视该假说并检验其普适性。虽然在特定设置下验证了表征相似性的主要结论,但在不同数据集上结果不一致。此外,我们发现预测行为不会随鲁棒性增强而收敛,因此不具备普遍性。差异源于分类层,但通过简单的分类器重训可实现更一致的预测行为。总体而言,本工作支持神经网络在特定场景下的部分普遍性,否定严格普遍性的观点。

原文摘要 · Abstract (English)

The modified universality hypothesis proposed by Jones et al. (2022) suggests that adversarially robust models trained for a given task are highly similar. We revisit the hypothesis and test its generality. While we verify Jones' main claim of high representational similarity in specific settings, results are not consistent across different datasets. We also discover that predictive behavior does not converge with increasing robustness and thus is not universal. We find that differing predictions originate in the classification layer, but show that more universal predictive behavior can be achieved with simple retraining of the classifiers. Overall, our work points towards partial universality of neural networks in specific settings and away from notions of strict universality.

鲁棒性模型相似性分类器设计

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。