arXiv:2502.12377cs.CV2025-02中稿 · International Work…被引 2

研究发现,模型对人类视觉的对齐程度影响其抗攻击能力。

Alignment and Adversarial Robustness: Are More Human-Like Models More Secure?

  • 分析114个模型在105个任务中的对齐与抗攻击表现。
  • 纹理和形状选择性对齐能强预测模型鲁棒性。
  • 为构建更安全、符合人眼感知的视觉模型提供新思路。

少量研究表明,与人类视觉更对齐的机器学习模型也表现出更高的对抗鲁棒性,引发疑问:人眼感知是否能提升模型安全性?若普遍成立,将开辟新路径。本文开展大规模实证分析,系统考察表征对齐与对抗鲁棒性的关系。评估了114个跨越不同架构与训练范式的模型,测量其神经与行为对齐、105个基准的任务性能及通过AutoAttack的对抗鲁棒性。结果表明,整体平均对齐与鲁棒性相关性较弱,但特定对齐基准(尤其纹理或形状选择性)可作为对抗鲁棒性的强预测因子。这说明不同形式的对齐在模型鲁棒性中扮演不同角色,推动进一步探索如何利用对齐机制构建更安全、基于感知的视觉模型。

原文摘要 · Abstract (English)

A small but growing body of work has shown that machine learning models which better align with human vision have also exhibited higher robustness to adversarial examples, raising the question: can human-like perception make models more secure? If true generally, such mechanisms would offer new avenues toward robustness. In this work, we conduct a large-scale empirical analysis to systematically investigate the relationship between representational alignment and adversarial robustness. We evaluate 114 models spanning diverse architectures and training paradigms, measuring their neural and behavioral alignment and engineering task performance across 105 benchmarks as well as their adversarial robustness via AutoAttack. Our findings reveal that while average alignment and robustness exhibit a weak overall correlation, specific alignment benchmarks serve as strong predictors of adversarial robustness, particularly those that measure selectivity toward texture or shape. These results suggest that different forms of alignment play distinct roles in model robustness, motivating further investigation into how alignment-driven approaches can be leveraged to build more secure and perceptually-grounded vision models.

对抗鲁棒性模型对齐视觉感知

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。