arXiv:2603.03507cs.LGcond-mat.dis-nn2026-03被引 2

神经网络的感知空间比人类概念空间大得多,导致对抗样本普遍存在。

Solving adversarial examples requires solving exponential misalignment

  • 用感知流形分析机器与人类对概念的认知差异
  • 发现机器感知流形维度高出人类数个数量级
  • 揭示对抗鲁棒性难提升的根本原因是维度不匹配

对抗攻击——人类无法察觉但能欺骗神经网络的输入扰动——仍是机器学习中的顽固问题,其成因仍不明。本文定义并分析了神经网络对某一类别概念的感知流形(PM),即网络自信分类为该类的所有输入构成的空间。研究发现,神经网络的感知流形维度比自然的人类概念维度高几个数量级。由于体积随维度呈指数增长,这表明机器与人类之间存在指数级认知错位,有大量输入被机器自信归类但人类无法感知。这一现象为对抗样本的产生提供了自然几何解释:由于网络感知流形占据输入空间极大区域,任意输入都极可能靠近某个类别的感知流形。因此,本文提出对抗鲁棒性需通过降低机器与人类感知流形的维度差距来实现,并预测:鲁棒准确率与到任一感知流形的距离应与感知流形维度呈负相关。我们在18种不同网络上验证了这些预测。关键发现是,即使最鲁棒的网络仍存在指数级错位,只有少数维度接近人类概念的感知流形才表现出与人类感知的对齐。该研究连接了对齐与对抗样本问题,表明高维感知流形是实现对抗鲁棒性的主要障碍。

原文摘要 · Abstract (English)

Adversarial attacks - input perturbations imperceptible to humans that fool neural networks - remain both a persistent failure mode in machine learning, and a phenomenon with mysterious origins. To shed light, we define and analyze a network's perceptual manifold (PM) for a class concept as the space of all inputs confidently assigned to that class by the network. We find, strikingly, that the dimensionalities of neural network PMs are orders of magnitude higher than those of natural human concepts. Since volume typically grows exponentially with dimension, this suggests exponential misalignment between machines and humans, with exponentially many inputs confidently assigned to concepts by machines but not humans. Furthermore, this provides a natural geometric hypothesis for the origin of adversarial examples: because a network's PM fills such a large region of input space, any input will be very close to any class concept's PM. Our hypothesis thus suggests that adversarial robustness cannot be attained without dimensional alignment of machine and human PMs, and therefore makes strong predictions: both robust accuracy and distance to any PM should be negatively correlated with the PM dimension. We confirmed these predictions across 18 different networks of varying robust accuracy. Crucially, we find even the most robust networks are still exponentially misaligned, and only the few PMs whose dimensionality approaches that of human concepts exhibit alignment to human perception. Our results connect the fields of alignment and adversarial examples, and suggest the curse of high dimensionality of machine PMs is a major impediment to adversarial robustness.

对抗样本感知流形模型对齐高维问题

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。