arXiv:2605.04443q-bio.NCcs.AI2026-05被引 1

神经对齐提升模型抗攻击能力,但并非因偏向低频信息。

Dissociating spatial frequency reliance from adversarial robustness advantages in neurally guided deep convolutional neural networks

  • 通过神经对齐使模型更依赖人类视觉的中频通道和低频信息。
  • 仅偏倚低频可轻微提升鲁棒性,但效果远不及神经对齐。
  • 真正关键的是整体表征相似性,而非空间频率偏好。

深度卷积神经网络(DCNN)在诸多视觉任务上已接近甚至超越人类,但仍易受微小扰动攻击。近期研究发现,将DCNN表征与人类视觉皮层活动对齐可提升抗攻击鲁棒性,但机制尚不明确。有假设认为这是由于神经对齐使模型远离脆弱的高频细节,转而关注低空间频率(LSF)。然而,人类物体识别实际上高度依赖特定中频带(“人眼通道”),而此前研究中该频段部分保留。本文探究了神经对齐带来的鲁棒性优势是否源于对LSF或人眼通道的偏好。结果表明:神经对齐模型同时增强对LSF和人眼通道的依赖;但直接引导模型聚焦这两个频段时,无论单独或联合,均未提升鲁棒性,甚至导致下降;仅偏倚于LSF虽带来有限改善,但增幅远小于神经对齐模型所引发的空间频率变化;且这些频段偏倚模型与人类神经表征几何的相似性并无显著提升。因此,空间频率依赖的改变更可能是学习类人表征的副产物,而非鲁棒性提升的根本原因,提示未来需探索表征属性的其他维度。

原文摘要 · Abstract (English)

Deep convolutional neural networks (DCNNs) have rivaled humans on many visual tasks, yet they remain vulnerable to near-imperceptible perturbations generated by adversarial attacks. Recent work shows that aligning DCNN representations with human visual cortex activity improves adversarial robustness, but the mechanisms driving this advantage are unclear. One hypothesis suggests that neural alignment confers robustness by biasing models away from brittle high-frequency details and towards the low spatial frequencies (LSF). However, recent work shows that human object recognition critically depends on a narrow, mid-frequency "human channel". Interestingly, this band was partially preserved in prior LSF-focused studies. Here, we investigate whether a spectral bias towards the LSF or the human channel is the primary driver of the adversarial robustness observed in neurally aligned DCNNs. We first show that DCNNs aligned to higher-order regions of the human ventral visual stream systematically increase reliance on both LSF and the human channel. However, directly steering DCNNs towards these bands revealed a clear dissociation. Biasing models towards the human channel, either alone or together with LSF, does not improve robustness and even impairs it. LSF bias produced some robustness gains, but such improvements are modest despite inducing much larger shifts in spatial-frequency reliance than neurally aligned models. Spatial-frequency-biased models overall show little, if any, increase in similarity to human neural representational geometry. Together, our results suggest that altered spatial-frequency reliance is likely an emergent property of learning more human-like representations rather than the primary mechanism by which neural alignment confers adversarial robustness, and motivate the need for future research examining representational properties beyond spatial-frequency profiles.

对抗鲁棒性神经启发视觉表征频域分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。