arXiv:2509.23333q-bio.NCcs.CV2025-09被引 1

用微小扰动检测脑模型的编码稳定性,发现鲁棒模型更像真实大脑。

Targeted perturbations reveal brain-like local coding axes in robustified, but not standard, ANN-based brain models

  • 通过对抗性扰动分析模型局部表征几何结构。
  • 标准模型对微扰极敏感,而鲁棒模型响应稳定且可迁移。
  • 鲁棒模型生成的扰动具语义意义,适合指导实验验证。

人工神经网络(ANN)已成为模拟人类视觉系统的主要工具,主要因其在预测神经反应方面的成功。然而,随着众多模型达到相似预测精度,需要更强的评估标准。本文利用小规模对抗性探针,刻画多个高预测力的ANN脑模型的局部表征几何结构。我们报告四项关键发现:第一,大多数现代ANN脑模型出人意料地脆弱,尽管预测得分高,其响应对微小、不可察觉的扰动极为敏感,暴露出不稳定的局部编码方向。第二,模型对对抗性探针的敏感度比预测准确率更能区分候选神经编码模型。第三,标准模型依赖于不跨架构转移的独立局部编码方向。第四,鲁棒化模型产生的对抗性探针引发可泛化的、语义有意义的变化,表明其捕捉到了视觉系统的局部编码维度。综上,局部表征几何提供了更强的脑模型评估标准。我们还为偏好鲁棒模型提供了实证依据,其更稳定的编码轴不仅与神经选择性更一致,还能生成具体可检验的未来实验预测。

原文摘要 · Abstract (English)

Artificial neural networks (ANNs) have become the de facto standard for modeling the human visual system, primarily due to their success in predicting neural responses. However, with many models now achieving similar predictive accuracy, we need a stronger criterion. Here, we use small-scale adversarial probes to characterize the local representational geometry of many highly predictive ANN-based brain models. We report four key findings. First, we show that most contemporary ANN-based brain models are unexpectedly fragile. Despite high prediction scores, their response predictions are highly sensitive to small, imperceptible perturbations, revealing unreliable local coding directions. Second, we demonstrate that a model's sensitivity to adversarial probes can better discriminate between candidate neural encoding models than prediction accuracy alone. Third, we find that standard models rely on distinct local coding directions that do not transfer across model architectures. Finally, we show that adversarial probes from robustified models produce generalizable and semantically meaningful changes, suggesting that they capture the local coding dimensions of the visual system. Together, our work shows that local representational geometry provides a stronger criterion for brain model evaluation. We also provide empirical grounds for favoring robust models, whose more stable coding axes not only align better with neural selectivity but also generate concrete, testable predictions for future experiments.

脑模型对抗攻击表征几何神经编码

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。