评估分辨率影响神经网络与大脑相似性比较,小图训练大图评估会误导结果。
Evaluation Resolution Confounds Learning-Rule Comparisons in Model-Brain RSA of Early Visual Cortex
- 用不同分辨率评估网络,发现未训练网络在高分辨率下表现更像大脑
- 224像素评估时,未训练网络与反向传播差距达+0.044,32像素时仅为-0.001
- 该现象在人脑fMRI和猴脑电生理中均成立,提示需报告评估分辨率
表示相似性分析(RSA)常用于比较学习规则生成的卷积网络表征是否接近大脑。由于反馈对齐、预测编码和STDP等生物合理规则难以扩展,相关研究通常在小图像(如32×32 CIFAR)上训练小型网络,再与自然刺激下的大脑反应比较。我们发现,常见结论——未训练或局部训练网络在早期视觉皮层表现优于反向传播——强烈依赖于评估分辨率。当评估分辨率从32像素提升至224像素,未训练网络与反向传播网络在V1的差异从-0.001±0.007扩大至+0.044±0.006,且在六个分辨率上单调增长(n=5种子)。该现象在人类fMRI中成立,在单种子猕猴电生理中方向一致,且适用于224像素训练的ImageNet ResNet-50和Swin-Tiny Transformer。我们检验四种可能机制均无法解释:训练/评估分辨率匹配、低级结构(Gabor、像素)、未训练基线归一化状态、池化特征收敛至全局亮度统计。第五项实验表明:将图像细节限制在训练分辨率但允许池化位置扩大12倍,可消除约90%效应,说明其根源在图像内容维度。控制实验显示:每幅图像仅一个亮度值即可达到rho=0.074,与未训练网络的0.075相当,限定了此任务在V1的分辨能力。唯一跨分辨率稳定的效应出现在枕叶面孔区(LOC)。早期视觉皮层的比较必须控制并报告评估分辨率。
原文摘要 · Abstract (English)
Representational similarity analysis (RSA) is increasingly used to ask which learning rules give convolutional networks brain-like representations. Because biologically plausible rules such as feedback alignment, predictive coding and STDP do not scale, studies that include them train small networks on small images (typically 32x32 CIFAR) and then compare them to brain responses recorded for naturalistic stimuli modeled at far higher resolution. We find that a common result here -- that untrained or locally trained networks rival or beat backpropagation at early visual cortex -- depends strongly on the resolution at which the network is evaluated. The V1 gap between an untrained and a backprop-trained network widens from -0.001 +/- 0.007 at the 32 px training resolution to +0.044 +/- 0.006 at 224 px, growing monotonically across six resolutions (n = 5 seeds). It holds in human fMRI and, directionally, in single-seed macaque electrophysiology, along the training trajectory, and for an ImageNet ResNet-50 and a Swin-Tiny transformer trained at 224 px. We test four candidate mechanisms and none accounts for it: train/eval resolution matching, low-level Gabor and pixel structure, the normalization state of the untrained baseline, and convergence of the pooled descriptor toward a global brightness statistic. A fifth experiment locates it: capping image detail at the training resolution while letting the pooled positions grow 12-fold removes about 90% of the effect, so the dependence lives on the image-content axis. One control result is worth stating separately: a single scalar luminance value per image reaches rho = 0.074 against V1, matching the untrained network's 0.075, bounding what this comparison can resolve at V1 here. The one learning effect that holds across resolution sits at LOC. Comparisons at early visual cortex must control, and report, the evaluation resolution.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。