未训练的CNN比反向传播模型更接近人脑视觉皮层表征。
Untrained CNNs Match Backpropagation at V1: A Systematic RSA Comparison of Four Learning Rules Against Human fMRI
- 用四种学习规则训练相同卷积网络,对比其与人脑fMRI数据的相似性。
- 未训练网络在初级视皮层表现优于反向传播,差异达0.044(p<0.001)。
- 研究揭示网络架构比训练方式对表征影响更大,适合神经科学与认知计算研究者。
计算神经科学的核心问题之一是:训练神经网络所用的学习规则是否决定其内部表征与人类视觉皮层的匹配程度。本研究系统比较了四种学习规则(反向传播、反馈对齐、预测编码、脉冲时间依赖可塑性)在相同卷积架构上的表现,使用代表相似性分析(RSA)评估其与人脑fMRI数据的一致性,数据来自THINGS-fMRI数据集(720个刺激,3名受试者),所有模型在224×224分辨率下处理输入,结果取5次随机种子平均值。关键发现:未训练的随机权重模型在V1/V2区域的表现优于反向传播(rho=0.076 vs. 0.034,Δrho=+0.044,p<0.001);在枕叶物体区(LOC),仅反向传播显著优于随机基线(rho=0.012 vs. -0.005,p<0.001);在颞叶物体区(IT),所有五种条件表现趋同(rho=0.008–0.014),无显著差异。部分RSA分析证实上述效应在像素相似性控制后依然成立。
原文摘要 · Abstract (English)
CORRECTION (August 2026): an evaluation-mode defect affected the predictive-coding and STDP conditions of this study; those results should not be used pending re-computation. At V1 and 224px, predictive coding falls from rho = 0.056 to 0.016 and STDP from 0.064 to 0.037, so the claims that STDP leads among trained rules and that PC and STDP lead at V1/V2 are not supported. The random, backpropagation and feedback-alignment conditions, which carry the untrained-versus-trained claim, are unchanged to within 0.0013. The result is however strongly dependent on the evaluation resolution, held fixed at 224px here; see arXiv:2608.12408. See the correction note on page 1; the original abstract below and the body are unchanged from v1. A central question in computational neuroscience is whether the learning rule used to train a neural network determines how well its internal representations align with those of the human visual cortex. We present a systematic comparison of four learning rules (backpropagation (BP), feedback alignment (FA), predictive coding (PC), and spike-timing-dependent plasticity (STDP)) applied to identical convolutional architectures and evaluated against human fMRI data from the THINGS-fMRI dataset (720 stimuli, 3 subjects) using Representational Similarity Analysis (RSA). All models process stimuli at 224 x 224 resolution; results are averaged across 5 random seeds. Crucially, we include an untrained random-weights baseline that reveals the dominant role of architecture. At V1/V2, the untrained baseline exceeds backpropagation (rho = 0.076 vs. rho = 0.034; Delta-rho = +0.044, p < 0.001). At LOC, only BP reliably exceeds the random baseline (rho = 0.012 vs. -0.005, p < 0.001). At IT, all five conditions converge (rho = 0.008-0.014) with no significant pairwise differences among trained rules. Partial RSA confirms all effects survive pixel-similarity control.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。