大模型越做越大,关键神经元反而越来越专一。
Neuron Populations Exhibit Divergent Selectivity with Scale

- 发现共享神经元随规模增长呈亚线性增长,占比下降。
- 神经元愈发专一,非共享神经元则保持泛化特性。
- 适合关注模型可解释性与神经元演化规律的研究者。
我们研究了神经网络中神经元群体随模型规模变化是否具有可预测性,将缩放定律扩展至损失等宏观可观测量之外。通过分析语言模型(最大300亿参数)和视觉模型(最大50亿参数)中的罗塞塔神经元(Rosetta Neurons),发现其绝对数量虽随规模增长,但占总神经元比例持续下降,符合亚线性幂律。同时观察到神经元极化效应:罗塞塔神经元的特异性不断增强,趋于单义性,而不断扩大的非罗塞塔神经元群体仍保持较低选择性。一个平衡特征效用与神经元容量限制的解析模型可解释该幂律及极化现象。此外,罗塞塔神经元随规模增强领域专业化,通过针对性数据过滤的继续预训练案例验证其选择性。结果揭示了可解释、共享的神经元层级结构存在缩放规律,连接模型规模与神经元通用性、选择性和专业化程度的系统性变化。
原文摘要 · Abstract (English)
We investigate whether neuron populations within neural networks evolve predictably with scale, extending scaling laws beyond macroscopic observables such as loss. To probe this question, we study Rosetta Neurons, a previously characterized class of neurons whose activation patterns are similar across independently trained models (Dravid et al., 2023). In separate analyses of language models up to 30B parameters and vision models up to 5B parameters, we observe that the population of Rosetta Neurons follows a sublinear power law in model size, growing in absolute number but occupying a shrinking fraction of the total neuron count. We further observe a Neuron Polarization Effect: Rosetta Neurons become more selective and increasingly monosemantic with scale, separating from a growing non-Rosetta population that remains less selective. An analytical model balancing feature utility against limited neuron capacity explains the sublinear power-law scaling and this polarization effect. Finally, we find that Rosetta Neurons become more domain-specialized with scale and illustrate their selectivity through a targeted data-filtering case study for continued pretraining. Our results point to a scaling law for interpretable, shared neuron-level structure, linking model size to systematic changes in neuron universality, selectivity, and specialization.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。