发现GPT-2模型中普遍存在稳定激活的通用神经元。
Universal Neurons in GPT-2: Emergence, Persistence, and Functional Impact
- 通过五组模型在五次训练节点分析激活相关性,识别出通用神经元。
- 剔除通用神经元后,交叉熵损失显著上升,影响模型预测性能。
- 早期和深层神经元中的通用神经元具有高度稳定性,随训练持续存在。
我们研究了独立训练的GPT-2 Small模型中神经元普遍性现象,考察这些在不同模型间激活始终相关的通用神经元如何在训练过程中出现并演化。通过对五个GPT-2模型在五个检查点上的分析,基于500万标记数据集的激活进行成对相关性计算,识别出通用神经元。消融实验显示,去除通用神经元会导致交叉熵损失显著增加,表明其对模型预测有重要功能影响。此外,我们量化了神经元持久性,发现通用神经元在训练各阶段具有高稳定性,尤其在早期和深层网络中表现突出。结果表明,语言模型训练过程中会涌现出稳定且普遍的表征结构。
原文摘要 · Abstract (English)
We investigate the phenomenon of neuron universality in independently trained GPT-2 Small models, examining these universal neurons-neurons with consistently correlated activations across models-emerge and evolve throughout training. By analyzing five GPT-2 models at five checkpoints, we identify universal neurons through pairwise correlation analysis of activations over a dataset of 5 million tokens. Ablation experiments reveal significant functional impacts of universal neurons on model predictions, measured via cross entropy loss. Additionally, we quantify neuron persistence, demonstrating high stability of universal neurons across training checkpoints, particularly in early and deeper layers. These findings suggest stable and universal representational structures emerge during language model training.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。