发现语言模型生成重复的‘元凶’神经元,揭示其复制前文的机制。
Repetition Neurons: How Do Language Models Produce Repetitions?
- 通过对比生成文本前后激活值,定位重复相关神经元。
- 神经元在重复持续时逐步更强,像在执行重复任务。
- 跨中英日三模型均现相似模式,具通用性。
本文提出重复神经元(repetition neurons),认为它们是导致文本生成中重复问题的关键神经单元。这些神经元在重复持续过程中激活强度逐渐增强,表明模型将重复视为需不断复制先前上下文的任务,类似于上下文学习。通过比较近期预训练语言模型生成文本中重复开始前后的激活值,我们识别出这些重复神经元。在三个英文和一个日文预训练语言模型中分析发现,重复神经元表现出相似的激活模式,显示出跨模型的一致性。
原文摘要 · Abstract (English)
This paper introduces repetition neurons, regarded as skill neurons responsible for the repetition problem in text generation tasks. These neurons are progressively activated more strongly as repetition continues, indicating that they perceive repetition as a task to copy the previous context repeatedly, similar to in-context learning. We identify these repetition neurons by comparing activation values before and after the onset of repetition in texts generated by recent pre-trained language models. We analyze the repetition neurons in three English and one Japanese pre-trained language models and observe similar patterns across them.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。