发现诱导头毒性是大模型重复生成的元凶,提出抑制策略改善输出多样性。
Induction Head Toxicity Mechanistically Explains Repetition Curse in Large Language Models

- 通过分析注意力机制,揭示诱导头在重复生成中主导输出的机制。
- 诱导头过度主导会排除其他注意力头,导致输出循环重复。
- 提出正则化方法降低诱导头权重,适合优化模型生成质量的研究者。
重复诅咒是大语言模型生成重复或循环标记序列的现象,虽广泛存在,但机制尚不明确。本文研究了诱导头——一种擅长上下文学习的注意力头——在驱动该行为中的作用。我们定义了诱导头的‘毒性’:在重复过程中,其倾向于主导模型输出的逻辑值,从而排斥其他注意力头的贡献。研究结果表明,诱导头是重复诅咒的关键驱动因素,提供了机制性解释,并提出了通过注意力头正则化来减弱其主导性的方法,以促进更丰富、连贯的输出。该发现对大模型的设计与训练具有重要启示。
原文摘要 · Abstract (English)
Repetition curse is a phenomenon where Large Language Models (LLMs) generate repetitive sequences of tokens or cyclic sequences. While the repetition curse has been widely observed, its underlying mechanisms remain poorly understood. In this work, we investigate the role of induction heads--a specific type of attention head known for their ability to perform in-context learning--in driving this repetitive behavior. Specifically, we focus on the "toxicity" of induction heads, which we define as their tendency to dominate the model's output logits during repetition, effectively excluding other attention heads from contributing to the generation process. Our findings have important implications for the design and training of LLMs. By identifying induction heads as a key driver of the repetition curse, we provide a mechanistic explanation for this phenomenon and suggest potential avenues for mitigation. We also propose a technique with attention head regularization that could be employed to reduce the dominance of induction heads during generation, thereby promoting more diverse and coherent outputs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。