arXiv:2504.01100cs.CLcs.LG2025-04被引 10

不同机制导致语言模型重复,自然与刻意复制行为本质不同。

Repetitions are not all alike: distinct mechanisms sustain repetition in language models

  • 对比自然提示与上下文学习诱导的重复,发现两种机制不同
  • 上下文学习重复由专用注意力头逐步形成,自然重复早期出现且无固定结构
  • 自然重复聚焦低信息词,是上下文无法召回时的退化行为

大型语言模型(LLMs)有时会陷入重复循环,持续生成相同词序列。由于自然人类语言中极少出现重复,而此类现象在多种任务和场景中频繁出现,令人困惑。本文探究行为上相似的重复模式是否源于不同内在机制,以及这些机制如何随训练发展。我们对比了自然文本提示引发的重复与显式要求复制行为的上下文学习(ICL)设置所诱发的重复。分析表明,ICL诱导的重复依赖于一个随训练逐步专门化的注意力头网络,而自然发生的重复则早期出现且缺乏明确电路结构。注意力检视显示,自然重复集中于低信息词,暗示在无法检索相关上下文时的退化行为。结果表明,表面相似的重复行为源自质性不同的内部过程,反映语言模型中不同类型的失败与适应模式。

原文摘要 · Abstract (English)

Large Language Models (LLMs) can sometimes degrade into repetitive loops, persistently generating identical word sequences. Because repetition is rare in natural human language, its frequent occurrence across diverse tasks and contexts in LLMs remains puzzling. Here we investigate whether behaviorally similar repetition patterns arise from distinct underlying mechanisms and how these mechanisms develop during model training. We contrast two conditions: repetitions elicited by natural text prompts with those induced by in-context learning (ICL) setups that explicitly require copying behavior. Our analyses reveal that ICL-induced repetition relies on a dedicated network of attention heads that progressively specialize over training, whereas naturally occurring repetition emerges early and lacks a defined circuitry. Attention inspection further shows that natural repetition focuses disproportionately on low-information tokens, suggesting a fallback behavior when relevant context cannot be retrieved. These results indicate that superficially similar repetition behaviors originate from qualitatively different internal processes, reflecting distinct modes of failure and adaptation in language models.

语言模型重复机制注意力分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。