发现控制模型冲突响应的关键神经元,揭示大模型如何平衡记忆与上下文。
Context Copying Modulation: The Role of Entropy Neurons in Managing Parametric and Contextual Knowledge Conflicts
- 识别出抑制上下文抄袭的熵神经元,影响模型输出熵但不改变预测排名。
- 移除这些神经元后,模型更倾向于直接复制上下文内容。
- 适用于研究大模型内部机制、矛盾信息处理的学者。
当大型语言模型(LLMs)面对与内部参数知识相冲突的上下文信息时,其行为表现不一致,且尚无公认解释来说明预期结果分布。近期研究在自回归Transformer模型中发现一类被称为熵神经元的神经元,它们对模型输出熵有显著影响,但对预测标记的排名整体影响较小。本文探究了这些神经元是否参与抑制Transformer中的上下文复制行为,以解决上下文与参数化知识之间的冲突。结果显示,熵神经元负责抑制多种大模型中的上下文复制行为,且删除这些神经元会显著改变生成过程。该结果深化了我们对大模型在处理冲突信息时内部动态的理解。
原文摘要 · Abstract (English)
The behavior of Large Language Models (LLMs) when facing contextual information that conflicts with their internal parametric knowledge is inconsistent, with no generally accepted explanation for the expected outcome distribution. Recent work has identified in autoregressive transformer models a class of neurons -- called entropy neurons -- that produce a significant effect on the model output entropy while having an overall moderate impact on the ranking of the predicted tokens. In this paper, we investigate the preliminary claim that these neurons are involved in inhibiting context copying behavior in transformers by looking at their role in resolving conflicts between contextual and parametric information. We show that entropy neurons are responsible for suppressing context copying across a range of LLMs, and that ablating them leads to a significant change in the generation process. These results enhance our understanding of the internal dynamics of LLMs when handling conflicting information.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。