发现语言模型推理与记忆的切换由单一方向控制。
The Reasoning-Memorization Interplay in Language Models Is Mediated by a Single Direction
- 在残差流中找到控制推理与记忆平衡的线性特征。
- 操纵这些特征可显著提升模型推理表现。
- 适合研究模型内部机制或提升生成可靠性的人参考。
大型语言模型(LLMs)在多种推理基准上表现优异,但先前研究指出其在面对未见问题时可能因过度依赖训练数据中的记忆而难以泛化。然而,模型在文本生成过程中何时从推理转向记忆仍不明确。本文通过分析模型残差流中的线性特征,揭示了控制推理与记忆平衡的机制。这些特征不仅能区分推理任务与记忆密集型任务,还可被干预以因果方式影响模型在推理任务上的表现。此外,干预这些推理特征有助于模型更准确地激活相关解题能力。研究为理解大模型中的推理与记忆机制提供了新视角,并为构建更稳健、可解释的生成式AI系统铺平道路。
原文摘要 · Abstract (English)
Large language models (LLMs) excel on a variety of reasoning benchmarks, but previous studies suggest they sometimes struggle to generalize to unseen questions, potentially due to over-reliance on memorized training examples. However, the precise conditions under which LLMs switch between reasoning and memorization during text generation remain unclear. In this work, we provide a mechanistic understanding of LLMs' reasoning-memorization dynamics by identifying a set of linear features in the model's residual stream that govern the balance between genuine reasoning and memory recall. These features not only distinguish reasoning tasks from memory-intensive ones but can also be manipulated to causally influence model performance on reasoning tasks. Additionally, we show that intervening in these reasoning features helps the model more accurately activate the most relevant problem-solving capabilities during answer generation. Our findings offer new insights into the underlying mechanisms of reasoning and memory in LLMs and pave the way for the development of more robust and interpretable generative AI systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。