发现大模型检索头随生成过程动态变化,揭示其内在规划机制。
Retrieval Heads are Dynamic
- 从动态视角分析生成过程中的检索头行为,突破静态统计局限。
- 动态检索头在关键步骤不可替代,性能显著优于静态头。
- 隐藏状态能预测未来检索模式,暗示模型具备内部规划能力。
近期研究发现大语言模型中的“检索头”负责从输入上下文中提取信息。然而,以往工作多依赖跨数据集的静态统计,仅识别平均表现良好的检索头,忽略了自回归生成过程中的细粒度时序动态。本文从动态视角探究检索头,通过大量分析提出三个核心结论:(1) 动态性:检索头随生成时间步动态变化;(2) 不可替代性:每个时间步的动态检索头具有特定作用,无法被静态检索头有效替代;(3) 相关性:模型隐藏状态编码了未来检索头模式的预测信号,表明存在内部规划机制。我们在针堆任务和多跳问答任务上验证了这些发现,并在动态检索增强生成框架中量化了动态与静态检索头的性能差异。本研究为理解大模型内部机制提供了新视角。
原文摘要 · Abstract (English)
Recent studies have identified "retrieval heads" in Large Language Models (LLMs) responsible for extracting information from input contexts. However, prior works largely rely on static statistics aggregated across datasets, identifying heads that perform retrieval on average. This perspective overlooks the fine-grained temporal dynamics of autoregressive generation. In this paper, we investigate retrieval heads from a dynamic perspective. Through extensive analysis, we establish three core claims: (1) Dynamism: Retrieval heads vary dynamically across timesteps; (2) Irreplaceability: Dynamic retrieval heads are specific at each timestep and cannot be effectively replaced by static retrieval heads; and (3) Correlation: The model's hidden state encodes a predictive signal for future retrieval head patterns, indicating an internal planning mechanism. We validate these findings on the Needle-in-a-Haystack task and a multi-hop QA task, and quantify the differences on the utility of dynamic and static retrieval heads in a Dynamic Retrieval-Augmented Generation framework. Our study provides new insights into the internal mechanisms of LLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。