发现语言模型中专门处理时间信息的注意力头,可精准编辑时间知识。
Does Time Have Its Place? Temporal Heads: Where Language Models Recall Time-specific Information
- 通过电路分析识别出专用于时间知识的注意力头
- 关闭这些头会降低时间相关问答能力但不影响其他任务
- 支持数值与文本形式的时间触发,适合时间敏感应用
尽管语言模型提取事实的能力已广受研究,其对随时间变化的事实的处理仍缺乏探索。我们通过电路分析发现‘时间头’——一类主要处理时间知识的特定注意力头,并确认其在多个模型中存在,位置可能不同,响应随知识类型和对应年份而异。禁用这些头会削弱模型回忆时间特定知识的能力,同时保持其他通用能力及非时间相关问答性能不受影响。此外,这些头不仅响应数值条件(如“2004年”),还对文本别名(如“那一年”)激活,表明其编码的时间维度超越简单数值表示。进一步地,我们展示了通过调整这些头的值来编辑时间知识的可行性。
原文摘要 · Abstract (English)
While the ability of language models to elicit facts has been widely investigated, how they handle temporally changing facts remains underexplored. We discover Temporal Heads, specific attention heads that primarily handle temporal knowledge, through circuit analysis. We confirm that these heads are present across multiple models, though their specific locations may vary, and their responses differ depending on the type of knowledge and its corresponding years. Disabling these heads degrades the model's ability to recall time-specific knowledge while maintaining its general capabilities without compromising time-invariant and question-answering performances. Moreover, the heads are activated not only numeric conditions ("In 2004") but also textual aliases ("In the year ..."), indicating that they encode a temporal dimension beyond simple numerical representation. Furthermore, we expand the potential of our findings by demonstrating how temporal knowledge can be edited by adjusting the values of these heads.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。