揭示大模型如何处理事实与反事实信息的冲突机制
Tracing Facts or just Copies? A critical investigation of the Competitions of Mechanisms in Large Language Models
- 通过注意力头强度分析事实输出比例,研究信息竞争机制
- 发现增强注意力头会抑制正确事实,说明是泛化复制抑制而非选择性压制
- 大模型注意力模式具领域特异性,越大的模型越精细敏感
本文开展可复现性研究,探究大语言模型(LLMs)在面对事实与反事实信息竞争时的处理机制,重点关注注意力头的作用。我们尝试复现并整合Ortu等、Yu等、Merullo等及McDougall等近期三项研究,这些研究利用机制可解释性工具探讨模型所学事实与上下文矛盾信息之间的竞争关系。本研究具体考察注意力头强度与事实输出比例的关系,评估关于注意力头抑制机制的多种假设,并探究此类注意力模式的领域特异性。结果表明,促进事实输出的注意力头并非通过选择性抑制反事实信息实现,而是通过泛化复制抑制机制,因为增强这些头反而会抑制正确的事实。此外,我们发现注意力头行为具有显著领域依赖性,更大模型表现出更专业化、类别敏感的注意力模式。
原文摘要 · Abstract (English)
This paper presents a reproducibility study examining how Large Language Models (LLMs) manage competing factual and counterfactual information, focusing on the role of attention heads in this process. We attempt to reproduce and reconcile findings from three recent studies by Ortu et al., Yu, Merullo, and Pavlick and McDougall et al. that investigate the competition between model-learned facts and contradictory context information through Mechanistic Interpretability tools. Our study specifically examines the relationship between attention head strength and factual output ratios, evaluates competing hypotheses about attention heads' suppression mechanisms, and investigates the domain specificity of these attention patterns. Our findings suggest that attention heads promoting factual output do so via general copy suppression rather than selective counterfactual suppression, as strengthening them can also inhibit correct facts. Additionally, we show that attention head behavior is domain-dependent, with larger models exhibiting more specialized and category-sensitive patterns.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。