通过电路分析揭示大模型知识掩盖机制,助力减少幻觉。
Pierce the Mists, Greet the Sky: Decipher Knowledge Overshadowing via Knowledge Circuit Analysis
- 提出知识电路分析法,追踪模型内部知识激活路径
- 发现注意力动态变化是知识掩盖的根源之一
- 适合研究大模型幻觉与训练机制的学者使用
大语言模型尽管能力强大,但仍受幻觉困扰。其中一种难以处理的类型——知识掩盖,指某一激活知识无意中遮蔽了另一相关知识,导致即使在高质量训练数据下仍产生错误输出。当前对掩盖现象的理解多停留在推理阶段观察,缺乏对其训练过程中成因和内在机制的深入认识。为此,本文提出PhantomCircuit框架,通过创新性的知识电路分析,剖析关键组件功能及注意力模式动态如何促成掩盖现象及其在训练过程中的演变。大量实验验证了该方法在识别此类现象上的有效性,为这一隐蔽幻觉提供了新视角,并为研究社区提供了潜在缓解手段的新方法论工具。
原文摘要 · Abstract (English)
Large Language Models (LLMs), despite their remarkable capabilities, are hampered by hallucinations. A particularly challenging variant, knowledge overshadowing, occurs when one piece of activated knowledge inadvertently masks another relevant piece, leading to erroneous outputs even with high-quality training data. Current understanding of overshadowing is largely confined to inference-time observations, lacking deep insights into its origins and internal mechanisms during model training. Therefore, we introduce PhantomCircuit, a novel framework designed to comprehensively analyze and detect knowledge overshadowing. By innovatively employing knowledge circuit analysis, PhantomCircuit dissects the function of key components in the circuit and how the attention pattern dynamics contribute to the overshadowing phenomenon and its evolution throughout the training process. Extensive experiments demonstrate PhantomCircuit's effectiveness in identifying such instances, offering novel insights into this elusive hallucination and providing the research community with a new methodological lens for its potential mitigation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。