通过神经元级归因定位关键推理路径,提升大模型多跳事实召回的编辑效果。
ACE: Attribution-Controlled Knowledge Editing for Multi-hop Factual Recall
- 基于神经元归因识别推理链中的查询-值激活路径,精准定位需编辑节点。
- 在GPT-J上提升9.44%,在Qwen3-8B上提升37.46%,显著优于现有方法。
- 适合关注大模型知识更新机制与可解释性研究的研究者。
大型语言模型需要高效的知识编辑(KE)来更新事实信息,但现有方法在多跳事实召回任务中表现显著下降,尤其当编辑涉及推理链中的隐含中间主体时更为严重。通过因果分析,我们发现这一局限源于对链式知识在神经元层面动态表征与利用的忽视。研究发现,在多跳推理过程中,隐含主体充当查询神经元,依次激活各变换器层中的对应值神经元以累积信息至最终答案,而这一动态先验被以往的KE工作所忽略。受此启发,我们提出ACE:一种用于多跳事实召回的归因控制知识编辑框架,利用神经元级归因识别并编辑这些关键的查询-值(Q-V)路径。ACE提供了基于机制理解的多跳知识编辑解决方案,在GPT-J上比当前最优方法提升9.44%,在Qwen3-8B上提升37.46%。进一步分析揭示了Qwen3中更精细的激活模式,并表明值神经元的语义可解释性由查询驱动的信息积累所协调。这些发现为基于内部推理机制原理推进知识编辑能力开辟了新路径。
原文摘要 · Abstract (English)
Large Language Models (LLMs) require efficient knowledge editing (KE) to update factual information, yet existing methods exhibit significant performance decay in multi-hop factual recall. This failure is particularly acute when edits involve intermediate implicit subjects within reasoning chains. Through causal analysis, we reveal that this limitation stems from an oversight of how chained knowledge is dynamically represented and utilized at the neuron level. We discover that during multi hop reasoning, implicit subjects function as query neurons, which sequentially activate corresponding value neurons across transformer layers to accumulate information toward the final answer, a dynamic prior KE work has overlooked. Guided by this insight, we propose ACE: Attribution-Controlled Knowledge Editing for Multi-hop Factual Recall, a framework that leverages neuron-level attribution to identify and edit these critical query-value (Q-V) pathways. ACE provides a mechanistically grounded solution for multi-hop KE, empirically outperforming state-of-the-art methods by 9.44% on GPT-J and 37.46% on Qwen3-8B. Our analysis further reveals more fine-grained activation patterns in Qwen3 and demonstrates that the semantic interpretability of value neurons is orchestrated by query-driven accumulation. These findings establish a new pathway for advancing KE capabilities based on the principled understanding of internal reasoning mechanisms.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。