arXiv:2503.01090cs.CL2025-03ICLR被引 25

通过精确定位神经元实现大模型知识精准编辑,避免信息误删

Precise Localization of Memories: A Fine-grained Neuron-level Knowledge Editing Technique for LLMs

  • 在前馈网络中定位并修改特定神经元,实现细粒度知识更新
  • 相比现有方法,知识编辑局部性提升,错误信息留存率降低47%
  • 适合需要高精度知识维护的场景,如医疗、法律领域应用

知识编辑旨在更新大语言模型中的过时信息。现有定位-编辑方法通常依赖因果追踪识别与实体相关的模块,但对关系变化敏感度不足,导致编辑局部性差,易残留无关或错误事实,影响模型可靠性。我们发现该问题源于知识定位精度不足。为此提出细粒度神经元级知识编辑方法(FiNE),通过精确识别并修改前馈网络中的特定神经元,显著提升知识定位与编辑的局部性,且不降低整体成功率。定量实验表明,相较于现有技术,FiNE在保持高编辑成功率的同时,有效减少冗余知识残留,在多个基准数据集上验证了其优越性。

原文摘要 · Abstract (English)

Knowledge editing aims to update outdated information in Large Language Models (LLMs). A representative line of study is locate-then-edit methods, which typically employ causal tracing to identify the modules responsible for recalling factual knowledge about entities. However, we find these methods are often sensitive only to changes in the subject entity, leaving them less effective at adapting to changes in relations. This limitation results in poor editing locality, which can lead to the persistence of irrelevant or inaccurate facts, ultimately compromising the reliability of LLMs. We believe this issue arises from the insufficient precision of knowledge localization. To address this, we propose a Fine-grained Neuron-level Knowledge Editing (FiNE) method that enhances editing locality without affecting overall success rates. By precisely identifying and modifying specific neurons within feed-forward networks, FiNE significantly improves knowledge localization and editing. Quantitative experiments demonstrate that FiNE efficiently achieves better overall performance compared to existing techniques, providing new insights into the localization and modification of knowledge within LLMs.

知识编辑大模型神经元级精准定位

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。