解决大模型知识编辑时的因果依赖破坏问题,提升编辑效果。
$μ$KE: Matryoshka Unstructured Knowledge Editing of Large Language Models
- 采用套娃式目标函数与自适应损失系数,保护早期记忆更新对后续输出的影响。
- 在四个基准上最高提升12.33%的编辑效果,优于当前最优方法。
- 适用于多种格式的知识编辑,适合需要精准修改模型知识的开发者。
大语言模型虽具强大知识储备,但受限于静态训练数据,常出现幻觉与安全风险。通过定位-编辑范式修改模型内部知识,相比重训练更具成本效益。然而,现有无结构方法(尤其是基于窗口的自回归方法)常破坏早期记忆更新与后期输出之间的因果依赖关系。本文首先理论分析此局限性,提出一种名为 $μ$KE(Matryoshka Unstructured Knowledge Editing)的新记忆更新机制,通过套娃式目标函数与自适应损失系数,有效保留此类依赖。在两个模型、四个基准上的实证评估表明,$μ$KE 相比当前最佳方法,编辑有效性最高提升12.33%,且在多种格式编辑下仍保持鲁棒性,展现出在大语言模型无结构知识编辑中的巨大潜力。
原文摘要 · Abstract (English)
Large language models (LLMs) have emerged as powerful knowledge bases yet are limited by static training data, leading to issues such as hallucinations and safety risks. Editing a model's internal knowledge through the locate-and-edit paradigm has proven a cost-effective alternative to retraining, though current unstructured approaches, especially window-based autoregressive methods, often disrupt the causal dependency between early memory updates and later output tokens. In this work, we first theoretically analyze these limitations and then introduce Matryoshka Unstructured Knowledge Editing ($μ$KE), a novel memory update mechanism that preserves such dependencies via a Matryoshka-style objective and adaptive loss coefficients. Empirical evaluations on two models across four benchmarks demonstrate that $μ$KE improves edit efficacy by up to 12.33% over state-of-the-art methods, and remains robust when applied to diverse formatted edits, underscoring its potential for effective unstructured knowledge editing in LLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。