arXiv:2502.15826cs.CLcs.AI2025-02NAACL被引 11

通过删去旧知识提升大模型纠错能力,避免新旧信息冲突。

CoME: An Unlearning-based Approach to Conflict-free Model Editing

  • 用反学习技术精准删除过时信息,减少干扰
  • 在GPT-J和LLaMA-3上编辑准确率显著提升
  • 适合需要高可靠性更新知识的场景

大语言模型常保留预训练中的过时或错误信息,影响可靠性。尽管已有模型编辑方法可在不重新训练的情况下修正错误,但常因新旧知识冲突导致效果下降。本文提出无冲突模型编辑框架CoME,通过选择性删除过时知识,利用反学习缓解知识干扰,使新信息能有效融入而不损害语言特征。在GPT-J和LLaMA-3上使用Counterfact和ZsRE数据集的实验表明,CoME可显著提升编辑准确率与模型可靠性。结果证明,针对性删除旧知识是提升编辑效果与保持生成性能的关键。

原文摘要 · Abstract (English)

Large language models (LLMs) often retain outdated or incorrect information from pre-training, which undermines their reliability. While model editing methods have been developed to address such errors without full re-training, they frequently suffer from knowledge conflicts, where outdated information interferes with new knowledge. In this work, we propose Conflict-free Model Editing (CoME), a novel framework that enhances the accuracy of knowledge updates in LLMs by selectively removing outdated knowledge. CoME leverages unlearning to mitigate knowledge interference, allowing new information to be integrated without compromising relevant linguistic features. Through experiments on GPT-J and LLaMA-3 using Counterfact and ZsRE datasets, we demonstrate that CoME improves both editing accuracy and model reliability when applied to existing editing methods. Our results highlight that the targeted removal of outdated knowledge is crucial for enhancing model editing effectiveness and maintaining the model's generative performance.

模型编辑反学习大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。