用模型编辑技术实现大模型删知识,效果比传统方法更好
Investigating Model Editing for Unlearning in Large Language Models
- 将Rome、Ike等编辑算法改造用于信息删除
- 在特定场景下遗忘效果优于传统方法
- 适合需要精准删去敏感信息的研究者
机器去学习旨在从模型中移除不想要的信息,但许多方法在参数量大的大语言模型上效率低下,或无法完全删除目标信息而不损害应保留的知识。模型编辑算法虽也处理模型内信息修改,但侧重于将输入导向新目标而非彻底删除。本文研究了ROME、IKE和WISE等编辑算法,并为其设计适用于去学习场景的新编辑目标。实验表明,在特定设置下,模型编辑方法在遗忘质量上超过基线去学习方法。然而,与传统去学习技术一样,它们仍难以在不损害整体性能的前提下准确界定需删除信息的范围。
原文摘要 · Abstract (English)
Machine unlearning aims to remove unwanted information from a model, but many methods are inefficient for LLMs with large numbers of parameters or fail to fully remove the intended information without degrading performance on knowledge that should be retained. Model editing algorithms solve a similar problem of changing information in models, but they focus on redirecting inputs to a new target rather than removing that information altogether. In this work, we explore the editing algorithms ROME, IKE, and WISE and design new editing targets for an unlearning setting. Through this investigation, we show that model editing approaches can exceed baseline unlearning methods in terms of quality of forgetting depending on the setting. Like traditional unlearning techniques, they struggle to encapsulate the scope of what is to be unlearned without damage to the overall model performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。