微调会破坏知识编辑效果,影响模型安全与效率。
Can Fine-Tuning Erase Your Edits? On the Fragile Coexistence of Knowledge Editing and Adaptation
- 测试254种配置,发现微调后知识编辑普遍失效
- 仅微调编辑层即可有效清除编辑,性能损失小
- 非编辑层微调反而比全层微调更易破坏编辑
知识编辑(KE)为更新大语言模型提供了轻量级替代方案,而微调仍是适配新领域和任务的主流方法。尽管两者广泛应用,但长期被孤立研究。本文系统量化了254种实验配置下微调对知识编辑的影响。结果表明,多数情况下编辑在微调后显著衰减:在GPT-J上使用AlphaEdit进行zsRE基准测试时,25.27%的成功编辑在微调后失效。进一步发现,仅微调编辑层即可有效移除编辑,且对下游任务性能影响较小;令人意外的是,仅微调非编辑层导致的编辑衰减比全层微调更严重。激活空间分析显示,微调引发的表征变化在幅度和方向上均比知识编辑更强烈、更一致。研究强调需在模型应用全流程中评估知识编辑的有效性。
原文摘要 · Abstract (English)
Knowledge editing (KE) offers a lightweight alternative to retraining for updating large language models (LLMs). Meanwhile, fine-tuning remains the default operation for adapting LLMs to new domains and tasks. Despite their widespread adoption, these two post-training interventions have been studied in isolation, leaving open a crucial question: if we fine-tune an edited model, do the edits survive? This question is motivated by practical objectives: removing covert or malicious edits, and preserving beneficial edits. If fine-tuning impairs edits (Fig.1), current KE methods become less efficient, as a newly fine-tuned model requires re-editing; if edits persist, fine-tuned models risk propagating hidden malicious edits, raising serious safety concerns. To this end, we systematically quantify edit decay after fine-tuning across 254 experimental configurations. Our results show that in general, edits decay substantially after subsequent fine-tuning. AlphaEdit exhibits the greatest decay on the zsRE benchmark when applied to GPT-J, where 25.27% of previously successful edits become unsuccessful after fine-tuning. We further find that fine-tuning only the edited layers is sufficient to effectively remove edits, while incurring only modest degradation in downstream performance. Surprisingly, fine-tuning non-edited layers leads to greater edit decay than all-layer fine-tuning. Besides, our activation space analysis reveals that fine-tuning produces a larger and more coherent representational shift, both in magnitude and direction, than KE. Overall, our study underscores the necessity of evaluating KE within the broader LLM application pipeline.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。