重新评估微调在无结构知识编辑中的有效性,发现优化后效果超越当前最佳方法。
Is Fine-Tuning an Effective Solution? Reassessing Knowledge Editing for Unstructured Data
- 构建新数据集,系统测试模型编辑后的局部性表现。
- 优化微调设置后,批量编辑性能提升6.78%至10.80%。
- 为未来研究提供可复现的微调训练方案,适合知识更新任务者参考。
无结构知识编辑(UKE)对更新大语言模型的知识至关重要,尤其针对长文本或自由格式的现实知识。现有方法虽有效,但存在两大问题:缺乏对编辑局部性的评估,以及微调(FT)方法在实际中异常失效。为此,我们扩展两个现有数据集,构建了UnKEBench-Loc和AKEW-Loc(CF),引入来自无结构与结构视角的局部性测试数据,实现对编辑后模型局部性的系统评估。进一步分析发现四种影响微调性能的关键因素,并据此设计实验,明确高性能微调方法的最优训练方式,形成可复用的训练范式。实验表明,在最优配置下,微调方法(FT-UKE)表现惊人,超越现有最先进水平;在批量编辑场景中,其优势随批次增大而增强,平均指标提升从+6.78%扩大至+10.80%。
原文摘要 · Abstract (English)
Unstructured Knowledge Editing (UKE) is crucial for updating the relevant knowledge of large language models (LLMs). It focuses on unstructured inputs, such as long or free-form texts, which are common forms of real-world knowledge. Although previous studies have proposed effective methods and tested them, some issues exist: (1) Lack of Locality evaluation for UKE, and (2) Abnormal failure of fine-tuning (FT) based methods for UKE. To address these issues, we first construct two datasets, UnKEBench-Loc and AKEW-Loc (CF), by extending two existing UKE datasets with locality test data from the unstructured and structured views. This enables a systematic evaluation of the Locality of post-edited models. Furthermore, we identify four factors that may affect the performance of FT-based methods. Based on these factors, we conduct experiments to determine how the well-performing FT-based methods should be trained for the UKE task, providing a training recipe for future research. Our experimental results indicate that the FT-based method with the optimal setting (FT-UKE) is surprisingly strong, outperforming the existing state-of-the-art (SOTA). In batch editing scenarios, FT-UKE shows strong performance as well, with its advantage over SOTA methods increasing as the batch size grows, expanding the average metric lead from +6.78% to +10.80%
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。