用嵌入空间扰动拓展模型知识评估范围,提升编辑后知识保留能力。
Beyond Local Edits: Embedding-Virtualized Knowledge for Broader Evaluation and Preservation of Model Editing
- 通过嵌入空间扰动虚拟化知识区域,扩展评估范围。
- 发现传统评测忽略的知识漂移现象,平均漂移率提升18.7%。
- 可插拔模块有效抑制知识漂移,适合模型编辑与长期维护场景。
大型语言模型的知识编辑方法通常依赖预定义基准测试,仅评估特定事实及其有限关联知识。此类评估受限于数据集样本,难以全面理解编辑对模型知识系统的影响。为此,我们提出嵌入虚拟化知识(EVK),通过嵌入空间中的受控扰动表征模型知识,探索远超显式标注范围的虚拟知识区域。基于EVK,构建了嵌入级评估基准EVK-Bench,量化编辑引发的知识漂移,揭示传统样本指标未捕捉的效应。进一步提出可插拔的EVK-Align模块,在不牺牲编辑精度的前提下,有效约束嵌入级知识漂移。实验表明,该方法实现更全面评估,并显著提升知识保留能力。
原文摘要 · Abstract (English)
Knowledge editing methods for large language models are commonly evaluated using predefined benchmarks that assess edited facts together with a limited set of related or neighboring knowledge. While effective, such evaluations remain confined to finite, dataset-bounded samples, leaving the broader impact of editing on the model's knowledge system insufficiently understood. To address this gap, we introduce Embedding-Virtualized Knowledge (EVK) that characterizes model knowledge through controlled perturbations in embedding space, enabling the exploration of a substantially broader and virtualized knowledge region beyond explicit data annotations. Based on EVK, we construct an embedding-level evaluation benchmark EVK-Bench that quantifies potential knowledge drift induced by editing, revealing effects that are not captured by conventional sample-based metrics. Furthermore, we propose a plug-and-play EVK-Align module that constrains embedding-level knowledge drift during editing and can be seamlessly integrated into existing editing methods. Experiments demonstrate that our approach enables more comprehensive evaluation while significantly improving knowledge preservation without sacrificing editing accuracy.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。