现有大模型编辑评估方法不靠谱,新协议更精准衡量知识保留能力
Are We Evaluating the Edit Locality of LLM Model Editing Properly?
- 提出新评估框架,解决旧方法对编辑局部性的误判问题
- 新指标与正则强度强相关,能更好区分不同编辑方法效果
- 适用于各类大模型和编辑技术,可灵活调整评估严格度
模型编辑已成为高效更新大语言模型知识的热门范式。核心目标是平衡编辑有效性(成功注入目标知识)与特定性(即编辑局部性,保留非目标知识)。然而,我们发现现有特定性评估协议存在根本缺陷。系统分析了三大问题:概念上的不一致、现有指标与特定性正则化强度弱相关,且灵敏度不足,难以区分不同方法的性能差异。为此,我们提出一种建设性评估协议,消除开放式模型与确定答案假设之间的冲突,避免查询无关的流畅性偏差,并支持评估严格度在近连续空间内平滑调节。在多种大模型、数据集和编辑方法上的实验表明,基于新协议的指标对正则强度变化更敏感,与正则强度高度相关,能更精细地分辨不同方法的知识保留能力。
原文摘要 · Abstract (English)
Model editing has recently emerged as a popular paradigm for efficiently updating knowledge in LLMs. A central desideratum of updating knowledge is to balance editing efficacy, i.e., the successful injection of target knowledge, and specificity (also known as edit locality), i.e., the preservation of existing non-target knowledge. However, we find that existing specificity evaluation protocols are inadequate for this purpose. We systematically elaborated on the three fundamental issues it faces. Beyond the conceptual issues, we further empirically demonstrate that existing specificity metrics are weakly correlated with the strength of specificity regularizers. We also find that current metrics lack sufficient sensitivity, rendering them ineffective at distinguishing the specificity performance of different methods. Finally, we propose a constructive evaluation protocol. Under this protocol, the conflict between open-ended LLMs and the assumption of determined answers is eliminated, query-independent fluency biases are avoided, and the evaluation strictness can be smoothly adjusted within a near-continuous space. Experiments across various LLMs, datasets, and editing methods show that metrics derived from the proposed protocol are more sensitive to changes in the strength of specificity regularizers and exhibit strong correlation with them, enabling more fine-grained discrimination of different methods' knowledge preservation capabilities.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。