提出新评估方法DCUE,更真实地衡量大模型删书能力。
Towards Evaluation for Real-World LLM Unlearning
- 基于分布修正思想,用验证集校准关键词置信度
- 通过柯尔莫哥洛夫-斯米尔诺夫检验量化评估结果
- 适合研究大模型数据删除与可信评估的学者
本文分析了现有大模型遗忘评估指标在实际场景中的实用性、精确性和鲁棒性局限。为克服这些缺陷,提出一种新指标DCUE(基于分布修正的遗忘评估)。该方法通过验证集识别核心关键词,并校正其置信度分布偏差,最终采用柯尔莫哥洛夫-斯米尔诺夫检验量化评估结果。实验表明,DCUE有效克服了现有指标的不足,未来可指导更实用、可靠的遗忘算法设计。
原文摘要 · Abstract (English)
This paper analyzes the limitations of existing unlearning evaluation metrics in terms of practicality, exactness, and robustness in real-world LLM unlearning scenarios. To overcome these limitations, we propose a new metric called Distribution Correction-based Unlearning Evaluation (DCUE). It identifies core tokens and corrects distributional biases in their confidence scores using a validation set. The evaluation results are quantified using the Kolmogorov-Smirnov test. Experimental results demonstrate that DCUE overcomes the limitations of existing metrics, which also guides the design of more practical and reliable unlearning algorithms in the future.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。