用认知诊断评估大模型删知识,更准更细。
Beyond Single-Value Metrics: Evaluating and Enhancing LLM Unlearning with Cognitive Diagnosis
- 用认知诊断模型细拆大模型的有害知识残留
- 在8种方法、2个基座模型上验证有效
- 适合安全可控研究者和模型审计人员
由于大模型广泛应用及日益突出的伦理与安全问题,已开发出多种大模型遗忘方法以消除有害知识和不当能力。然而,现有评估多依赖问答准确率等单一指标,难以捕捉有害知识的细微残留,导致无法真实评估遗忘效果。为此,我们提出UNCD(基于认知诊断的遗忘评估)框架,利用认知诊断建模实现对大模型遗忘的细粒度评估。专用基准UNCD-Cyber可详细检验危险能力的移除情况。此外,我们提出UNCD-Agent,通过诊断知识残余并生成针对性遗忘数据来优化遗忘过程。在8种遗忘方法和2个基座模型上的大量实验表明,UNCD不仅提升了评估精度,还有效促进了有害能力的清除。
原文摘要 · Abstract (English)
Due to the widespread use of LLMs and the rising critical ethical and safety concerns, LLM unlearning methods have been developed to remove harmful knowledge and undesirable capabilities. In this context, evaluations are mostly based on single-value metrics such as QA accuracy. However, these metrics often fail to capture the nuanced retention of harmful knowledge components, making it difficult to assess the true effectiveness of unlearning. To address this issue, we propose UNCD (UNlearning evaluation via Cognitive Diagnosis), a novel framework that leverages Cognitive Diagnosis Modeling for fine-grained evaluation of LLM unlearning. Our dedicated benchmark, UNCD-Cyber, provides a detailed assessment of the removal of dangerous capabilities. Moreover, we introduce UNCD-Agent, which refines unlearning by diagnosing knowledge remnants and generating targeted unlearning data. Extensive experiments across eight unlearning methods and two base models demonstrate that UNCD not only enhances evaluation but also effectively facilitates the removal of harmful LLM abilities.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。