提出概念级遗忘评估基准,测试大模型在有害与良性场景中精准删除知识的能力。
ConceptGuard: Benchmarking Context-Sensitive Unlearning in Large Language Models

- 基于双用途概念构建新基准,实现上下文敏感的遗忘评估。
- 现有方法在概念级控制上表现差,遗忘与可用性矛盾明显。
- 适合关注模型安全、可控性的研究人员和开发者使用。
大型语言模型日益需要选择性移除有害或敏感知识,即遗忘能力。然而,现有方法和基准未能全面评估此能力。当前方法依赖独立的事实集进行遗忘与保留测试,并通过简单直接的事实召回衡量效果,无法捕捉遗忘的核心要求:消除有害行为的同时保留有益知识。我们主张有效遗忘应作用于概念层面,确保危险用法被彻底清除,而正确且有用的用法得以保留,从而实现概念上有意义的完整遗忘。为此,我们引入双用途概念——既可用于有害也可用于良性场景的概念。基于这些概念,构建名为ConceptGuard的基准,其遗忘与保留集合在概念使用上显式互补。该基准首次支持以概念为单位评估遗忘,评价方式注重意图敏感性,目标是最大化上下文分离以促进更安全行为。实验表明,现有遗忘技术在此设置下表现不佳,表现出弱上下文分离、低ROUGE分数及概念级指标差。结果揭示强烈的遗忘-效用权衡、有限的上下文敏感性提升以及不同方法间概念控制的一致性差,为更符合现实安全需求的遗忘方法提供思路。数据集已公开。
原文摘要 · Abstract (English)
Large Language Models (LLMs) increasingly require selective removal of harmful or sensitive knowledge, called unlearning, yet existing methods and benchmarks fail to evaluate this capability completely. Current approaches rely on disjoint forget and retain sets composed of independent facts, and measure success using simple and direct factual recall. This framing fails to capture a key requirement of unlearning, namely the ability to eliminate harmful behaviors while preserving benign and beneficial knowledge. We argue that effective unlearning must operate at the level of concepts, ensuring complete removal of unsafe applications while maintaining their correct and useful usage, thereby achieving conceptually meaningful and complete unlearning. To better evaluate unlearning techniques from such a practical viewpoint, we introduce the notion of dual-use concepts: concepts that can be used in both harmful and benign contexts. Building on these concepts, we construct a benchmark called ConceptGuard where forget and retain sets are explicitly complementary in concept usage. Our benchmark uniquely enables unlearning to be explored and gauged at the level of concepts, instead of sparse facts, and evaluation is intent-sensitive with the goal of maximizing contextual separation to promote safer behavior. We demonstrate that current unlearning techniques perform poorly under this setting, showing weak contextual separation alongside poor performance in ROUGE and concept-level metrics. Our results reveal strong forgetting-utility trade-offs, limited gains in contextual sensitivity, and poor consistency in concept-level control across methods, and provide ideas for unlearning approaches that better align with real-world safety requirements. Our dataset is publicly available.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。