arXiv:2508.13650cs.CL2025-08ACL被引 9

用稀疏自编码器永久删除大模型有害知识,防篡改且不伤性能。

CRISP: Persistent Concept Unlearning via Sparse Autoencoders

  • 通过多层稀疏自编码器识别关键特征并抑制激活。
  • 在WMDP基准上成功移除有害知识,保留通用与领域能力。
  • 实现语义清晰的特征分离,适合安全敏感场景应用。

随着大语言模型在真实场景中广泛应用,如何选择性移除不需要的知识同时保持模型可用性变得至关重要。现有方法虽利用稀疏自编码器(SAEs)对单一语义特征进行精准干预,但多数仅在推理时生效,无法持久改变模型参数,易被拥有参数访问权限的恶意方绕过或逆转。本文提出CRISP,一种基于SAE的参数高效持久概念去学习方法。CRISP自动识别跨多层的关键SAE特征并抑制其激活。我们在两种LLM上进行实验,结果表明该方法在WMDP基准的安全关键去学习任务中优于先前方法,成功移除有害知识的同时保留通用与领域内能力。特征级分析显示,CRISP实现了目标概念与良性概念间的语义一致分离,可精准抑制目标特征。

原文摘要 · Abstract (English)

As large language models (LLMs) are increasingly deployed in real-world applications, the need to selectively remove unwanted knowledge while preserving model utility has become paramount. Recent work has explored sparse autoencoders (SAEs) to perform precise interventions on monosemantic features. However, most SAE-based methods operate at inference time, which does not create persistent changes in the model's parameters. Such interventions can be bypassed or reversed by malicious actors with parameter access. We introduce CRISP, a parameter-efficient method for persistent concept unlearning using SAEs. CRISP automatically identifies salient SAE features across multiple layers and suppresses their activations. We experiment with two LLMs and show that our method outperforms prior approaches on safety-critical unlearning tasks from the WMDP benchmark, successfully removing harmful knowledge while preserving general and in-domain capabilities. Feature-level analysis reveals that CRISP achieves semantically coherent separation between target and benign concepts, allowing precise suppression of the target features.

概念去学习稀疏自编码器大模型安全

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。