arXiv:2410.19278cs.LGcs.AI2024-10被引 60

用稀疏自编码器清除语言模型中的生物学知识,效果有限且易引发副作用。

Applying sparse autoencoders to unlearn knowledge in language models

  • 通过可解释的特征激活负向调节实现知识移除。
  • 能有效消除部分生物相关问答,但其他领域影响小。
  • 多特征联合干预副作用大,不如现有微调方法可靠。

我们研究了稀疏自编码器(SAEs)是否可用于从语言模型中移除知识。采用武器扩散代理数据集的生物学子集,在gemma-2b-it和gemma-2-2b-it模型上进行测试。结果表明,可通过个体可解释的生物学相关SAE特征,有效移除部分WMDP-Bio问答内容,对非生物学领域影响极小。但仅零化特征无效,必须进行负向缩放才有效。同时使用多个SAE特征可移除多个主题,但副作用与现有表示误导法相当或更大。当前SAE质量或干预技术仍需提升,才能使基于SAE的知识移除达到微调方法的水平。

原文摘要 · Abstract (English)

We investigate whether sparse autoencoders (SAEs) can be used to remove knowledge from language models. We use the biology subset of the Weapons of Mass Destruction Proxy dataset and test on the gemma-2b-it and gemma-2-2b-it language models. We demonstrate that individual interpretable biology-related SAE features can be used to unlearn a subset of WMDP-Bio questions with minimal side-effects in domains other than biology. Our results suggest that negative scaling of feature activations is necessary and that zero ablating features is ineffective. We find that intervening using multiple SAE features simultaneously can unlearn multiple different topics, but with similar or larger unwanted side-effects than the existing Representation Misdirection for Unlearning technique. Current SAE quality or intervention techniques would need to improve to make SAE-based unlearning comparable to the existing fine-tuning based techniques.

知识移除稀疏编码语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。