arXiv:2504.08192cs.LGcs.AI2025-04被引 24

动态稀疏自编码器让大模型精准删知识,效果远超传统方法。

SAEs $\textit{Can}$ Improve Unlearning: Dynamic Sparse Autoencoder Guardrails for Precision Unlearning in LLMs

  • 用动态筛选特征的稀疏自编码器,精准定位需删除的知识节点。
  • 在序列化删知识和零样本场景下,效率与稳定性显著优于现有方法。
  • 适合关注模型安全、可解释性及抗反学习攻击的开发者使用。

机器去学习是一种提升大语言模型安全性的有效途径,通过移除模型中不希望保留的知识。然而,现有的基于梯度的去学习方法存在计算成本高、超参数不稳定、难以支持序列化去学习、易受重学攻击、数据效率低以及缺乏可解释性等问题。稀疏自编码器(SAEs)因其能实现基于激活的靶向去学习,具备改进上述问题的潜力,但此前研究显示其性能仍不及基于梯度的方法。本文证明,若采用动态策略,SAEs 可显著提升去学习效果。我们提出动态稀疏自编码器防护机制(Dynamic DAE Guardrails, DSG),一种基于合理特征选择与动态分类器的精确去学习新方法。实验表明,DSG 在多项指标上显著超越主流去学习方法,实现了更优的遗忘-效用权衡。该方法克服了梯度法在计算效率、稳定性、序列去学习能力、抗重学攻击性、数据效率(包括零样本场景)以及可解释性方面的缺陷。

原文摘要 · Abstract (English)

Machine unlearning is a promising approach to improve LLM safety by removing unwanted knowledge from the model. However, prevailing gradient-based unlearning methods suffer from issues such as high computational costs, hyperparameter instability, poor sequential unlearning capability, vulnerability to relearning attacks, low data efficiency, and lack of interpretability. While Sparse Autoencoders are well-suited to improve these aspects by enabling targeted activation-based unlearning, prior approaches underperform gradient-based methods. This work demonstrates that, contrary to these earlier findings, SAEs can significantly improve unlearning when employed dynamically. We introduce $\textbf{Dynamic DAE Guardrails}$ (DSG), a novel method for precision unlearning that leverages principled feature selection and a dynamic classifier. Our experiments show DSG substantially outperforms leading unlearning methods, achieving superior forget-utility trade-offs. DSG addresses key drawbacks of gradient-based approaches for unlearning -- offering enhanced computational efficiency and stability, robust performance in sequential unlearning, stronger resistance to relearning attacks, better data efficiency including zero-shot settings, and more interpretable unlearning.

模型安全去学习稀疏编码可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。