用稀疏自编码器定位并消除大模型中的有害知识。
Don't Forget It! Conditional Sparse Autoencoder Clamping Works for Unlearning
- 通过稀疏自编码器识别模型内部的有害概念
- 在保留正常功能前提下,使模型无法回答危险问题
- 为可控删除大模型有害知识提供新思路
大型语言模型(LLM)虽具强大能力,但也带来生物武器、高级化学或网络攻击等知识滥用风险。由于其内部机制如黑箱般难以理解,开发者难以控制模型行为。近期,稀疏自编码器(SAEs)被用于解析LLM内部概念表征,实现对隐藏激活的直接调控。本文在Gemma-2-2b模型中,利用SAEs从武器与大规模杀伤性武器代理数据集(WMDP)中识别出有害概念,并通过特征调制降低模型生成有害内容的能力,同时保持其在无害任务上的性能表现。结果表明,基于SAE的显式知识删减技术具备可行性,重燃了该方向的研究希望。
原文摘要 · Abstract (English)
Recent developments in Large Language Model (LLM) capabilities have brought great potential but also posed new risks. For example, LLMs with knowledge of bioweapons, advanced chemistry, or cyberattacks could cause violence if placed in the wrong hands or during malfunctions. Because of their nature as near-black boxes, intuitive interpretation of LLM internals remains an open research question, preventing developers from easily controlling model behavior and capabilities. The use of Sparse Autoencoders (SAEs) has recently emerged as a potential method of unraveling representations of concepts in LLMs internals, and has allowed developers to steer model outputs by directly modifying the hidden activations. In this paper, we use SAEs to identify unwanted concepts from the Weapons of Mass Destruction Proxy (WMDP) dataset within gemma-2-2b internals and use feature steering to reduce the model's ability to answer harmful questions while retaining its performance on harmless queries. Our results bring back optimism to the viability of SAE-based explicit knowledge unlearning techniques.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。