激活操控可精准提取已删除的敏感信息,暴露当前去记忆技术漏洞。
Extracting Unlearned Information from LLMs with Activation Steering
- 通过生成可控激活向量,精确引导模型输出目标信息。
- 在多个数据集上成功恢复通用知识,但难以提取具体私人细节。
- 揭示了现有去记忆方法的严重安全隐患,适合安全与隐私研究者参考。
大型语言模型(LLMs)在大规模预训练中无意记忆了训练数据片段,可能包含敏感或受版权保护的信息。近年来,去记忆技术被用于在训练后移除敏感知识。然而,已有研究表明,恶意攻击者仍可通过多种方式提取已被删除的信息。现有攻击仅能生成候选输出集合,无法精确定位包含目标信息的真实答案。本文提出激活操控方法,实现对已去记忆模型中信息的精确检索。我们引入一种名为匿名激活操控的新方法生成操控向量,并结合简单的词频统计法,在候选集中定位正确答案。在多种去记忆技术与数据集上的评估表明,激活操控可成功恢复通用知识(如广为人知的虚构角色),但在提取特定信息(如非公众人物的细节)方面存在局限。总体结果表明,从已去记忆模型中精确检索信息是可行的,凸显了当前去记忆技术的重大安全隐患。
原文摘要 · Abstract (English)
An unintended consequence of the vast pretraining of Large Language Models (LLMs) is the verbatim memorization of fragments of their training data, which may contain sensitive or copyrighted information. In recent years, unlearning has emerged as a solution to effectively remove sensitive knowledge from models after training. Yet, recent work has shown that supposedly deleted information can still be extracted by malicious actors through various attacks. Still, current attacks retrieve sets of possible candidate generations and are unable to pinpoint the output that contains the actual target information. We propose activation steering as a method for exact information retrieval from unlearned LLMs. We introduce a novel approach to generating steering vectors, named Anonymized Activation Steering. Additionally, we develop a simple word frequency method to pinpoint the correct answer among a set of candidates when retrieving unlearned information. Our evaluation across multiple unlearning techniques and datasets demonstrates that activation steering successfully recovers general knowledge (e.g., widely known fictional characters) while revealing limitations in retrieving specific information (e.g., details about non-public individuals). Overall, our results demonstrate that exact information retrieval from unlearned models is possible, highlighting a severe vulnerability of current unlearning techniques.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。