arXiv:2506.17279cs.CRcs.AI2025-06被引 6

用分步推理攻击暴露大模型删不掉的知识,揭示删除漏洞。

Step-by-Step Reasoning Attack: Revealing 'Erased' Knowledge in Large Language Models

  • 设计分步推理提示,系统性挖掘被隐藏的知识
  • 62.5%攻击提示成功复现被删除的哈利·波特信息
  • 适合关注模型隐私与安全的研究者和开发者

大型语言模型(LLMs)中的知识擦除对遵守数据与人工智能法规、保护用户隐私、减少偏见和误导信息至关重要。现有去学习方法旨在高效且有效地移除特定知识,同时保持整体模型性能,尤其保留应留存的信息。然而,观察发现这些去学习技术往往只是压制知识,使其深藏于表面之下,可通过恰当提示重新获取。本文表明,分步推理可作为恢复这些隐藏信息的后门。我们提出一种基于分步推理的黑盒攻击Sleek,通过三个核心组件:(1)利用大模型生成查询构建的分步推理对抗提示生成策略;(2)有效召回被删除内容的攻击机制;(3)将提示分类为直接、间接和隐含三类,以识别最能利用去学习弱点的查询类型。在四种先进去学习技术及两个主流大模型上的广泛评估显示,现有方法无法确保可靠的知识删除。生成的对抗提示中,62.5%成功从经WHP去学习的Llama模型中恢复了被遗忘的哈利·波特事实,50%暴露了对保留知识的不公平压制。本工作凸显了信息泄露的持续风险,强调需要更稳健的去学习策略。

原文摘要 · Abstract (English)

Knowledge erasure in large language models (LLMs) is important for ensuring compliance with data and AI regulations, safeguarding user privacy, mitigating bias, and misinformation. Existing unlearning methods aim to make the process of knowledge erasure more efficient and effective by removing specific knowledge while preserving overall model performance, especially for retained information. However, it has been observed that the unlearning techniques tend to suppress and leave the knowledge beneath the surface, thus making it retrievable with the right prompts. In this work, we demonstrate that \textit{step-by-step reasoning} can serve as a backdoor to recover this hidden information. We introduce a step-by-step reasoning-based black-box attack, Sleek, that systematically exposes unlearning failures. We employ a structured attack framework with three core components: (1) an adversarial prompt generation strategy leveraging step-by-step reasoning built from LLM-generated queries, (2) an attack mechanism that successfully recalls erased content, and exposes unfair suppression of knowledge intended for retention and (3) a categorization of prompts as direct, indirect, and implied, to identify which query types most effectively exploit unlearning weaknesses. Through extensive evaluations on four state-of-the-art unlearning techniques and two widely used LLMs, we show that existing approaches fail to ensure reliable knowledge removal. Of the generated adversarial prompts, 62.5% successfully retrieved forgotten Harry Potter facts from WHP-unlearned Llama, while 50% exposed unfair suppression of retained knowledge. Our work highlights the persistent risks of information leakage, emphasizing the need for more robust unlearning strategies for erasure.

知识擦除模型安全提示攻击大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。