arXiv:2505.17160cs.CLcs.AI2025-05EMNLP被引 6

用对抗性后缀探测被遗忘模型中的隐藏知识。

Harry Potter is Still Here! Probing Knowledge Leakage in Targeted Unlearned Large Language Models via Automated Adversarial Prompting

  • 设计对抗性后缀自动触发模型残留知识
  • 未完全删除的哈利·波特知识仍可被激活
  • 适合评估模型隐私安全与算法可靠性

本文提出LURK(Latent UnleaRned Knowledge)框架,通过对抗性后缀提示探测已遗忘的大语言模型中隐含保留的知识。该方法自动生成针对哈利·波特领域的对抗性提示后缀,以激发模型中残余的特定信息。实验表明,即使被判定为成功遗忘的模型,在针对性对抗条件下仍会泄露个性化知识,暴露出当前遗忘评估标准的关键缺陷。LURK通过间接探测手段,为评估遗忘算法的鲁棒性提供了更严格、更具诊断性的工具。所有代码将公开。

原文摘要 · Abstract (English)

This work presents LURK (Latent UnleaRned Knowledge), a novel framework that probes for hidden retained knowledge in unlearned LLMs through adversarial suffix prompting. LURK automatically generates adversarial prompt suffixes designed to elicit residual knowledge about the Harry Potter domain, a commonly used benchmark for unlearning. Our experiments reveal that even models deemed successfully unlearned can leak idiosyncratic information under targeted adversarial conditions, highlighting critical limitations of current unlearning evaluation standards. By uncovering latent knowledge through indirect probing, LURK offers a more rigorous and diagnostic tool for assessing the robustness of unlearning algorithms. All code will be publicly available.

模型遗忘知识泄露对抗攻击安全评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。