测试发现部分去记忆方法在提示攻击下失效,知识未真正删除。
Prompt Attacks Reveal Superficial Knowledge Removal in Unlearning Methods
- 用提示词攻击测试八种去记忆方法的漏洞
- 特定提示可恢复57.3%的被删知识准确率
- 适合关注模型安全与真实去记忆的研究者
本文揭示某些机器去记忆方法在简单提示攻击下可能失效。我们系统评估了三种模型族中的八种去记忆技术,通过输出、逻辑值和探测分析,检验所谓已删除的知识能否被重新获取。尽管RMU和TAR表现出强鲁棒性,但ELM对特定提示攻击仍脆弱(如在原提示前加印地语填充文本,可恢复57.3%准确率)。逻辑值分析进一步表明,去记忆模型不太可能通过改变输出格式隐藏知识,因输出与逻辑值准确率存在强相关性。研究挑战了当前对去记忆有效性的普遍假设,强调需建立可靠评估框架以区分真实知识删除与表面输出抑制。为促进后续研究,我们公开发布评估框架,便于测试提示技术以恢复被删知识。
原文摘要 · Abstract (English)
In this work, we demonstrate that certain machine unlearning methods may fail under straightforward prompt attacks. We systematically evaluate eight unlearning techniques across three model families using output-based, logit-based, and probe analysis to assess the extent to which supposedly unlearned knowledge can be retrieved. While methods like RMU and TAR exhibit robust unlearning, ELM remains vulnerable to specific prompt attacks (e.g., prepending Hindi filler text to the original prompt recovers 57.3% accuracy). Our logit analysis further indicates that unlearned models are unlikely to hide knowledge through changes in answer formatting, given the strong correlation between output and logit accuracy. These findings challenge prevailing assumptions about unlearning effectiveness and highlight the need for evaluation frameworks that can reliably distinguish between genuine knowledge removal and superficial output suppression. To facilitate further research, we publicly release our evaluation framework to easily evaluate prompting techniques to retrieve unlearned knowledge.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。