提出新方法实现真正删除模型中的特定知识,而非伪装掩盖。
Unlearning vs. Obfuscation: Are We Truly Removing Knowledge?
- 用多选题和KL散度让模型对目标信息产生随机拒绝
- 在探测测试中拒绝率超90%,远高于混淆方法
- 适合关注数据隐私与合规的AI研发人员
去学习已成为大语言模型支持数据隐私、合规及伦理部署的关键能力。现有方法常通过注入错误或无关信息来隐藏知识,实为知识添加而非真实删除,易被探测攻击。本文首次明确区分去学习与混淆,并提出基于探测的评估框架,验证现有方法的真实有效性。进一步提出DF-MCQ方法,通过自动构建多选题并利用KL散度使模型预测分布平坦化,有效消除对目标个体的知识,触发合理拒绝行为。实验表明,该方法在探测问题上实现超过90%的拒绝率,且不确定性水平远高于混淆方法。
原文摘要 · Abstract (English)
Unlearning has emerged as a critical capability for large language models (LLMs) to support data privacy, regulatory compliance, and ethical AI deployment. Recent techniques often rely on obfuscation by injecting incorrect or irrelevant information to suppress knowledge. Such methods effectively constitute knowledge addition rather than true removal, often leaving models vulnerable to probing. In this paper, we formally distinguish unlearning from obfuscation and introduce a probing-based evaluation framework to assess whether existing approaches genuinely remove targeted information. Moreover, we propose DF-MCQ, a novel unlearning method that flattens the model predictive distribution over automatically generated multiple-choice questions using KL-divergence, effectively removing knowledge about target individuals and triggering appropriate refusal behaviour. Experimental results demonstrate that DF-MCQ achieves unlearning with over 90% refusal rate and a random choice-level uncertainty that is much higher than obfuscation on probing questions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。