arXiv:2511.00030cs.LGcs.AI2025-11被引 4

发现大模型删知识时会漏掉无辜信息,且现有测试方法看不见这种损失。

Probing Knowledge Holes in Unlearned LLMs

  • 通过分析被删内容的邻近和潜在影响区域,生成测试用例探测知识漏洞
  • 98.7%的测试用例中,删过知识的模型给出无关或荒谬回答
  • 提醒研究者:不能只靠常规基准测试评估知识保留能力

机器去学习作为一项主流技术,可在不重新训练的前提下选择性删除预训练期间吸收的不当知识。尽管近期去学习方法能有效移除不良内容而不会严重损害标准基准上的表现,我们发现它们可能无意中造成“知识漏洞”——即标准基准未能捕捉到的良性知识意外丢失。为探测去学习模型的知识漏洞位置,我们提出一种测试用例生成框架,探索被删内容的直接邻域及更广泛的潜在失败区域。评估结果显示去学习存在显著隐藏代价:高达98.7%的测试用例中,去学习后的模型给出无关或无意义响应,而原始预训练模型可正确回答。这些发现要求重新思考当前对去学习中知识保留的评估方式,摆脱依赖标准静态基准的惯性。

原文摘要 · Abstract (English)

Machine unlearning has emerged as a prevalent technical solution for selectively removing unwanted knowledge absorbed during pre-training, without requiring full retraining. While recent unlearning techniques can effectively remove undesirable content without severely compromising performance on standard benchmarks, we find that they may inadvertently create ``knowledge holes'' -- unintended losses of benign knowledge that standard benchmarks fail to capture. To probe where unlearned models reveal knowledge holes, we propose a test case generation framework that explores both immediate neighbors of unlearned content and broader areas of potential failures. Our evaluation demonstrates significant hidden costs of unlearning: up to 98.7\% of the test cases yield irrelevant or nonsensical responses from unlearned models, despite being answerable by the pretrained model. These findings necessitate rethinking the conventional approach to evaluating knowledge preservation in unlearning, moving beyond standard, static benchmarks.

大模型去学习知识漏洞模型评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。