arXiv:2607.19442cs.LGcs.AI2026-07

提出新方法评估模型遗忘能力,发现现有标准可能误判保留信息。

Unlearning as Distribution Restoration: A Controlled Counterfactual Study, a Validated Selective Screen, and the Limits of Oracle-Free Certification

论文配图:Unlearning as Distribution Restoration: A Controlled Counterfactual Study, a Validated Selective Screen, and the Limits of Oracle-Free Certification
图 1 · 摘自论文原文
  • 将遗忘视为分布恢复,以匹配参考模型为基准评估
  • 45组实验中仅15组可被无证书认证,多数方法无法通过验证
  • 推荐使用筛选性测试,适合研究模型隐私与数据遗忘机制

机器遗忘常通过重训练模型与原模型在探测任务上的表现匹配来评估。在可控的虚假事实测试环境中,我们发现该标准可能偏好保留未学习知识的方法:被判定合格的方法仍比从未学习的水平多保留2.82纳特信息(置信区间[-3.16, -2.48])。我们将遗忘重新定义为向匹配参考分布的恢复,并在跨五类开源架构的45个模型-种子组合中检验无证书筛选与证书式标准。参考模型自身即否定了绝对保留/回溯证书的有效性:注入模型在41/45次中未能通过固定保留阈值,31/45次回溯失败;而参考模型仅在1/45次完全通过。基于基础锚定的未见样本筛选仍具强效,作为必要筛选测试,在封闭挑战集上45次全拒注入模型,44次接受参考模型,并部分检测出实体路由抑制(35/45)。以参考模型自身工作点为基准的相对损益校准,仅在15/45次中可认证,其通过项位于重训练噪声范围(0.80纳特)内,而传统探测标准偏差达5.17纳特。固定幅度逻辑抑制攻击在12/45次中击穿全前向测试,说明仅前向认证不可靠;本方法为实际生成方法的实证筛选测试。一个可识别性定理限定了哪些事实能实现无证书遗忘阈值,预测TOFU为边界情况。

原文摘要 · Abstract (English)

Machine unlearning is commonly evaluated by matching a retrained oracle on trained probes. In a controlled nonce-fact testbed with a matched retraining reference, we find this criterion can favor methods that retain held-out knowledge: candidates it rates adequate score held-out forget facts $-2.82$ nats below the never-learned level (cluster CI $[-3.16,-2.48]$). We recast unlearning as restoration to the matched reference and audit oracle-free screens and certificate-style criteria across 45 model-seed cells spanning five open architecture families. The reference itself falsifies an absolute retain/round-trip certificate: the injected model, which retains the retain set by construction, fails the fixed retain threshold in 41/45 cells and its own round trip in 31/45, and the reference fully certifies in only 1/45. A base-anchored held-out screen remains strong as a selective necessary test: on a sealed challenge suite it rejects the injected model in 45/45 cells, accepts the reference in 44/45, and partially detects entity-routing suppression (35/45); it is a necessary test with measured sensitivity, not a sufficiency certificate. A damage-relative recalibration anchored to the reference's own operating point certifies a small subset in 15/45 cells; where it does not abstain, its picks lie within retraining noise (0.80 nats) on the axes it optimizes, while the common trained-probe criterion sits 5.17 nats away (a supporting comparison, not a head-to-head benchmark). A fixed-magnitude logit-suppression attack defeats the full forward battery in 12/45 cells, so forward-only certification is not sound; our method is an empirical selective test for methods-as-produced. An identifiability theorem delimits which facts admit an oracle-free forget threshold at all, with TOFU as the predicted boundary case.

模型遗忘隐私保护无证书认证反事实评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。