arXiv:2510.12981cs.LG2025-10被引 3

现有评估方法骗人,新指标能真实检测模型是否忘掉数据

Reference-Specific Unlearning Metrics Can Hide the Truth: A Reality Check

  • 用双向似然比对生成样本,评估模型输出分布是否一致
  • 传统方法得分高时,新指标显示模型反而更不接近理想状态
  • 适合研究模型遗忘能力或安全性的研究人员使用

当前生成模型的遗忘评估依赖特定参考响应或分类器输出,而非核心目标:未学习模型是否与从未见过不良数据的模型在行为上无法区分。这种参考依赖方法存在系统性盲点,允许模型在看似成功时仍保留可通过其他提示或攻击获取的敏感知识。为此,我们提出功能对齐分布等价性(FADE)新指标,通过比较生成样本上的双向似然分配,衡量未学习模型与参考模型的分布相似性。不同于依赖预设参考的传统方法,FADE捕捉整个输出分布的功能对齐,提供对真实遗忘的合理评估。在LLM遗忘的TOFU基准和文本到图像扩散模型遗忘的UnlearnCanvas基准上的实验表明,许多在传统指标上表现接近最优的方法,在新指标下反而比遗忘前更远离理想状态。这些发现揭示了现有评估体系的根本缺陷,并证明FADE为开发和评估真正有效的遗忘方法提供了更可靠的基线。

原文摘要 · Abstract (English)

Current unlearning metrics for generative models evaluate success based on reference responses or classifier outputs rather than assessing the core objective: whether the unlearned model behaves indistinguishably from a model that never saw the unwanted data. This reference-specific approach creates systematic blind spots, allowing models to appear successful while retaining unwanted knowledge accessible through alternative prompts or attacks. We address these limitations by proposing Functional Alignment for Distributional Equivalence (FADE), a novel metric that measures distributional similarity between unlearned and reference models by comparing bidirectional likelihood assignments over generated samples. Unlike existing approaches that rely on predetermined references, FADE captures functional alignment across the entire output distribution, providing a principled assessment of genuine unlearning. Our experiments on the TOFU benchmark for LLM unlearning and the UnlearnCanvas benchmark for text-to-image diffusion model unlearning reveal that methods achieving near-optimal scores on traditional metrics fail to achieve distributional equivalence, with many becoming more distant from the gold standard than before unlearning. These findings expose fundamental gaps in current evaluation practices and demonstrate that FADE provides a more robust foundation for developing and assessing truly effective unlearning methods.

模型遗忘评估指标生成模型安全评测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。