提出新评测方法,检测模型是否仍能恢复被遗忘的关联
Association Restoration Test: Revealing Restorable Shortcuts after Unlearning

- 通过放大残差成分,测试遗忘后关联能否被原分类器重新利用
- 在多个数据集上发现现有评估指标无法捕捉功能性的关联残留
- 适合关注模型真实遗忘效果的研究者,尤其对关联性偏见敏感任务
关联遗忘旨在消除模型对标签-属性捷径的学习,同时保持任务性能。现有评估主要衡量输出鲁棒性或探测冻结特征中捷径属性是否仍可读,但均无法判断保留的关联是否仍可被原始分类器实际使用。本文提出关联恢复测试(ART),一种事后诊断方法,用于评估功能性的捷径可恢复性。ART估计类别条件下的关联方向,放大残差成分,并用原始分类器头评估修改后的特征。在Waterbirds、CelebA、SpuCoDogs以及ISIC时间戳伪影扩展数据集上,我们表明输出指标、表示探测和ART分别刻画了捷径缓解的不同方面。这些发现推动了针对学习到的关联而非单个类别或概念的遗忘与捷径缓解方法的恢复感知评估。
原文摘要 · Abstract (English)
Association unlearning aims to disable learned label-attribute shortcuts while preserving task performance. Existing evaluations mainly measure output-level robustness or probe whether shortcut attributes remain readable in frozen features, but neither test determines whether a retained association remains functionally usable by the original classifier. We propose the Association Restoration Test (ART), a post-hoc diagnostic for functional shortcut restorability. ART estimates class-conditional association directions, amplifies residual components, and evaluates the modified features with the original classifier head. Across Waterbirds, CelebA, SpuCoDogs, and an ISIC timestamp-artifact extension, we show that output metrics, representation probes, and ART characterize distinct aspects of shortcut mitigation. These findings motivate restoration-aware evaluation for unlearning and shortcut-mitigation methods that target learned associations rather than individual classes or concepts.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。