提出通用数据删除方法REM,能有效清除各类错误数据影响。
Redirection for Erasing Memory (REM): Towards a universal unlearning method for corrupted data
- 用新增神经元重定向错误数据,实现精准清除
- 在不同数据污染场景下均优于现有最强方法
- 适合需要可靠数据清理的视觉模型部署场景
机器学习中的数据删除研究涉及多种任务,但现有方法多针对特定任务,难以系统比较。为此,我们提出一个概念空间,用于刻画视觉分类器中各类错误数据删除任务,其维度包括发现率(删除时已知错误数据的比例)和错误数据的统计规律性(从随机样本到共享概念)。此前方法仅适用于该空间的部分区域,且在其他区域表现失败。本文提出新方法红移记忆擦除(REM),核心是引入新神经元将错误数据重定向至其上,再将其丢弃或关闭,从而抑制错误数据影响。REM在全任务空间表现稳健,显著优于以往最优方法。
原文摘要 · Abstract (English)
Machine unlearning is studied for a multitude of tasks, but specialization of unlearning methods to particular tasks has made their systematic comparison challenging. To address this issue, we propose a conceptual space to characterize diverse corrupted data unlearning tasks in vision classifiers. This space is described by two dimensions, the discovery rate (the fraction of the corrupted data that are known at unlearning time) and the statistical regularity of the corrupted data (from random exemplars to shared concepts). Methods proposed previously have been targeted at portions of this space and-we show-fail predictably outside these regions. We propose a novel method, Redirection for Erasing Memory (REM), whose key feature is that corrupted data are redirected to dedicated neurons introduced at unlearning time and then discarded or deactivated to suppress the influence of corrupted data. REM performs strongly across the space of tasks, in contrast to prior SOTA methods that fail outside the regions for which they were designed.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。