arXiv:2504.06658cs.LGcs.AI2025-04被引 2

提出量化样本遗忘难易度的新指标,提升大模型隐私删除效率

A Neuro-inspired Interpretation of Unlearning in Large Language Models through Sample-level Unlearning Difficulty

  • 受神经科学启发,定义样本级遗忘难度度量MRD
  • 发现难删样本具特定特征,易删样本可优先处理
  • 基于MRD加权采样,显著提升现有遗忘算法效果

为应对隐私保护法规,大语言模型(LLM)的遗忘能力日益受到关注。然而,现有研究常忽视遗忘过程的可解释性,尤其缺乏对样本级遗忘难度的分析。多数研究假设所有样本遗忘难度一致,这一简化可能将算法性能归因于样本选择而非算法设计,误导发展方向。为此,本文研究LLM遗忘与样本特征的关系,聚焦遗忘难度。受神经科学启发,提出记忆移除难度(MRD)度量,用于量化样本级遗忘难度。利用MRD,分析难删与易删样本的特征差异,并提出基于MRD的加权采样方法,优先选择易遗忘样本以优化现有遗忘算法。在公开基准和数据集上验证,结果表明该度量与方法有效。

原文摘要 · Abstract (English)

Driven by privacy protection laws and regulations, unlearning in Large Language Models (LLMs) is gaining increasing attention. However, current research often neglects the interpretability of the unlearning process, particularly concerning sample-level unlearning difficulty. Existing studies typically assume a uniform unlearning difficulty across samples. This simplification risks attributing the performance of unlearning algorithms to sample selection rather than the algorithm's design, potentially steering the development of LLM unlearning in the wrong direction. Thus, we investigate the relationship between LLM unlearning and sample characteristics, with a focus on unlearning difficulty. Drawing inspiration from neuroscience, we propose a Memory Removal Difficulty ($\mathrm{MRD}$) metric to quantify sample-level unlearning difficulty. Using $\mathrm{MRD}$, we analyze the characteristics of hard-to-unlearn versus easy-to-unlearn samples. Furthermore, we propose an $\mathrm{MRD}$-based weighted sampling method to optimize existing unlearning algorithms, which prioritizes easily forgettable samples, thereby improving unlearning efficiency and effectiveness. We validate the proposed metric and method using public benchmarks and datasets, with results confirming its effectiveness.

大模型遗忘学习可解释性样本筛选

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。