arXiv:2509.05316cs.LGcs.AI2025-09被引 1

提出模块化方法提升大模型删忆的可靠性和效果

Standard vs. Modular Sampling: Best Practices for Reliable LLM Unlearning

  • 引入多样邻居集,平衡遗忘与模型能力
  • 标准1:1采样效率低,结果差,易隐藏性能缺陷
  • 提出MELU模块化策略,稳定实现有效删忆

传统大模型删忆通常分为'遗忘'和'保留'两部分,目标是移除遗忘集中的知识同时保留其余知识。在隐私导向的删忆研究中,保留集常被细分为与遗忘目标直接或间接关联的邻居集,并辅以通用知识集。现有基准普遍仅采用单一邻居集,未能反映真实数据复杂性。删忆通常采用1:1采样或循环迭代采样,但这些标准做法的有效性和稳定性尚未被充分评估。本研究系统评估了常见实践,发现单一邻居集次优,标准采样会掩盖性能权衡。基于此,我们提出并验证了一套最佳实践:(1) 引入多样化邻居集以平衡遗忘效果与模型效用;(2) 标准1:1采样效率低且表现不佳;(3) 提出模块化实体级删忆(MELU)作为循环采样的替代方案。实验证明,该模块化方法结合稳健算法,能提供清晰稳定的高效删忆路径。

原文摘要 · Abstract (English)

A conventional LLM Unlearning setting consists of two subsets -"forget" and "retain", with the objectives of removing the undesired knowledge from the forget set while preserving the remaining knowledge from the retain. In privacy-focused unlearning research, a retain set is often further divided into neighbor sets, containing either directly or indirectly connected to the forget targets; and augmented by a general-knowledge set. A common practice in existing benchmarks is to employ only a single neighbor set, with general knowledge which fails to reflect the real-world data complexities and relationships. LLM Unlearning typically involves 1:1 sampling or cyclic iteration sampling. However, the efficacy and stability of these de facto standards have not been critically examined. In this study, we systematically evaluate these common practices. Our findings reveal that relying on a single neighbor set is suboptimal and that a standard sampling approach can obscure performance trade-offs. Based on this analysis, we propose and validate an initial set of best practices: (1) Incorporation of diverse neighbor sets to balance forget efficacy and model utility, (2) Standard 1:1 sampling methods are inefficient and yield poor results, (3) Our proposed Modular Entity-Level Unlearning (MELU) strategy as an alternative to cyclic sampling. We demonstrate that this modular approach, combined with robust algorithms, provides a clear and stable path towards effective unlearning.

大模型删忆采样策略隐私保护

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。