arXiv:2503.04795cs.CLcs.AI2025-03ACL

通过调整模型权重实现精准删去敏感信息,同时保留有用知识。

Cyber for AI at SemEval-2025 Task 4: Forgotten but Not Lost: The Balancing Act of Selective Unlearning in Large Language Models

  • 用全局权重修改方法平衡删减效果与知识保留。
  • 7B和1B模型在测试集上得分分别达0.409和0.389。
  • 适合关注AI隐私保护与模型可删除性的研究者。

大型语言模型在处理敏感或过时数据时面临隐私、伦理与合规挑战,需选择性移除特定内容。重新训练模型成本过高,难以实施,因此亟需高效替代方案。作为SemEval 2025任务4的一部分,本文聚焦于大模型中的选择性遗忘技术。我们提出基于全局权重修改的方法,在遗忘效果、知识保留与模型后续可用性之间取得平衡。论文还详细介绍了任务特异性评估机制、实验结果与挑战。所提算法在7B和1B目标模型的测试集上分别获得0.409和0.389的综合得分,验证了可验证的大模型遗忘能力,展现了良好前景。

原文摘要 · Abstract (English)

Large Language Models (LLMs) face significant challenges in maintaining privacy, ethics, and compliance, when sensitive or obsolete data must be selectively removed. Retraining these models from scratch is computationally infeasible, necessitating efficient alternatives. As part of the SemEval 2025 Task 4, this work focuses on the application of selective unlearning in LLMs to address this challenge. In this paper, we present our experiments and findings, primarily leveraging global weight modification to achieve an equilibrium between effectiveness of unlearning, knowledge retention, and target model's post-unlearning utility. We also detail the task-specific evaluation mechanism, results, and challenges. Our algorithms have achieved an aggregate score of 0.409 and 0.389 on the test set for 7B and 1B target models, respectively, demonstrating promising results in verifiable LLM unlearning.

大模型隐私保护选择性遗忘

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。