发现语法相似是遗忘失败的主因,提出改写查询语句来提升模型遗忘效果。
Rethinking Benign Relearning: Syntax as the Hidden Driver of Unlearning Failures
- 用语法多样性改写遗忘请求,降低模型对旧内容的恢复倾向。
- 在多个基准上验证,语法相似数据即使主题不同也会触发遗忘失效。
- 既加速遗忘又保持模型性能,适合需要安全可控的场景使用。
机器遗忘旨在移除训练模型中的特定内容,同时保持整体性能。然而,良性重学习现象——即从无害微调数据中重新出现被遗忘的信息——表明现有遗忘方法仍存在根本脆弱性。传统解释归因于主题相关性,但本文通过系统分析发现,语法相似性才是主要驱动因素:在多个基准测试中,语法结构相似的数据即便无主题重叠,也因表示和梯度对齐而持续引发信息恢复。基于此,我们提出语法多样化策略,在遗忘前将原始遗忘请求改写为结构异构的形式。该方法有效抑制良性重学习,加速遗忘过程,并显著缓解遗忘效果与模型可用性之间的权衡。
原文摘要 · Abstract (English)
Machine unlearning aims to remove specific content from trained models while preserving overall performance. However, the phenomenon of benign relearning, in which forgotten information reemerges even from benign fine-tuning data, reveals that existing unlearning methods remain fundamentally fragile. A common explanation attributes this effect to topical relevance, but we find this account insufficient. Through systematic analysis, we demonstrate that syntactic similarity, rather than topicality, is the primary driver: across benchmarks, syntactically similar data consistently trigger recovery even without topical overlap, due to their alignment in representations and gradients with the forgotten content. Motivated by this insight, we introduce syntactic diversification, which paraphrases the original forget queries into heterogeneous structures prior to unlearning. This approach effectively suppresses benign relearning, accelerates forgetting, and substantially alleviates the trade-off between unlearning efficacy and model utility.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。