提出双粒度数据合成方法,让大模型更彻底地遗忘特定内容。
From Domains to Instances: Dual-Granularity Data Synthesis for LLM Unlearning
- 用模型自身生成遗忘数据,通过提示词引导和对抗性策略匹配其知识分布。
- 在哈利·波特领域,相关性提升20%,多样性提高0.05,数据量减半。
- 适合关注模型隐私保护与可控遗忘的开发者和研究人员。
尽管机器遗忘对移除大语言模型中的隐私、有害或受版权保护内容至关重要,但现有基准常无法真实反映模型的实际遗忘范围。本文定义了两种不同遗忘粒度:域级和实例级,并提出 extit{BiForget} 自动化框架,用于合成高质量遗忘数据集。与依赖外部生成器的先前方法不同, extit{BiForget} 利用目标模型自身,通过种子引导和对抗性提示,激发与其内部知识分布一致的数据。在多个基准上的实验表明,该方法在相关性、多样性和效率之间取得更优平衡。定量结果表明,在哈利·波特领域,相关性提升约20,多样性提升约0.05,同时总数据规模减半。最终,该方法实现了更强的遗忘效果与更好的模型性能保留,为评估大模型遗忘能力提供了更严谨的基础。
原文摘要 · Abstract (English)
Although machine unlearning is essential for removing private, harmful, or copyrighted content from LLMs, current benchmarks often fail to faithfully represent the true ``forgetting scope'' learned by the model. We formalize two distinct unlearning granularities, domain-level and instance-level, and propose \BiForget, an automated framework for synthesizing high-quality forget sets. Unlike prior work relying on \emph{external} generators, \BiForget exploits the target model per se to elicit data that matches its internal knowledge distribution through seed-guided and adversarial prompting. Our experiments across diverse benchmarks show that it achieves a superior balance of relevance, diversity, and efficiency. Quantitatively, in the Harry Potter domain, it improves relevance by ${\sim}20$ and diversity by ${\sim}$0.05 while \emph{halving} the total data size compared to SOTAs. Ultimately, it facilitates more robust forgetting and better utility preservation, providing a more rigorous foundation for evaluating LLM unlearning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。