无需保留数据集即可实现大模型知识删除,且不损害模型性能。
Classifier-free guidance in LLMs Safety
- 用合成数据直接训练,结合修正的无分类器引导策略。
- 在不降低模型能力前提下,显著提升知识遗忘效果。
- 适合关注大模型安全与隐私保护的研究者和开发者。
本文提出一种无需保留数据集的大语言模型知识删除方法,采用基于ORPO的强化学习框架,并在推理阶段引入改进的无分类器引导技术。通过在符合无分类器引导机制的训练设置中直接使用合成替代数据进行训练,实现了显著的遗忘效果,同时未造成模型性能下降。该工作为大模型安全与隐私保护提供了新思路,是NeurIPS 2024 LLM-PC投稿的扩展版本,曾获二等奖。
原文摘要 · Abstract (English)
The paper describes LLM unlearning without a retaining dataset, using the ORPO reinforcement learning method with inference enhanced by modified classifier-free guidance. Significant improvement in unlearning, without degradation of the model, is achieved through direct training on synthetic replacement data in CFG-aware training regime, with classifier-free guidance applied during the inference. This article is an extended version of the NeurIPS 2024 LLM-PC submission, which was awarded second prize.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。