研究大模型在低质遗忘数据下的去记忆能力,发现语义核心保留时去记忆仍有效。
LLM Unlearning on Noisy Forget Sets: A Study of Incomplete, Rewritten, and Watermarked Data
- 针对被篡改、水印或低质量的遗忘数据,测试大模型去记忆效果
- 即使表面形式变化,只要核心语义保留,去记忆仍有效
- 适合关注大模型安全与隐私的开发者和研究人员
大型语言模型虽具备强大生成能力,但会记忆敏感信息、强化偏见并生成有害内容,引发伦理与安全担忧。为此,研究者提出大模型去记忆任务,旨在从预训练模型中移除与不良数据相关的知识。然而,现有方法多假设遗忘数据为清晰、完整样本,而真实场景中的遗忘数据常为低质量、合成重写或带水印的数据,影响去记忆可靠性。本文首次系统研究在受扰或低保真遗忘数据(即噪声遗忘集)下的去记忆表现。通过在典型去记忆方法RMU和NPO上基准测试,发现只要核心语义信号得以保留,去记忆仍具显著鲁棒性。我们提出基于显著性的解释:驱动遗忘的关键语义成分在表面形式剧烈变化下依然保持影响力。这表明去记忆算法主要依赖深层语义线索而非表层词汇模式。
原文摘要 · Abstract (English)
Large language models (LLMs) exhibit remarkable generative capabilities but raise ethical and security concerns by memorizing sensitive data, reinforcing biases, and producing harmful content. These risks have spurred interest in LLM unlearning, the task of removing knowledge associated with undesirable data from pre-trained models. However, most existing methods assume access to clean, well-defined forget data samples, whereas real-world forget data could often be low-quality, synthetically rewritten, or watermarked, casting doubt on the reliability of unlearning. This work presents the first study of unlearning under perturbed or low-fidelity forget data, referred to as noisy forget sets. By systematically benchmarking state-of-the-art LLM unlearning methods, RMU and NPO, on such noisy forget sets, we find that unlearning remains surprisingly robust to perturbations, provided that core semantic signals are preserved. To explain this robustness, we propose a saliency-based interpretation: key semantic components that drive forgetting remain consistently influential despite substantial variation in surface form. This suggests that unlearning algorithms are primarily guided by deep semantic cues rather than shallow lexical patterns.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。