提出统一表征与双向蒸馏方法,实现大模型多目标安全删减。
Harmonizing Multi-Objective LLM Unlearning via Unified Domain Representation and Bidirectional Logit Distillation

- 统一数据表示减少领域差异,协同优化多目标
- 双向蒸馏同时抑制有害行为并保留有用知识
- 兼顾隐私删除、鲁棒性与边界行为,适合高安全场景
大型语言模型(LLMs)的去学习对消除模型中的危险或泄露隐私的信息至关重要。实际应用中需同时满足多个挑战性目标:移除不良知识、保持通用能力、避免对邻近概念的过度拒绝,以及确保对对抗性探测攻击的鲁棒性。然而,现有方法通常仅关注部分目标,如去学习效果和能力保留,而忽视了鲁棒性与边界行为。简单扩展这些方法至多目标场景可能导致任务干扰。本文提出一种新型多目标去学习框架,通过数据与优化协同设计实现目标调和:将训练语料统一为一致的数据表示以缩小领域差距,并引入双向蒸馏机制,从上下文引导的教师模型中提取期望行为,同时在学生模型中抑制不良行为。理论与实证分析表明,该方法能对齐领域分布,将看似无关的去学习任务转化为协同优化。实验显示其达到当前最优性能,在多种严苛要求下实现均衡可靠的去学习。
原文摘要 · Abstract (English)
Large Language Models (LLMs) unlearning is crucial for removing hazardous or privacy-leaking information from the model. Practical LLM unlearning demands satisfying multiple challenging objectives simultaneously: removing undesirable knowledge, preserving general utility, avoiding over-refusal of neighboring concepts, and, crucially, ensuring robustness against adversarial probing attacks. However, existing unlearning methods primarily focus on a limited subset of these goals, typically unlearning efficacy and utility preservation while overlooking robustness and boundary behaviors. Naively extending these methods to multi-objective settings may lead to unlearning task interference. We propose a novel multi-objective unlearning framework that harmonizes multiple unlearning objectives through a data and optimization co-design: We standardize training corpora into a unified data representation to reduce the domain gap, and then introduce a bidirectional distillation method that simultaneously elicits desired behavior from a context-instructed teacher while suppressing undesirable behavior in the student model. Theoretical and empirical analyses show that our method aligns domain distributions and converts seemingly irrelevant unlearning tasks into cooperative optimization. Evaluation demonstrates state-of-the-art performance, which enables balanced and reliable unlearning across diverse, challenging requirements.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。