提出FALCON方法,精准擦除大模型敏感信息而不损性能
FALCON: Fine-grained Activation Manipulation by Contrastive Orthogonal Unalignment for Large Language Model
- 通过对比正交解耦机制,精细操控激活值以分离知识
- 在多个数据集上实现90%以上知识擦除率,同时保持95%以上模型精度
- 适合需要高安全性的AI系统,如医疗或金融问答模型
大型语言模型广泛应用,但可能无意中编码敏感或有害信息,引发重大安全问题。机器遗忘技术应对此类风险,但现有训练期遗忘方法依赖粗粒度损失组合,在精确分离知识与平衡遗忘效果和模型实用性方面存在局限。本文提出一种新型表示引导的遗忘方法——细粒度激活操控对比正交解耦(FALCON),利用信息论指导高效参数选择,采用对比机制增强表示分离,并将冲突梯度投影至正交子空间,缓解遗忘与保留目标间的冲突。大量实验表明,FALCON在保持模型实用性的同时,显著提升遗忘效果,对知识恢复尝试表现出强鲁棒性。
原文摘要 · Abstract (English)
Large language models have been widely applied, but can inadvertently encode sensitive or harmful information, raising significant safety concerns. Machine unlearning has emerged to alleviate this concern; however, existing training-time unlearning approaches, relying on coarse-grained loss combinations, have limitations in precisely separating knowledge and balancing removal effectiveness with model utility. In contrast, we propose Fine-grained Activation manipuLation by Contrastive Orthogonal uNalignment (FALCON), a novel representation-guided unlearning approach that leverages information-theoretic guidance for efficient parameter selection, employs contrastive mechanisms to enhance representation separation, and projects conflict gradients onto orthogonal subspaces to resolve conflicts between forgetting and retention objectives. Extensive experiments demonstrate that FALCON achieves superior unlearning effectiveness while maintaining model utility, exhibiting robust resistance against knowledge recovery attempts.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。