提出新方法,精准删除大模型敏感信息且不影响正常功能。
Feature-Selective Representation Misdirection for Machine Unlearning
- 通过激活重要性图引导有方向的干扰向量,选择性压制有害表征。
- 在20%-30%数据重叠下仍有效,性能优于现有方法且损失极小。
- 适合需要安全合规、隐私保护的AI系统部署场景。
随着大语言模型在安全关键和监管领域日益普及,其对敏感或禁止知识的保留带来了隐私泄露、合规风险及滥用等威胁。近期研究表明,机器遗忘可帮助模型满足不断变化的法律、安全与治理要求。然而,现有遗忘技术假设遗忘与保留数据集完全分离,在实际应用中分布高度纠缠时难以实现。基于扰动的方法常导致模型通用能力下降或无法保障安全。为此,我们提出特征选择性表征误导方法(SRMU),一种基于激活编辑的原理性框架,实现特征感知且方向可控的扰动。不同于无差别权重扰动,SRMU采用结构化误导向量与激活重要性图,可选择性抑制有害表征同时保留良性表征的实用性。在广泛使用的WMDP基准上,针对低熵与高熵配置进行实验,结果表明SRMU在最小实用性能损失下达到最先进遗忘效果,并在20%-30%数据重叠条件下仍有效,而现有基线方法在此情况下失效。SRMU为安全驱动的模型治理、隐私合规与可控知识移除提供了稳健基础。代码已公开于https://figshare.com/s/d5931192a8824de26aff。
原文摘要 · Abstract (English)
As large language models (LLMs) are increasingly adopted in safety-critical and regulated sectors, the retention of sensitive or prohibited knowledge introduces escalating risks, ranging from privacy leakage to regulatory non-compliance to to potential misuse, and so on. Recent studies suggest that machine unlearning can help ensure deployed models comply with evolving legal, safety, and governance requirements. However, current unlearning techniques assume clean separation between forget and retain datasets, which is challenging in operational settings characterized by highly entangled distributions. In such scenarios, perturbation-based methods often degrade general model utility or fail to ensure safety. To address this, we propose Selective Representation Misdirection for Unlearning (SRMU), a novel principled activation-editing framework that enforces feature-aware and directionally controlled perturbations. Unlike indiscriminate model weights perturbations, SRMU employs a structured misdirection vector with an activation importance map. The goal is to allow SRMU selectively suppresses harmful representations while preserving the utility on benign ones. Experiments are conducted on the widely used WMDP benchmark across low- and high-entanglement configurations. Empirical results reveal that SRMU delivers state-of-the-art unlearning performance with minimal utility losses, and remains effective under 20-30\% overlap where existing baselines collapse. SRMU provides a robust foundation for safety-driven model governance, privacy compliance, and controlled knowledge removal in the emerging LLM-based applications. We release the replication package at https://figshare.com/s/d5931192a8824de26aff.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。