arXiv:2509.25743cs.LGcs.CL2025-09ACL

提出旋转控制方法,让大模型持续删除数据时保持性能不下降。

Rotation Control Unlearning: Quantifying and Controlling Continuous Unlearning for LLM with The Cognitive Rotation Space

  • 用旋转空间量化每次删数据的程度,精确控制删除强度。
  • 无需保留原始训练数据,仍达当前最优效果。
  • 适合需要频繁更新数据安全性的实际应用系统。

随着大语言模型广泛应用,其安全漏洞日益受关注。机器遗忘被用于消除不良数据的影响,但现有方法不仅依赖保留数据以维持模型性能,还在连续遗忘请求下出现累积性性能崩溃。为此,我们提出旋转控制遗忘(RCU)方法,利用旋转显著性权重量化并控制连续遗忘过程中的遗忘程度。设计斜对称损失构建认知旋转空间,通过旋转角度变化模拟连续遗忘过程。进一步引入正交旋转轴正则化,强制不同遗忘请求间旋转方向相互垂直,有效减少干扰,缓解累积性能损失。多数据集实验表明,本方法在无需保留数据集的情况下达到当前最优性能。

原文摘要 · Abstract (English)

As Large Language Models (LLMs) become increasingly prevalent, their security vulnerabilities have already drawn attention. Machine unlearning is introduced to seek to mitigate these risks by removing the influence of undesirable data. However, existing methods not only rely on the retained dataset to preserve model utility, but also suffer from cumulative catastrophic utility loss under continuous unlearning requests. To solve this dilemma, we propose a novel method, called Rotation Control Unlearning (RCU), which leverages the rotational salience weight of RCU to quantify and control the unlearning degree in the continuous unlearning process. The skew symmetric loss is designed to construct the existence of the cognitive rotation space, where the changes of rotational angle can simulate the continuous unlearning process. Furthermore, we design an orthogonal rotation axes regularization to enforce mutually perpendicular rotation directions for continuous unlearning requests, effectively minimizing interference and addressing cumulative catastrophic utility loss. Experiments on multiple datasets confirm that our method without retained dataset achieves SOTA performance.

大模型安全机器遗忘连续学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。