arXiv:2601.11042cs.CLcs.AI2026-01ACL被引 2

通过分析模型权重的谱特性,提出新方法稳定大模型持续知识编辑

Spectral Characterization and Mitigation of Sequential Knowledge Editing Collapse

  • 发现模型通用能力与预训练权重的主要奇异方向密切相关
  • 在20,000次连续编辑下仍保持良好性能,显著缓解灾难性遗忘
  • 无需修改模型结构,可直接插入现有编辑流程使用

大语言模型中连续知识编辑常导致模型通用能力的灾难性崩溃,尤其在参数修改类方法中更为明显。现有方法依赖启发式约束来缓解该问题,但对退化机制理解不足。本文通过谱分析发现,模型通用能力与预训练权重矩阵的主要奇异方向紧密相关,这些方向对扰动极为敏感,且在多次编辑中被逐步破坏,其变化趋势与编辑效果和通用性能下降高度一致。基于此,我们提出REVIVE——一种即插即用框架,通过显式保护主要奇异子空间来稳定连续编辑。REVIVE在原始权重的谱基下表示参数更新,并过滤会干扰保护区域的成分。大量实验表明,REVIVE在多种模型和基准上均能持续提升编辑效果,同时在长达20,000次的连续编辑中显著保留通用能力。

原文摘要 · Abstract (English)

Sequential knowledge editing in large language models often causes catastrophic collapse of the model's general abilities, especially for parameter-modifying methods. Existing approaches mitigate this issue through heuristic constraints on parameter updates, yet the mechanisms underlying such degradation remain insufficiently understood. In this work, we present a spectral analysis of sequential knowledge editing and show that a model's general abilities are closely associated with dominant singular directions of pretrained weight matrices. These directions are highly sensitive to perturbations and are progressively disrupted by repeated edits, closely tracking the collapse in both editing efficacy and general performance. Building on this insight, we propose REVIVE, a plug-and-play framework that stabilizes sequential editing by explicitly preserving the dominant singular subspace. REVIVE represents parameter updates in the spectral basis of the original weights and filters components that would interfere with the protected region. Extensive experiments across multiple models and benchmarks show that REVIVE consistently improves editing efficacy while substantially preserving general abilities under long-horizon sequential editing, including extreme settings with up to 20,000 edits.

知识编辑大模型谱分析稳定性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。