解决大模型同主题知识编辑后指令跟随失效问题
Beyond the Covariance Trap: Unlocking Generalization in Same-Subject Knowledge Editing for Large Language Models
- 通过几何对齐与分层知识融合提升编辑稳定性
- 在多个数据集上使指令遵循准确率提升15%-23%
- 适合需要可靠动态记忆的智能体系统开发者
尽管定位-编辑式知识更新能高效修改大语言模型(LLM)中的知识,但在实际同主题知识编辑场景中出现关键泛化失败:模型虽能回忆原始编辑内容,却无法在用户指令下正确调用。本文揭示该现象的几何根源在于提示变化引发的内部激活漂移超出模型编辑后的泛化容忍度。其不稳定性源于双重病理:(1) 正交梯度联合优化使解陷入尖锐极小值,稳定域狭窄;(2) 标准协方差约束反而形成‘协方差陷阱’,放大输入扰动。为此提出RoSE(Robust Same-subject Editing),采用各向同性几何对齐最小化表征偏移,并通过分层知识整合平滑优化景观。大量实验表明,RoSE显著提升指令遵循能力,为大模型智能体的鲁棒参数化记忆奠定基础。
原文摘要 · Abstract (English)
While locate-then-edit knowledge editing efficiently updates knowledge encoded within Large Language Models (LLMs), a critical generalization failure mode emerges in the practical same-subject knowledge editing scenario: models fail to recall the updated knowledge when following user instructions, despite successfully recalling it in the original edited form. This paper identifies the geometric root of this generalization collapse as a fundamental conflict where the inner activation drifts induced by prompt variations exceed the model's geometric tolerance for generalization after editing. We attribute this instability to a dual pathology: (1) The joint optimization with orthogonal gradients collapses solutions into sharp minima with narrow stability, and (2) the standard covariance constraint paradoxically acts as a Covariance Trap that amplifies input perturbations. To resolve this, we introduce RoSE (Robust Same-subject Editing), which employs Isotropic Geometric Alignment to minimize representational deviation and Hierarchical Knowledge Integration to smooth the optimization landscape. Extensive experiments demonstrate that RoSE significantly improves instruction-following capabilities, laying the foundation for robust interactive parametric memory of LLM agents.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。