arXiv:2503.01751cs.AIcs.CL2025-03ACL被引 12

用分布式激活调控,让大模型更稳定地更新知识。

SAKE: Steering Activations for Knowledge Editing

  • 将要编辑的事实建模为一组同义句和逻辑推论的分布
  • 在多个事实相关分布上实现更鲁棒的知识修改
  • 适合需要精准、泛化知识更新的研究者

由于大语言模型会记忆真实世界事实,因此需要以可控且高效的方式更新其知识。尽管现有知识编辑(KE)方法旨在修改特定事实,但存在上下文鲁棒性差和无法推广至相关逻辑推论等问题。为此,我们提出SAKE,一种通过最优传输建模事实分布的激活调控方法,将待编辑事实视为一组同义表达与逻辑推论构成的分布,从而在整体分布上调整模型行为。多个数值实验表明,相比现有方法,SAKE能实现更稳健的知识编辑效果。

原文摘要 · Abstract (English)

As Large Langue Models have been shown to memorize real-world facts, the need to update this knowledge in a controlled and efficient manner arises. Designed with these constraints in mind, Knowledge Editing (KE) approaches propose to alter specific facts in pretrained models. However, they have been shown to suffer from several limitations, including their lack of contextual robustness and their failure to generalize to logical implications related to the fact. To overcome these issues, we propose SAKE, a steering activation method that models a fact to be edited as a distribution rather than a single prompt. Leveraging Optimal Transport, SAKE alters the LLM behavior over a whole fact-related distribution, defined as paraphrases and logical implications. Several numerical experiments demonstrate the effectiveness of this method: SAKE is thus able to perform more robust edits than its existing counterparts.

知识编辑大模型激活调控最优传输

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。