arXiv:2512.16227cs.CLcs.AI2025-12被引 1

用信息瓶颈理论精准编辑大模型知识,避免错误扩散。

An Information-Theoretic Framework for Robust Large Language Model Editing

  • 基于信息瓶颈理论压缩关键知识,隔离更新范围。
  • 在多模型架构上实现高精度、广泛适用的编辑效果。
  • 适合需要安全更新知识的大模型应用者使用。

大语言模型在科学、技术和社会领域已不可或缺,但其内部错误或过时信息会降低准确性并限制安全部署。现有编辑方法常难以泛化至宽域,导致意外后果。本文提出基于信息瓶颈理论的新框架,精确压缩并隔离修正所需的核心信息,最小化对无关行为的干扰。在此基础上,我们设计了信息瓶颈知识编辑器(IBKE),利用紧凑的隐空间表示引导梯度更新,实现鲁棒且广泛适用的模型编辑。我们在多种大模型架构和标准基准任务上验证了IBKE的有效性,结果表明其在编辑精度、泛化性和特异性方面均达到当前最优水平。该工作建立了一个理论严谨且实用的开放域知识编辑范式,推动大模型在真实场景中的可用性与可信度。

原文摘要 · Abstract (English)

Large Language Models (LLMs) have become indispensable tools in science, technology, and society, enabling transformative advances across diverse fields. However, errors or outdated information within these models can undermine their accuracy and restrict their safe deployment. Developing efficient strategies for updating model knowledge without the expense and disruption of full retraining remains a critical challenge. Current model editing techniques frequently struggle to generalize corrections beyond narrow domains, leading to unintended consequences and limiting their practical impact. Here, we introduce a novel framework for editing LLMs, grounded in information bottleneck theory. This approach precisely compresses and isolates the essential information required for generalizable knowledge correction while minimizing disruption to unrelated model behaviors. Building upon this foundation, we present the Information Bottleneck Knowledge Editor (IBKE), which leverages compact latent representations to guide gradient-based updates, enabling robust and broadly applicable model editing. We validate IBKE's effectiveness across multiple LLM architectures and standard benchmark tasks, demonstrating state-of-the-art accuracy and improved generality and specificity of edits. These findings establish a theoretically principled and practical paradigm for open-domain knowledge editing, advancing the utility and trustworthiness of LLMs in real-world applications.

大模型编辑信息瓶颈知识更新

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。