arXiv:2608.17836cs.LG2026-08

利用知识编辑中的关联上下文检索,实现对大模型的白盒攻击

Leveraging Association Context Retrieval in Knowledge Edit- ing to Build White-Box Attacks on LLMs

论文配图:Leveraging Association Context Retrieval in Knowledge Edit- ing to Build White-Box Attacks on LLMs
图 1 · 摘自论文原文
  • 基于知识编辑的定位-修改思路,结合模型自生关联知识
  • 攻击成功率提升,且不显著损害模型通用能力
  • 适合研究大模型安全与对抗攻击的学者参考

随着大语言模型自主性增强,研究诱导其产生不安全行为的方法变得至关重要。本文提出一种新型白盒攻击方法,灵感来自知识编辑领域的‘定位-编辑’范式。观察发现,经此类编辑的模型会对编辑目标分配异常高的预测概率,这一特性在设计攻击时极具优势。我们通过从模型中检索关联知识,扩展了约束移除范围,使其覆盖整个主题类别,而非局限于预定义数据集中的提示。在多种模型架构上的实验表明,该方法相比现有方法攻击效果更优,且未对模型整体性能造成严重损害。

原文摘要 · Abstract (English)

As large language models (LLMs) are granted increasing autonomy, it is essential to investigate methods that can induce unsafe behavior. We propose a novel white-box attack inspired by locate-then-edit approaches from the field of Knowledge Editing. Our choice is motivated by the observation that models edited with such schemes tend to assign unusually high prediction probabilities to the edit target, a property that is particularly advantageous when designing attacks. We modify the editing framework by incorporating as- sociative knowledge retrieved from the model, thereby extending constraint removal to an entire thematic category rather than being limited to prompts from a predefined dataset. Experiments with various archi- tectures demonstrate improved attack effectiveness over competing methods without dealing critical damage to general model performance.

大模型安全白盒攻击知识编辑

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。