测试大模型如何消除错误叙事,发现删掉信息后反而可能恢复相关人物。
Between Suppression and Collapse: Evaluating Narrative Unlearning with LENS

- 设计多层级评估框架,检测模型对特定叙事的抑制效果。
- 部分模型在删去叙事后仍能复现其核心内容,且存在实体恢复现象。
- 适合关注模型偏见与安全性的研究人员参考。
大型语言模型可能以合理解释的形式重现与虚假信息一致的叙事框架,引发对现有机器遗忘算法能否有效抑制此类行为的疑问。本文提出基于上下文的叙事抑制评估协议LENS,用于测试目标叙事在直接、归因、对比和抽象抵抗四个层级上的再现情况。评估针对两个来源有据的叙事:一是将俄罗斯对乌克兰的战争归因于北约东扩,二是将美国对台湾的政策描述为利用或抛弃。实验涵盖四个近120亿参数的多语言指令模型:Lapa LLM、Gemma-12B、Qwen-14B 和 TAIDE-Gemma。引入抑制-崩溃效率(SCE)评分作为检查点选择指标,奖励目标叙事抑制,同时惩罚输出质量下降。结果表明,经筛选的检查点可降低叙事再现率,且抑制能力可能超越直接遗忘提示。此外,还观察到一种副作用:抽象的A/B/C类提示可能导致模型在遗忘后重新激活与目标叙事相关的现实人物。这些发现表明LENS是一个有效的诊断工具,可用于报告并指导叙事遗忘深层结构的后续研究。
原文摘要 · Abstract (English)
Large language models (LLMs) can reproduce disinformation-aligned narrative frames as plausible explanations, raising the question of whether existing machine-unlearning algorithms can suppress this behavior. We introduce Level-based Evaluation of Narrative Suppression (LENS), a contextualization based evaluation protocol for testing target narrative reproduction across direct, attributed, contrastive, and abstract resistance levels. We evaluate two source-grounded narratives: one framing Russia's war against Ukraine as forced by NATO expansion, and one framing the United States as exploiting or abandoning Taiwan. The experiments cover four near-12B multilingual instruction models: Lapa LLM, Gemma-12B, Qwen-14B, and TAIDE-Gemma. We introduce the Suppression-Collapse Efficiency (SCE) score as a checkpoint selection summary that rewards target-narrative suppression while penalizing degraded outputs. Our results shows that selected checkpoints can reduce narrative reproduction and suppression may transfer beyond direct forget prompts. We also report entity recovery as a separate side effect: abstract A/B/C prompts can cause models to recover the real-world actors associated with the target frame after unlearning. These findings demonstrate that LENS is a successful diagnostic protocol for both reporting and guiding the further study of the deeper structure of narrative unlearning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。