arXiv:2608.04519cs.AIcs.CL2026-08

提出新基准评估大模型去记忆的鲁棒性,发现现有方法易被多跳推理和恢复攻击泄露知识。

Leak-Resistant Unlearning: A New Benchmark for Evaluating Multi-Hop Reasoning Consistency and Recovery Robustness

论文配图:Leak-Resistant Unlearning: A New Benchmark for Evaluating Multi-Hop Reasoning Consistency and Recovery Robustness
图 1 · 摘自论文原文
  • 设计涵盖多跳推理路径的新评测基准,模拟复杂知识依赖场景。
  • 在3个模型、6种方法上测试,发现多数去记忆方法对恢复攻击不鲁棒。
  • 揭示遗忘质量、鲁棒性与模型性能间的权衡关系,适合安全与合规研究者。

评估机器去记忆方法对敏感知识是否彻底清除至关重要。当前基准主要包含单跳问题和有限的多跳问题,仍面临两大挑战:(1) 知识并非孤立存在,多样化的多跳推理路径可能引发比常规查询更严重的知识泄露;(2) 去记忆过程可能脆弱,未完全清除的知识可通过轻量级后去记忆适配等恢复攻击部分复现,导致静态评估不足。为此,本文提出 extit{Unlearning} 作为新基准,用于评估大语言模型在多种推理路径和恢复攻击下的知识移除鲁棒性。我们在3个模型、6种去记忆方法及2个精心构建的数据集上进行实验。结果表明,现有方法对多跳推理路径和恢复攻击均表现出脆弱性。我们进一步探索了去记忆质量、鲁棒性与模型效用之间的权衡关系。

原文摘要 · Abstract (English)

Benchmarking machine unlearning methods is critical to understand whether sensitive knowledge is removed from large language models (LLMs) or not. Current unlearning benchmarks include mainly single-hop questions and a narrow set of multi-hop questions. Although effective, they still face two challenges. (1) Knowledge is not isolated, whereby diverse multi-hop reasoning paths can potentially induce knowledge leakage than normal queries. (2) Unlearning may be fragile: unlearned knowledge can be partially recovered through recovery attacks such as lightweight post-unlearning adaptation, making static evaluation insufficient. Therefore, in this paper, we introduce \unlearning as a novel benchmark to understand robust LLM knowledge removal across diverse reasoning paths and recovery attacks. We experiment with this benchmark on 3 models, 6 unlearning methods, and 2 carefully curated datasets. Results show that existing methods are vulnerable to multi-hop reasoning paths and recovery attacks. We further explore the trade-off among forget quality, robustness, and model utility for LLM unlearning.

大模型去记忆安全性评测基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。