arXiv:2604.04255cs.LGcs.CR2026-04KDD

攻击大推理模型的遗忘机制,让其给出错误答案却有合理假推理。

Towards Unveiling Vulnerabilities of Large Reasoning Models in Machine Unlearning

  • 设计双层精确攻击,利用可微目标与关键词对齐实现有效遗忘攻击。
  • 在白盒与黑盒场景下均成功诱导模型输出错误结果,且推理过程看似合理。
  • 揭示了大推理模型在数据遗忘时的新安全风险,适合关注AI安全的研究者。

大型语言模型(LLMs)具备强大的语义理解能力,推动了数据挖掘的发展。大推理模型(LRMs)通过显式的多步推理轨迹进一步提升了性能。与此同时,用户“被遗忘的权利”催生了机器遗忘技术,旨在不重新训练模型的前提下消除特定数据的影响。然而,遗忘过程可能引入新的安全漏洞,暴露额外的攻击面。尽管已有诸多关于遗忘攻击的研究,但针对LRMs的工作仍属空白。本文首次提出针对LRM的遗忘攻击,使模型在生成看似合理但误导性的推理链的同时输出错误最终答案。该目标因逻辑约束不可微、长推理链优化效果弱及遗忘数据集选择离散而极具挑战性。为此,我们提出一种双层精确遗忘攻击方法,包含可微目标函数、影响词对齐机制和松弛指示策略。为验证攻击的有效性与泛化能力,我们设计了新型优化框架,并在白盒与黑盒设置下开展全面实验,旨在揭示大推理模型遗忘流程中的新兴威胁。

原文摘要 · Abstract (English)

Large language models (LLMs) possess strong semantic understanding, driving significant progress in data mining applications. This is further enhanced by large reasoning models (LRMs), which provide explicit multi-step reasoning traces. On the other hand, the growing need for the right to be forgotten has driven the development of machine unlearning techniques, which aim to eliminate the influence of specific data from trained models without full retraining. However, unlearning may also introduce new security vulnerabilities by exposing additional interaction surfaces. Although many studies have investigated unlearning attacks, there is no prior work on LRMs. To bridge the gap, we first in this paper propose LRM unlearning attack that forces incorrect final answers while generating convincing but misleading reasoning traces. This objective is challenging due to non-differentiable logical constraints, weak optimization effect over long rationales, and discrete forget set selection. To overcome these challenges, we introduce a bi-level exact unlearning attack that incorporates a differentiable objective function, influential token alignment, and a relaxed indicator strategy. To demonstrate the effectiveness and generalizability of our attack, we also design novel optimization frameworks and conduct comprehensive experiments in both white-box and black-box settings, aiming to raise awareness of the emerging threats to LRM unlearning pipelines.

大模型安全机器遗忘推理攻击

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。