arXiv:2604.15847cs.CL2026-04ACL被引 1

让大模型忘记特定知识,同时保持推理能力

CiPO: Counterfactual Unlearning for Large Reasoning Models through Iterative Preference Optimization

论文配图:CiPO: Counterfactual Unlearning for Large Reasoning Models through Iterative Preference Optimization
图 1 · 摘自论文原文
  • 通过反事实推理迭代优化,精准干预模型思维链
  • 在复杂基准上完全清除中间推理和最终答案中的目标知识
  • 适合需要隐私保护且不牺牲推理性能的场景

机器遗忘近年来受到广泛关注,作为一种可选地从大规模人类数据训练的大型语言模型中移除不当隐私或版权信息的有前途技术。然而,强调长思维链(CoT)推理以应对复杂问题的大规模推理模型(LRMs)带来了遗忘难题:现有方法要么难以彻底消除思维链中的不良知识,要么因干扰推理过程而损害推理性能。为此,我们提出基于迭代偏好优化的反事实遗忘(CiPO),将遗忘重新定义为对LRM思维链的定向干预。具体而言,给定一个期望的遗忘目标答案,CiPO指导LRM生成逻辑有效的反事实推理路径用于偏好微调。随着LRM适应该反事实路径,CiPO迭代更新偏好学习数据以增强与原模型的差异。此迭代循环确保了理想的遗忘效果与平滑优化,有效缓解了上述困境。在多个挑战性基准上的实验表明,CiPO在遗忘方面表现卓越,能够完全移除思维链中间步骤及最终答案中的知识,同时保持LRM的推理能力。

原文摘要 · Abstract (English)

Machine unlearning has gained increasing attention in recent years, as a promising technique to selectively remove unwanted privacy or copyrighted information from Large Language Models that are trained on a massive scale of human data. However, the emergence of Large Reasoning Models (LRMs), which emphasize long chain-of-thought (CoT) reasoning to address complex questions, presents a dilemma to unlearning: existing methods either struggle to completely eliminate undesired knowledge from the CoT traces or degrade the reasoning performances due to the interference with the reasoning process. To this end, we introduce Counterfactual Unlearning through iterative Preference Optimization (CiPO), a novel framework that redefines unlearning as the targeted intervention of the CoT reasoning in LRMs. More specifically, given a desired unlearning target answer, CiPO instructs LRMs to generate a logically valid counterfactual reasoning trace for preference tuning. As the LRM adjusts to the counterfactual trace, CiPO iteratively updates the preference learning data to increase the discrepancy from the original model. This iterative loop ensures both desirable unlearning and smooth optimization, effectively mitigating the dilemma. Experiments on challenging benchmarks demonstrate that CiPO excels at unlearning, completely removing knowledge from both the intermediate CoT steps and the final answer, while preserving the reasoning abilities of LRMs.

模型遗忘推理链偏好优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。