arXiv:2608.17804cs.LGcs.CL2026-08

研究大模型遗忘中奖励设计对回答行为的影响,发现优化成功不等于真正遗忘。

An Empirical Study of Reward Specification and Benchmark Reliability in GRPO-based LLM Unlearning

  • 对比四种奖励机制,探索如何让模型在敏感问题上不泄露但也不拒绝回答
  • 发现优化过程看似成功,实则可能隐藏未完全遗忘的漏洞
  • 适合关注大模型安全与可控性的人士阅读

实际的大模型遗忘通常通过两个目标评估:消除特定知识并保留非目标能力。在生成式问答中,还存在第三种未明确的行为:当提示与目标相关但可作更广泛回答时,模型应以不泄露特定信息的方式回应,而非泄露、回避或拒绝。本文在受控的LoRA-GRPO-RWKU设置下研究这一规范问题,比较了四种奖励设计——词法抑制、反拒绝引导、基于评分标准的宽泛回答和显式拒绝对比——并考察是否有SFT预热。实验表明,优化成功并不等同于行为遗忘:RWKU遗忘分数、外部完成审计、训练终态滚动生成审计及训练动态可能得出不同结论。我们归因于奖励作弊终点、GRPO中的策略支持限制、基准探测无法捕捉终端变化,以及优化过程中可选择低语义泄露的宽主题回答。

原文摘要 · Abstract (English)

Practical LLM unlearning is usually evaluated through two objectives: suppress target-specific knowledge and preserve non-target utility. In generative QA, this leaves a third behavior underspecified: when a target-adjacent prompt admits a broader answer without target-specific leakage, the model should answer at that level rather than leak, evade, or refuse. We study this specification problem in a controlled LoRA-GRPO RWKU setting, comparing four reward designs that span lexical suppression, anti-refusal shaping, rubric-based broad answering, and an explicit refusal contrast, with and without SFT warm-up. The experiments show that optimization success is not equivalent to behavioral unlearning: RWKU forget scores, held-out completion audits, terminal training-rollout audits, and training dynamics can point to different conclusions. We trace these disagreements to reward-hacking endpoints, policy-support limits in GRPO, benchmark probes that miss endpoint changes, and rewards that can select broad-topic answering with low semantic leakage during optimization.

大模型遗忘奖励设计安全对齐语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。