arXiv:2603.00823cs.CLcs.AI2026-03被引 1

测试大模型在对话中遗忘效果是否稳定,发现静态评估可能高估了实际效果。

A Comprehensive Evaluation of LLM Unlearning Robustness under Multi-Turn Interaction

  • 设计交互场景模拟真实对话中的自我修正与条件提问
  • 静态评估中已遗忘的知识在交互中可被恢复,暴露评估局限性
  • 强调需在动态交互下验证遗忘稳定性,适合关注模型安全的开发者

机器遗忘旨在不从头训练的情况下移除特定训练数据对预训练模型的影响,随着大型语言模型面临安全、隐私和法律问题,这一技术日益重要。尽管先前研究多集中在静态、单轮设置下的遗忘评估,但真实交互环境中的遗忘鲁棒性仍缺乏探索。本文通过分析两种常见交互模式——自我修正和对话条件查询——考察遗忘在交互环境中的稳定性。结果发现,静态评估中看似已遗忘的知识常可通过交互恢复。虽然更强的遗忘策略提升了表面鲁棒性,却往往导致行为僵化而非真正知识清除。研究提示:静态评估可能高估实际应用效果,强调需在交互设置下确保遗忘的稳定性。

原文摘要 · Abstract (English)

Machine unlearning aims to remove the influence of specific training data from pre-trained models without retraining from scratch, and is increasingly important for large language models (LLMs) due to safety, privacy, and legal concerns. Although prior work primarily evaluates unlearning in static, single-turn settings, forgetting robustness under realistic interactive use remains underexplored. In this paper, we study whether unlearning remains stable in interactive environments by examining two common interaction patterns: self-correction and dialogue-conditioned querying. We find that knowledge appearing forgotten in static evaluation can often be recovered through interaction. Although stronger unlearning improves apparent robustness, it often results in behavioral rigidity rather than genuine knowledge erasure. Our findings suggest that static evaluation may overestimate real-world effectiveness and highlight the need for ensuring stable forgetting under interactive settings.

大模型遗忘交互

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。