arXiv:2604.06154cs.CL2026-04

通过保留特定知识来遗忘所有其他内容,提升大模型安全性

Exclusive Unlearning

  • 只保留所需知识,主动遗忘其余所有内容
  • 可抵御各类越狱攻击并安全应对医疗数学等指令
  • 适合高风险场景下需全面消除有害输出的模型

将大语言模型应用于医疗、教育等工业场景时,生成有害内容的风险日益突出。现有机器遗忘方法虽能删除特定有害知识与表达,但面对多样化的有害内容仍难全面清除。本文提出专属遗忘(Exclusive Unlearning, EU),不再逐项列出需遗忘的目标,而是通过广泛遗忘除保留知识外的所有内容,实现对多种有害输出的泛化性移除。实验表明,采用该方法可获得对各类输入(包括越狱攻击)均具安全性的模型,同时保持对医学、数学等特定领域多样化指令的响应能力。

原文摘要 · Abstract (English)

When introducing Large Language Models (LLMs) into industrial applications, such as healthcare and education, the risk of generating harmful content becomes a significant challenge. While existing machine unlearning methods can erase specific harmful knowledge and expressions, diverse harmful content makes comprehensive removal difficult. In this study, instead of individually listing targets for forgetting, we propose Exclusive Unlearning (EU), which aims for broad harm removal by extensively forgetting everything except for the knowledge and expressions we wish to retain. We demonstrate that through Exclusive Unlearning, it is possible to obtain a model that ensures safety against a wide range of inputs, including jailbreaks, while maintaining the ability to respond to diverse instructions related to specific domains such as medicine and mathematics.

大模型安全遗忘学习AI伦理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。