arXiv:2501.04952cs.LGcs.AI2025-01被引 58

提出机器遗忘在AI安全中的局限性与未解难题

Open Problems in Machine Unlearning for AI Safety

  • 识别遗忘技术在敏感领域中难以避免副作用
  • 揭示遗忘危险知识可能损害有益用途
  • 指出评估与安全机制兼容性等关键挑战

随着人工智能在网络安全、生物研究和医疗等关键领域日益自主,确保其安全性与人类价值观对齐至关重要。机器遗忘——选择性删除或抑制特定知识的能力——已在隐私保护和数据移除任务中展现潜力,但其在AI安全中的应用仍面临挑战。本文聚焦于双重用途知识(如网络安全、化学生物辐射核安全领域)的管理问题:同一信息既可造福人类又可被滥用。模型可能组合看似无害的信息实现有害目的,此时强制遗忘将严重影响有益应用。论文系统梳理了当前遗忘技术的内在约束与开放问题,包括遗忘对模型的广泛副作用、与现有安全机制之间的新矛盾,以及评估、鲁棒性与遗忘过程中安全特征保持等难题。通过明确这些限制,旨在引导未来研究在更广阔的AI安全框架中合理运用遗忘技术,承认其边界并探索替代方案。

原文摘要 · Abstract (English)

As AI systems become more capable, widely deployed, and increasingly autonomous in critical areas such as cybersecurity, biological research, and healthcare, ensuring their safety and alignment with human values is paramount. Machine unlearning -- the ability to selectively forget or suppress specific types of knowledge -- has shown promise for privacy and data removal tasks, which has been the primary focus of existing research. More recently, its potential application to AI safety has gained attention. In this paper, we identify key limitations that prevent unlearning from serving as a comprehensive solution for AI safety, particularly in managing dual-use knowledge in sensitive domains like cybersecurity and chemical, biological, radiological, and nuclear (CBRN) safety. In these contexts, information can be both beneficial and harmful, and models may combine seemingly harmless information for harmful purposes -- unlearning this information could strongly affect beneficial uses. We provide an overview of inherent constraints and open problems, including the broader side effects of unlearning dangerous knowledge, as well as previously unexplored tensions between unlearning and existing safety mechanisms. Finally, we investigate challenges related to evaluation, robustness, and the preservation of safety features during unlearning. By mapping these limitations and open challenges, we aim to guide future research toward realistic applications of unlearning within a broader AI safety framework, acknowledging its limitations and highlighting areas where alternative approaches may be required.

AI安全机器遗忘双重用途

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。