让大模型真正遗忘敏感信息,应对网络安全与隐私风险
LLM Unlearning for Cyber Defense: A Survey on Methods, Challenges, and Emerging Threats
- 通过梯度方法实现无需重训的模型知识删除
- 可有效抑制记忆信息泄露,避免数据提取与成员推断攻击
- 适合关注大模型安全与合规性的研究人员和开发者
大语言模型日益应用于医疗、金融、教育及决策支持等关键领域,但其无法遗忘特性带来严重的网络安全、隐私与安全风险。敏感个人信息、版权内容、危险领域知识及训练数据记忆长期存在于数十亿参数中,使模型易受信息提取、越狱攻击、成员推断等威胁,且可能引发法律与财务损失。真实案例显示,聊天机器人重现私密信息或虚构法律引用已造成实际损害。由于对十亿参数模型进行重新训练在计算上不可行,且知识分布于全局参数而非局部单元,因此无需重训的大模型遗忘技术成为主要防御手段,旨在移除或抑制特定知识而不影响其他有用能力。然而核心问题仍存:现有方法是否真正抹除知识,还是仅抑制常规提示下的表达?本综述从安全、鲁棒性与可验证遗忘角度审视大模型遗忘,重点关注与现有训练流程兼容且可扩展至十亿级参数的基于梯度的方法。
原文摘要 · Abstract (English)
LLMs are increasingly deployed in security-critical systems across healthcare, finance, education, and decision support, yet their inability to forget creates serious cybersecurity, privacy, and safety risks. Sensitive personal information, copyrighted material, hazardous domain knowledge, and memorized training data remain encoded across billions of parameters long after deployment, leaving models vulnerable to extraction, jailbreak attacks, membership inference, and regulatory non-compliance. Real-world incidents, from chatbots regenerating private information to fabricated legal citations producing direct legal and financial cost, place the problem at the center of the emerging-threats landscape rather than the realm of speculation. Because retraining billion-parameter models on revised corpora is computationally infeasible, and because knowledge within an LLM is distributed and entangled across parameters rather than localized to identifiable units, LLM unlearning has emerged as the principal cyber defense response, aiming to remove or suppress targeted knowledge from a trained model without retraining and without eroding what the model should still know. A central question, however, remains unresolved. Do current methods genuinely remove knowledge, or do they only stop the model from expressing it under ordinary prompting conditions? This survey examines LLM unlearning through the lens of security, robustness, and verifiable forgetting, with primary focus on gradient-based methods, which have come to dominate the field due to their compatibility with existing training pipelines and their scalability to billion-parameter models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。