arXiv:2511.09855cs.LG2025-11被引 2

提出可验证的遗忘机制,让大模型安全删除隐私数据。

Unlearning Imperative: Securing Trustworthy and Responsible LLMs through Engineered Forgetting

  • 结合差分隐私与加密技术实现可验证遗忘
  • 现有方法仍难防对抗性数据恢复,需更强防护
  • 适合关注数据合规与可信AI的开发者与监管者

大语言模型在敏感领域应用日益广泛,但其无法确保私密信息被永久删除,成为关键短板。重新训练成本过高,现有遗忘方法分散、难以验证且易被恢复。本文综述了大模型机器遗忘研究进展,分析遗忘效果评估方法、模型对对抗攻击的鲁棒性,以及在模型复杂或专有性限制透明度时如何建立用户信任。探讨了差分隐私、同态加密、联邦学习和临时记忆等技术方案,同时评估审计实践与监管框架的协同作用。研究发现虽有进展,但可靠且可验证的遗忘机制仍未解决。未来需发展无需重训练的高效技术、强化对抗恢复防御,并建立问责治理结构,以保障大模型在敏感场景中的安全部署。本研究整合技术与制度视角,勾勒出可被强制遗忘的可信AI路径。

原文摘要 · Abstract (English)

The growing use of large language models in sensitive domains has exposed a critical weakness: the inability to ensure that private information can be permanently forgotten. Yet these systems still lack reliable mechanisms to guarantee that sensitive information can be permanently removed once it has been used. Retraining from the beginning is prohibitively costly, and existing unlearning methods remain fragmented, difficult to verify, and often vulnerable to recovery. This paper surveys recent research on machine unlearning for LLMs and considers how far current approaches can address these challenges. We review methods for evaluating whether forgetting has occurred, the resilience of unlearned models against adversarial attacks, and mechanisms that can support user trust when model complexity or proprietary limits restrict transparency. Technical solutions such as differential privacy, homomorphic encryption, federated learning, and ephemeral memory are examined alongside institutional safeguards including auditing practices and regulatory frameworks. The review finds steady progress, but robust and verifiable unlearning is still unresolved. Efficient techniques that avoid costly retraining, stronger defenses against adversarial recovery, and governance structures that reinforce accountability are needed if LLMs are to be deployed safely in sensitive applications. By integrating technical and organizational perspectives, this study outlines a pathway toward AI systems that can be required to forget, while maintaining both privacy and public trust.

大模型遗忘机制隐私保护可信AI

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。