无需重训练即可快速删除数据,且操作可验证。
TrustErase: Auditable Instant Machine Unlearning with Passport-Embedded Representations

- 用嵌入式密钥在模型中实现即刻遗忘。
- 在多个数据集上表现优于或媲美现有顶尖方法。
- 适合需要合规审计的隐私敏感场景。
隐私合规人工智能的需求推动了机器遗忘的发展;然而,现有的重训练或知识蒸馏方法难以验证且计算成本高。我们提出 TrustErase,一种可验证、无数据的遗忘框架,利用嵌入式护照表示实现即时、模块化和可审计的遗忘。通过将护照视为参数高效适配层中的加密密钥,TrustErase 可通过简单禁用实现特定类别或数据集的删除,无需重训练、微调或原始数据访问。基于奇异值分解的方法将护照隐藏于模型权重中,确保遗忘操作透明且可证明合规。在 MNIST、CIFAR10 和 CIFAR100 上的评估表明,TrustErase 在严格无数据条件下达到或超过 DELETE、L2UL 与 Boundary Shrink 等先进基准的表现。最终,TrustErase 建立了可信任、可问责、即时可遗忘的人工智能新范式。
原文摘要 · Abstract (English)
The demand for privacy-compliant AI has amplified the need for machine unlearning; yet, existing retraining or distillation-based methods remain unverifiable and computationally costly. We introduce TrustErase, a verifiable, data-free unlearning framework leveraging passport-embedded representations for instant, modular, and auditable forgetting. By treating passports as cryptographic keys within parameter-efficient adaptation layers, TrustErase enables the removal of specific classes or datasets through simple deactivation, without retraining, fine-tuning, or access to the original data. A singular value based decomposition conceals passports within model weights, ensuring that unlearning actions remain transparent and provably compliant. Evaluations on MNIST, CIFAR10 and CIFAR100 show that TrustErase matches or exceeds state-of-the-art benchmarks such as DELETE, L2UL, and Boundary Shrink, while operating in a strictly data-free regime. Ultimately, TrustErase establishes a new paradigm for trustworthy, accountable, and instantly forgettable AI systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。