arXiv:2502.10635cs.LGcs.CR2025-02

提出实用机器遗忘方法,实现数据移除与模型性能平衡。

Privacy Preservation through Practical Machine Unlearning

  • 采用重训练与SISA框架实现精确遗忘
  • DaRE框架在HSpam14上保持性能但计算开销大
  • 适合关注隐私合规与伦理AI的开发者

机器学习模型依赖海量数据持续优化预测与推荐能力。在隐私问题突出的背景下,机器遗忘成为可选数据删除的关键技术。本文评估了朴素重训练与基于SISA框架的精确遗忘方法,在HSpam14数据集上的计算成本、一致性与可行性。研究探索将遗忘机制融入正未标记学习(PU Learning),以应对部分标注数据的挑战。结果表明,如DaRE等遗忘框架可在保障模型性能的同时实现隐私合规,但伴随显著计算开销。该研究强调机器遗忘对构建伦理化AI和增强数据驱动系统信任的重要性。

原文摘要 · Abstract (English)

Machine Learning models thrive on vast datasets, continuously adapting to provide accurate predictions and recommendations. However, in an era dominated by privacy concerns, Machine Unlearning emerges as a transformative approach, enabling the selective removal of data from trained models. This paper examines methods such as Naive Retraining and Exact Unlearning via the SISA framework, evaluating their Computational Costs, Consistency, and feasibility using the $\texttt{HSpam14}$ dataset. We explore the potential of integrating unlearning principles into Positive Unlabeled (PU) Learning to address challenges posed by partially labeled datasets. Our findings highlight the promise of unlearning frameworks like $\textit{DaRE}$ for ensuring privacy compliance while maintaining model performance, albeit with significant computational trade-offs. This study underscores the importance of Machine Unlearning in achieving ethical AI and fostering trust in data-driven systems.

机器遗忘隐私保护模型鲁棒性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。