arXiv:2508.18502cs.LGcs.AI2025-08中稿 · SIBGRAPI'25被引 1

数据增强能显著提升模型删忘效果,降低隐私泄露风险。

Data Augmentation Improves Machine Unlearning

  • 设计合理的数据增强策略可提升删忘方法性能。
  • 使用TrivialAug增强后,平均差距指标降低40.12%。
  • 适合关注模型隐私保护与高效删忘的研究者。

机器删忘(MU)旨在移除特定数据对已训练模型的影响,同时保持模型在剩余数据上的性能。尽管已有研究指出记忆性与数据增强之间的关联,但系统性增强设计在删忘中的作用仍缺乏深入探讨。本文研究了不同数据增强策略对删忘方法(包括SalUn、随机标签和微调)的影响。在CIFAR-10和CIFAR-100上,于不同遗忘率条件下进行实验,结果表明合理设计的增强策略能显著提升删忘有效性,缩小与重新训练模型之间的性能差距。使用TrivialAug增强时,平均差距指标最多降低40.12%。结果表明,数据增强不仅能减少记忆现象,还在实现隐私保护与高效删忘中起关键作用。

原文摘要 · Abstract (English)

Machine Unlearning (MU) aims to remove the influence of specific data from a trained model while preserving its performance on the remaining data. Although a few works suggest connections between memorisation and augmentation, the role of systematic augmentation design in MU remains under-investigated. In this work, we investigate the impact of different data augmentation strategies on the performance of unlearning methods, including SalUn, Random Label, and Fine-Tuning. Experiments conducted on CIFAR-10 and CIFAR-100, under varying forget rates, show that proper augmentation design can significantly improve unlearning effectiveness, reducing the performance gap to retrained models. Results showed a reduction of up to 40.12% of the Average Gap unlearning Metric, when using TrivialAug augmentation. Our results suggest that augmentation not only helps reduce memorization but also plays a crucial role in achieving privacy-preserving and efficient unlearning.

机器删忘数据增强隐私保护

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。