arXiv:2409.01062cs.LGcs.CR2024-09中稿 · Transactions on Ma…被引 4

随机擦除可有效防御模型逆向攻击,且不损害模型性能。

Random Erasing vs. Model Inversion: A Promising Defense or a False Hope?

  • 用随机擦除训练模型,使逆向重建图像特征偏离真实数据。
  • 在37种配置下实现最佳隐私-效用平衡,攻击准确率显著下降。
  • 无需修改模型结构,可无缝集成现有隐私保护方法。

模型逆向(MI)攻击可通过机器学习模型重建私有训练数据,构成严重隐私威胁。现有防御多聚焦于模型层面,而数据对MI鲁棒性的影响仍待探索。本文研究传统用于提升模型抗遮挡泛化能力的随机擦除(RE),发现其在防御MI攻击方面出人意料地有效。新颖的特征空间分析表明,使用RE训练的模型中,MI重构图像的特征与真实私有数据特征存在显著差异,而真实数据特征仍保持类内紧凑、类间分离。这一特性共同导致重构质量下降,攻击准确率降低,同时维持合理的自然准确性。进一步分析显示,部分擦除和随机位置擦除是关键因素:前者阻止模型完整观测物体,后者实现强隐私-效用权衡。大量实验在37个设置中验证,该方法在隐私-效用权衡上达到当前最优,对多种攻击类型、网络架构和配置均表现优越。首次在某些配置下实现攻击准确率显著下降而不牺牲效用。

原文摘要 · Abstract (English)

Model Inversion (MI) attacks pose a significant privacy threat by reconstructing private training data from machine learning models. While existing defenses primarily concentrate on model-centric approaches, the impact of data on MI robustness remains largely unexplored. In this work, we explore Random Erasing (RE), a technique traditionally used for improving model generalization under occlusion, and uncover its surprising effectiveness as a defense against MI attacks. Specifically, our novel feature space analysis shows that models trained with RE-images introduce a significant discrepancy between the features of MI-reconstructed images and those of the private data. At the same time, features of private images remain distinct from other classes and well-separated from different classification regions. These effects collectively degrade MI reconstruction quality and attack accuracy while maintaining reasonable natural accuracy. Furthermore, we explore two critical properties of RE including Partial Erasure and Random Location. Partial Erasure prevents the model from observing entire objects during training. We find this has a significant impact on MI, which aims to reconstruct the entire objects. Random Location of erasure plays a crucial role in achieving a strong privacy-utility trade-off. Our findings highlight RE as a simple yet effective defense mechanism that can be easily integrated with existing privacy-preserving techniques. Extensive experiments across 37 setups demonstrate that our method achieves state-of-the-art (SOTA) performance in the privacy-utility trade-off. The results consistently demonstrate the superiority of our defense over existing methods across different MI attacks, network architectures, and attack configurations. For the first time, we achieve a significant degradation in attack accuracy without a decrease in utility for some configurations.

模型逆向隐私保护随机擦除数据增强

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。