系统梳理模型反演攻击与防御方法,揭示深度学习隐私风险。
Model Inversion Attacks: A Survey of Approaches and Countermeasures
- 从攻击者视角出发,分析不同接口和假设下的反演策略
- 总结图像、文本、图数据等多场景下攻击成功率与重建质量
- 适合关注隐私安全的开发者、研究人员参考
深度神经网络在欧几里得数据(如图像、文本)和非欧几里得数据(如图)上推动了众多研究与应用。由于这些网络可能处理敏感数据,其部署引发隐私泄露担忧。模型反演攻击(MIAs)利用对训练好的模型的访问,重构训练样本或推断模型所代表的隐私敏感特征。该攻击已在图像、文本和图等领域被验证有效,凸显神经网络的脆弱性,引起学术界对隐私泄露风险的关注。本综述基于威胁模型与假设,系统梳理攻击与防御方法,比较暴露接口、攻击者知识、重建先验与目标、失败模式、部署约束及隐私-效用权衡,突出建模原则、优化挑战与未来方向。同时维护一个持续更新的研究资源库:https://github.com/AndrewZhou924/Awesome-model-inversion-attack。
原文摘要 · Abstract (English)
Deep neural networks have enabled numerous studies and applications on both Euclidean data, such as images and text, and non-Euclidean data, such as graphs. Because these networks may process private data, their deployment raises concerns about privacy leakage. Model inversion attacks (MIAs) exploit access to a trained model to reconstruct training examples or infer privacy-sensitive characteristics represented by the model. The effectiveness of MIAs has been demonstrated in various domains, including images, text, and graphs. These attacks highlight the vulnerability of neural networks and raise awareness about the risk of privacy leakage within the research community. This survey provides a threat-model- and assumption-aware synthesis of attacks and defenses. We compare exposed interfaces, attacker knowledge, reconstruction priors and targets, failure modes, deployment constraints, and privacy-utility trade-offs, while highlighting modeling principles, optimization challenges, and future directions. We also maintain an evolving repository of relevant research at https://github.com/AndrewZhou924/Awesome-model-inversion-attack.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。