攻击者仅凭学生模型就能推断教师模型的训练数据,暴露迁移学习隐私漏洞。
Rethinking Membership Inference Attacks Against Transfer Learning
- 通过分析师生模型隐藏层差异,构建针对迁移学习的新型成员推理攻击
- 在4个数据集上验证成功,即使仅访问学生模型也能推断教师数据
- 揭示迁移学习中未被充分关注的隐私风险,适合研究模型安全者参考
迁移学习在跨任务知识迁移中表现优异,但面临成员推理攻击(MIAs)带来的重大隐私威胁。尽管此类攻击对模型训练数据构成显著风险,但在迁移学习领域仍缺乏深入研究。师生模型间的交互关系尚未在成员推理攻击中得到充分探索,可能遗漏了迁移学习中的隐私漏洞。本文提出一种新的迁移学习成员推理攻击方法,在白盒环境下仅访问学生模型,即可判断特定数据点是否用于训练教师模型。该方法深入分析学生模型与其影子模型在隐藏层表示上的差异,利用这些差异优化影子模型训练并辅助成员推理决策。在四个不同迁移学习任务的数据集上评估表明,即便攻击者仅能访问学生模型,教师模型的训练数据仍可能被成功推断。本工作揭示了迁移学习中未被充分认识的成员推理风险。
原文摘要 · Abstract (English)
Transfer learning, successful in knowledge translation across related tasks, faces a substantial privacy threat from membership inference attacks (MIAs). These attacks, despite posing significant risk to ML model's training data, remain limited-explored in transfer learning. The interaction between teacher and student models in transfer learning has not been thoroughly explored in MIAs, potentially resulting in an under-examined aspect of privacy vulnerabilities within transfer learning. In this paper, we propose a new MIA vector against transfer learning, to determine whether a specific data point was used to train the teacher model while only accessing the student model in a white-box setting. Our method delves into the intricate relationship between teacher and student models, analyzing the discrepancies in hidden layer representations between the student model and its shadow counterpart. These identified differences are then adeptly utilized to refine the shadow model's training process and to inform membership inference decisions effectively. Our method, evaluated across four datasets in diverse transfer learning tasks, reveals that even when an attacker only has access to the student model, the teacher model's training data remains susceptible to MIAs. We believe our work unveils the unexplored risk of membership inference in transfer learning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。