提出新方法破解多领域图模型隐私漏洞,验证其成员信息易泄露。
Privacy Auditing of Multi-domain Graph Pre-trained Model under Membership Inference Attacks
- 通过机器遗忘增强目标模型过拟合特征,放大隐私信号。
- 仅用少量影子图构建可靠影子模型,提升攻击可行性。
- 基于相似性比对正负样本,精准识别训练成员身份。
多领域图预训练已成为构建图基础模型的关键技术,虽显著提升图神经网络的泛化能力,但其在成员推断攻击(MIAs)下的隐私风险尚未充分研究。由于:(i) 多领域预训练削弱了过拟合特征,使传统攻击失效;(ii) 训练图多样性导致难以获取代表性影子数据集;(iii) 基于嵌入的输出信息量低于原始logits,导致成员信号弱化。针对上述挑战,本文提出MGP-MIA框架:首先设计成员信号放大机制,通过机器遗忘强化目标模型的过拟合特征;其次提出增量式影子模型构建方法,利用有限影子图实现可靠影子模型;最后引入基于相似性的推理机制,依据样本与正负例的相似度判断成员身份。大量实验验证了MGP-MIA的有效性,并揭示了多领域图预训练模型的隐私隐患。
原文摘要 · Abstract (English)
Multi-domain graph pre-training has emerged as a pivotal technique in developing graph foundation models. While it greatly improves the generalization of graph neural networks, its privacy risks under membership inference attacks (MIAs), which aim to identify whether a specific instance was used in training (member), remain largely unexplored. However, effectively conducting MIAs against multi-domain graph pre-trained models is a significant challenge due to: (i) Enhanced Generalization Capability: Multi-domain pre-training reduces the overfitting characteristics commonly exploited by MIAs. (ii) Unrepresentative Shadow Datasets: Diverse training graphs hinder the obtaining of reliable shadow graphs. (iii) Weakened Membership Signals: Embedding-based outputs offer less informative cues than logits for MIAs. To tackle these challenges, we propose MGP-MIA, a novel framework for Membership Inference Attacks against Multi-domain Graph Pre-trained models. Specifically, we first propose a membership signal amplification mechanism that amplifies the overfitting characteristics of target models via machine unlearning. We then design an incremental shadow model construction mechanism that builds a reliable shadow model with limited shadow graphs via incremental learning. Finally, we introduce a similarity-based inference mechanism that identifies members based on their similarity to positive and negative samples. Extensive experiments demonstrate the effectiveness of our proposed MGP-MIA and reveal the privacy risks of multi-domain graph pre-training.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。