arXiv:2502.19070cs.LGcs.CR2025-02AAAI被引 6

提出新评估指标与生成框架,提升模型反演攻击对单样本隐私的还原能力。

A Sample-Level Evaluation and Generative Framework for Model Inversion Attacks

  • 设计复合评分指标DDCS,精准衡量单个训练样本的重建质量。
  • 发现多数训练样本仍能抵抗先进攻击,验证隐私保护有效性。
  • 融合熵损失与自然梯度,增强攻击生成能力,适用于隐私研究者。

模型反演(MI)攻击可重构神经网络的训练数据,引发重大隐私担忧。现有攻击多聚焦标签级隐私,而对单个训练样本的隐私保护研究不足,主要受限于评估指标缺陷。本文提出针对训练样本分析的新型指标——多样性与距离复合评分(DDCS),综合多种攻击属性评估每个样本的重建保真度,显著提升样本级隐私评估精度。基于DDCS,我们发现即使在最先进攻击下,许多训练样本仍具韧性。为此,进一步提出一种迁移学习框架,通过引入熵损失与自然梯度下降,增强攻击者生成能力。大量实验验证该框架在DDCS、覆盖率和FID等指标上均优于现有方法。最后,证明DDCS可用于无监督识别易受攻击样本,具备防御应用潜力。

原文摘要 · Abstract (English)

Model Inversion (MI) attacks, which reconstruct the training dataset of neural networks, pose significant privacy concerns in machine learning. Recent MI attacks have managed to reconstruct realistic label-level private data, such as the general appearance of a target person from all training images labeled on him. Beyond label-level privacy, in this paper we show sample-level privacy, the private information of a single target sample, is also important but under-explored in the MI literature due to the limitations of existing evaluation metrics. To address this gap, this study introduces a novel metric tailored for training-sample analysis, namely, the Diversity and Distance Composite Score (DDCS), which evaluates the reconstruction fidelity of each training sample by encompassing various MI attack attributes. This, in turn, enhances the precision of sample-level privacy assessments. Leveraging DDCS as a new evaluative lens, we observe that many training samples remain resilient against even the most advanced MI attack. As such, we further propose a transfer learning framework that augments the generative capabilities of MI attackers through the integration of entropy loss and natural gradient descent. Extensive experiments verify the effectiveness of our framework on improving state-of-the-art MI attacks over various metrics including DDCS, coverage and FID. Finally, we demonstrate that DDCS can also be useful for MI defense, by identifying samples susceptible to MI attacks in an unsupervised manner.

模型反演隐私评估生成对抗安全防御

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。