用模型重编程放大隐私泄漏,高效实现成员推理攻击
ReproMIA: A Comprehensive Analysis of Model Reprogramming for Proactive Membership Inference Attacks
- 通过模型重编程主动放大隐私痕迹,无需训练影子模型
- 在低误报率下性能显著提升,大模型与扩散模型均超越基线
- 适用于大模型、扩散模型等多类架构,适合隐私安全研究者
深度学习模型在关键领域的广泛应用加剧了因数据记忆带来的隐私风险。尽管成员推理攻击(MIAs)是评估此类隐私漏洞的金标准,但传统方法受限于影子模型训练的高昂计算成本,且在低误报率约束下性能急剧下降。为此,本文提出一种新思路:利用模型重编程作为主动信号放大器以增强隐私泄露。基于此,我们构建了统一高效的前瞻性框架 exttt{ReproMIA},理论与实证均证明其可主动诱导并放大模型表征中隐含的隐私足迹。我们在多种架构(包括大语言模型、扩散模型和分类模型)上实现 exttt{ReproMIA} 的定制化版本。在十余个基准测试及多种模型上的全面评估显示, exttt{ReproMIA} 持续且显著优于现有最优方法,在低误报率场景下表现尤为突出:大模型平均提升 5.25\\% AUC 与 10.68\\" TPR@1\\"FPR;扩散模型分别提升 3.70\\" 和 12.40\\"。
原文摘要 · Abstract (English)
The pervasive deployment of deep learning models across critical domains has concurrently intensified privacy concerns due to their inherent propensity for data memorization. While Membership Inference Attacks (MIAs) serve as the gold standard for auditing these privacy vulnerabilities, conventional MIA paradigms are increasingly constrained by the prohibitive computational costs of shadow model training and a precipitous performance degradation under low False Positive Rate constraints. To overcome these challenges, we introduce a novel perspective by leveraging the principles of model reprogramming as an active signal amplifier for privacy leakage. Building upon this insight, we present \texttt{ReproMIA}, a unified and efficient proactive framework for membership inference. We rigorously substantiate, both theoretically and empirically, how our methodology proactively induces and magnifies latent privacy footprints embedded within the model's representations. We provide specialized instantiations of \texttt{ReproMIA} across diverse architectural paradigms, including LLMs, Diffusion Models, and Classification Models. Comprehensive experimental evaluations across more than ten benchmarks and a variety of model architectures demonstrate that \texttt{ReproMIA} consistently and substantially outperforms existing state-of-the-art baselines, achieving a transformative leap in performance specifically within low-FPR regimes, such as an average of 5.25\% AUC and 10.68\% TPR@1\%FPR increase over the runner-up for LLMs, as well as 3.70\% and 12.40\% respectively for Diffusion Models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。