首次揭示扩散语言模型的隐私泄露风险,提出高效会员攻击方法。
Membership Inference Attacks Against Fine-tuned Diffusion Language Models
- 设计新攻击方法SAMA,通过聚合不同掩码密度的信号提升检测能力。
- 在9个数据集上比最优基线提升30%相对AUC,低误报率下最高提升8倍。
- 适用于研究模型隐私或构建防御的开发者,尤其关注扩散模型安全者。
扩散语言模型(DLMs)作为自回归模型的有前景替代方案,采用双向掩码标记预测。然而其在会员推理攻击(MIA)下的隐私泄露风险尚未得到充分研究。本文首次系统性探究了DLMs的MIA脆弱性。与自回归模型固定的单一预测模式不同,DLMs的多重可掩码配置呈指数级增加攻击机会。这种对多个独立掩码的探测能力显著提升了检测概率。为此,我们提出SAMA(子集聚合会员攻击),通过鲁棒聚合解决信号稀疏问题。SAMA在渐进掩码密度下采样掩码子集,并采用在重尾噪声下仍有效的符号统计方法。通过逆权重聚合优先考虑稀疏掩码的清洁信号,SAMA将稀疏记忆检测转化为鲁棒投票机制。在九个数据集上的实验表明,SAMA相比最佳基线实现30%的相对AUC提升,且在低误报率下最高可达8倍提升。这些发现揭示了DLMs中此前未知的重大隐私漏洞,亟需开发针对性的隐私防护措施。
原文摘要 · Abstract (English)
Diffusion Language Models (DLMs) represent a promising alternative to autoregressive language models, using bidirectional masked token prediction. Yet their susceptibility to privacy leakage via Membership Inference Attacks (MIA) remains critically underexplored. This paper presents the first systematic investigation of MIA vulnerabilities in DLMs. Unlike the autoregressive models' single fixed prediction pattern, DLMs' multiple maskable configurations exponentially increase attack opportunities. This ability to probe many independent masks dramatically improves detection chances. To exploit this, we introduce SAMA (Subset-Aggregated Membership Attack), which addresses the sparse signal challenge through robust aggregation. SAMA samples masked subsets across progressive densities and applies sign-based statistics that remain effective despite heavy-tailed noise. Through inverse-weighted aggregation prioritizing sparse masks' cleaner signals, SAMA transforms sparse memorization detection into a robust voting mechanism. Experiments on nine datasets show SAMA achieves 30% relative AUC improvement over the best baseline, with up to 8 times improvement at low false positive rates. These findings reveal significant, previously unknown vulnerabilities in DLMs, necessitating the development of tailored privacy defenses.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。