arXiv:2607.12354cs.LG2026-07

模型隐私泄露主因是对抗非鲁棒特征,而非记忆训练数据。

Reducing information dependency does not cause training data privacy. Adversarially non-robust features do

论文配图:Reducing information dependency does not cause training data privacy. Adversarially non-robust features do
图 1 · 摘自论文原文
  • 用对抗性训练诱导非鲁棒特征以提升隐私保护
  • 即使只看到3%像素仍可被重建,隐私与记忆无关
  • 适合关注模型隐私与对抗鲁棒性的研究者

本文挑战了信息依赖(包括死记硬背)导致图像重构攻击中训练数据暴露的主流观点。我们发现,即便没有死记硬背,大规模数据泄露仍会持续,其根源在于与对抗鲁棒性的可调关联。三个意外发现:(1) 抑制模型倒置攻击(MIA)的近期防御方法在理想攻击者下有效,但并未降低信息依赖度(HSIC);(2) 极度记忆训练集的模型对MIA仍保持鲁棒;(3) 在97%训练像素未被见过的模型上,理论隐私界限极强,但仍可被MIA严重重构。我们提供因果证据表明,隐私源于对抗样本文献中的‘非鲁棒特征’(可泛化但不可感知且不稳定)。近期防御通过无意间使模型趋向此类特征而提升隐私。为此,我们提出反对抗训练(AT-AT),主动学习非鲁棒特征,实现优于现有防御的重建防御与更高准确率。研究重塑了训练数据暴露的理解,并揭示新的隐私-鲁棒性权衡。

原文摘要 · Abstract (English)

In this paper, we challenge the prevailing view that information dependency (including rote memorization) drives training data exposure to image reconstruction attacks. We show that extensive exposure can persist without rote memorization and is instead caused by a tunable connection to adversarial robustness. We begin by presenting three surprising results: (1) recent defenses that inhibit reconstruction by Model Inversion Attacks (MIAs), which evaluate leakage under an idealized attacker, do not reduce standard measures of information dependency (HSIC); (2) models that maximally memorize their training datasets remain robust to MIA reconstruction; and (3) models trained without seeing 97% of the training pixels, where recent information-theoretic bounds give arbitrarily strong privacy guarantees under standard assumptions, can still be devastatingly reconstructed by MIA. To explain these findings, we provide causal evidence that privacy under MIA arises from what the adversarial examples literature calls ``non-robust'' features (generalizable but imperceptible and unstable features). We further show that recent MIA defenses obtain their privacy improvements by unintentionally shifting models toward such features. To establish this causal relationship, we introduce Anti Adversarial Training (AT-AT), a training regime that intentionally learns non-robust features to obtain both superior reconstruction defense and higher accuracy than state-of-the-art defenses. Our results revise the prevailing understanding of training data exposure and reveal a new privacy-robustness tradeoff.

模型隐私对抗鲁棒性数据泄露

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。