arXiv:2509.20177cs.LG2025-09NeurIPS被引 1

揭示生成式模型逆向攻击高效背后的几何原理

Generative Model Inversion Through the Lens of the Manifold Hypothesis

  • 通过分析梯度方向,发现生成器流形能自动滤除噪声
  • 梯度与数据流形对齐程度越高,攻击越有效
  • 无需训练即可提升逆向攻击效果,适合安全研究者

模型逆向攻击(MIA)旨在从训练好的模型中重构类别代表性样本。近期基于生成对抗网络的生成式逆向攻击通过学习图像先验来引导重构过程,生成视觉质量高且与私有训练数据高度一致的结果。为探究其有效性原因,我们分析了逆向损失关于合成输入的梯度,发现这些梯度异常嘈杂。进一步分析表明,生成式逆向攻击通过将梯度投影到生成器流形的切空间,隐式地实现了去噪:保留与流形对齐的信息方向,同时过滤掉离流形的分量。实证测量显示,在标准监督训练的模型中,损失梯度常与数据流形存在显著角度偏差,表明其与类别相关方向对齐不佳。据此提出核心假设:当损失梯度更贴近生成器流形时,模型对逆向攻击更加脆弱。我们设计了一种新训练目标以显式促进该对齐,并在此基础上提出一种无需训练的增强方法,在逆向过程中提升梯度-流形对齐度,从而在多个基准上持续优于现有最先进生成式逆向攻击方法。

原文摘要 · Abstract (English)

Model inversion attacks (MIAs) aim to reconstruct class-representative samples from trained models. Recent generative MIAs utilize generative adversarial networks to learn image priors that guide the inversion process, yielding reconstructions with high visual quality and strong fidelity to the private training data. To explore the reason behind their effectiveness, we begin by examining the gradients of inversion loss with respect to synthetic inputs, and find that these gradients are surprisingly noisy. Further analysis reveals that generative inversion implicitly denoises these gradients by projecting them onto the tangent space of the generator manifold, filtering out off-manifold components while preserving informative directions aligned with the manifold. Our empirical measurements show that, in models trained with standard supervision, loss gradients often exhibit large angular deviations from the data manifold, indicating poor alignment with class-relevant directions. This observation motivates our central hypothesis: models become more vulnerable to MIAs when their loss gradients align more closely with the generator manifold. We validate this hypothesis by designing a novel training objective that explicitly promotes such alignment. Building on this insight, we further introduce a training-free approach to enhance gradient-manifold alignment during inversion, leading to consistent improvements over state-of-the-art generative MIAs.

模型逆向生成模型安全攻防流形假设

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。