arXiv:2412.08544cs.LGcs.CR2024-12CVPR被引 3

用优化方法重构训练数据,发现初始值影响结果真实性。

Training Data Reconstruction: Privacy due to Uncertainty?

  • 将重建问题建模为双层优化,提升重构稳定性。
  • 随机初始化常生成看似真实但非真实训练样本的图像。
  • 对隐私保护研究有启发,适合关注模型安全的读者。

从神经网络参数中重构训练数据是重大隐私隐患。已有研究证明在特定条件下可实现重构。本文通过实证分析,提出将重构问题重新表述为双层优化问题。实验表明,无论是新方法还是以往方法,其结果都高度依赖待重构图像 $x$ 的初始化。特别地,使用随机初始化时,可生成外观类似真实训练样本但实际不在训练集中的图像。针对仿射网络和单隐藏层网络的实验显示,攻击者无法确定重构图像是否曾出现在原始训练集中,这暗示了当前重构技术在判定数据归属方面存在不确定性。

原文摘要 · Abstract (English)

Being able to reconstruct training data from the parameters of a neural network is a major privacy concern. Previous works have shown that reconstructing training data, under certain circumstances, is possible. In this work, we analyse such reconstructions empirically and propose a new formulation of the reconstruction as a solution to a bilevel optimisation problem. We demonstrate that our formulation as well as previous approaches highly depend on the initialisation of the training images $x$ to reconstruct. In particular, we show that a random initialisation of $x$ can lead to reconstructions that resemble valid training samples while not being part of the actual training dataset. Thus, our experiments on affine and one-hidden layer networks suggest that when reconstructing natural images, yet an adversary cannot identify whether reconstructed images have indeed been part of the set of training samples.

隐私安全数据重构神经网络

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。