提出从贝叶斯后验和训练模型中重构训练数据的数学框架。
On Reconstructing Training Data From Bayesian Posteriors and Trained Models
- 构建了用后验分布重构训练数据的数学模型
- 发现最大均值差异可刻画易被重构的数据特征
- 首次在贝叶斯与非贝叶斯模型中统一实现数据重建
公开模型结构与训练参数会带来训练数据重建攻击风险,这是现代机器学习的主要安全隐患。本文提出三项核心贡献:建立问题的数学表达框架;通过最大均值差异(MMD)等价关系刻画训练数据中易被重构的特征;提出一种适用于贝叶斯与非贝叶斯模型的得分匹配重构框架,其中贝叶斯场景的重建为文献首次实现。
原文摘要 · Abstract (English)
Publicly releasing the specification of a model with its trained parameters means an adversary can attempt to reconstruct information about the training data via training data reconstruction attacks, a major vulnerability of modern machine learning methods. This paper makes three primary contributions: establishing a mathematical framework to express the problem, characterising the features of the training data that are vulnerable via a maximum mean discrepancy equivalance and outlining a score matching framework for reconstructing data in both Bayesian and non-Bayesian models, the former is a first in the literature.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。