用梯度特征和假预测防御无数据模型盗取,有效防住多种攻击。
Model-Guardian: Protecting against Data-Free Model Stealing Using Gradient Representations and Deceptive Predictions
- 利用合成样本的梯度特征检测异常查询行为
- 在7种攻击下准确率超11种现有防御方法,达新基准
- 适合云上部署模型的保密防护,尤其对抗先进生成模型攻击
模型盗取攻击正日益威胁云端部署机器学习模型的机密性。近期研究发现,攻击者可利用数据合成技术,在无真实数据场景下盗取模型,形成无数据模型盗取攻击。现有防御方法存在效果差、泛化能力弱、覆盖不全等问题。为此,本文提出新型防御框架 Model-Guardian,包含两个组件:无数据模型盗取检测器(DFMS-Detector)与假预测(DPreds)。该框架借助合成样本的特性及样本梯度表示,有效弥补现有防御缺陷。在七种主流无数据模型盗取攻击上的大量实验表明,Model-Guardian 具有显著优越的防御效果与泛化能力,优于十一种现有防御方法,达到新基准。值得注意的是,本文首次采用多种 GAN 与扩散模型生成高度逼真的查询样本进行攻击,而 Model-Guardian 仍能保持精准检测能力。
原文摘要 · Abstract (English)
Model stealing attack is increasingly threatening the confidentiality of machine learning models deployed in the cloud. Recent studies reveal that adversaries can exploit data synthesis techniques to steal machine learning models even in scenarios devoid of real data, leading to data-free model stealing attacks. Existing defenses against such attacks suffer from limitations, including poor effectiveness, insufficient generalization ability, and low comprehensiveness. In response, this paper introduces a novel defense framework named Model-Guardian. Comprising two components, Data-Free Model Stealing Detector (DFMS-Detector) and Deceptive Predictions (DPreds), Model-Guardian is designed to address the shortcomings of current defenses with the help of the artifact properties of synthetic samples and gradient representations of samples. Extensive experiments on seven prevalent data-free model stealing attacks showcase the effectiveness and superior generalization ability of Model-Guardian, outperforming eleven defense methods and establishing a new state-of-the-art performance. Notably, this work pioneers the utilization of various GANs and diffusion models for generating highly realistic query samples in attacks, with Model-Guardian demonstrating accurate detection capabilities.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。