通过账户级特征差异检测盗取行为,保护模型不被低成本复制。
Defense Against Model Stealing Based on Account-Aware Distribution Discrepancy
- 基于账户粒度的特征分布差异检测恶意查询
- 在软硬标签场景下均实现强防御且不影响正常用户
- 可即插即用,适合部署于各类图像分类服务
恶意用户通过查询响应训练克隆模型以低成本复现商业模型功能,难以及时防御。本文提出一种非参数化检测器Account-aware Distribution Discrepancy(ADD),利用账户级局部依赖关系识别恶意查询。将每类样本建模为特征空间中的多元正态分布(MVN),以加权类间分布差异之和作为恶意得分。ADD与基于随机预测投毒的防御机制结合,形成即插即用的D-ADD防御模块,应用于图像分类模型。大量实验表明,D-ADD在软标签和硬标签设置下均能有效抵御多种攻击,对良性用户的正常服务影响极小。
原文摘要 · Abstract (English)
Malicious users attempt to replicate commercial models functionally at low cost by training a clone model with query responses. It is challenging to timely prevent such model-stealing attacks to achieve strong protection and maintain utility. In this paper, we propose a novel non-parametric detector called Account-aware Distribution Discrepancy (ADD) to recognize queries from malicious users by leveraging account-wise local dependency. We formulate each class as a Multivariate Normal distribution (MVN) in the feature space and measure the malicious score as the sum of weighted class-wise distribution discrepancy. The ADD detector is combined with random-based prediction poisoning to yield a plug-and-play defense module named D-ADD for image classification models. Results of extensive experimental studies show that D-ADD achieves strong defense against different types of attacks with little interference in serving benign users for both soft and hard-label settings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。