提出高效模型提取方法E3,用极少查询实现更高精度复制。
Efficient and Effective Model Extraction
- 优化查询准备与训练流程,设计简单但高效
- 仅需0.5%查询量,准确率比现有方法高50%以上
- 适合评估MLaaS系统安全,可作基准测试工具
模型提取旨在通过机器学习即服务(MLaaS)API以最小开销创建功能相似的副本,通常用于非法获利或作为后续攻击的前奏,对MLaaS生态构成严重威胁。然而,近期研究显示,当目标任务分布不可知时,模型提取效率极低,即使大幅增加攻击预算也难以生成足够相似的复制品,削弱了攻击者动机。本文重新审视提取生命周期中的基本设计选择,提出一种简单却极为有效的算法——高效且有效的模型提取(E3),重点优化查询准备与训练过程。E3在保持极低计算成本的同时,显著提升泛化能力。例如,在CIFAR-10上,仅需0.005倍查询预算和不到0.2倍运行时间,其准确率绝对提升超过50%,远超基于生成模型的数据无关模型提取方法。研究结果表明模型提取仍是持续威胁,并建议其可作为未来安全评估的基准算法。
原文摘要 · Abstract (English)
Model extraction aims to create a functionally similar copy from a machine learning as a service (MLaaS) API with minimal overhead, typically for illicit profit or as a precursor to further attacks, posing a significant threat to the MLaaS ecosystem. However, recent studies have shown that model extraction is highly inefficient, particularly when the target task distribution is unavailable. In such cases, even substantially increasing the attack budget fails to produce a sufficiently similar replica, reducing the adversary's motivation to pursue extraction attacks. In this paper, we revisit the elementary design choices throughout the extraction lifecycle. We propose an embarrassingly simple yet dramatically effective algorithm, Efficient and Effective Model Extraction (E3), focusing on both query preparation and training routine. E3 achieves superior generalization compared to state-of-the-art methods while minimizing computational costs. For instance, with only 0.005 times the query budget and less than 0.2 times the runtime, E3 outperforms classical generative model based data-free model extraction by an absolute accuracy improvement of over 50% on CIFAR-10. Our findings underscore the persistent threat posed by model extraction and suggest that it could serve as a valuable benchmarking algorithm for future security evaluations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。