通过边界采样大幅降低模型窃取的查询次数,提升窃取模型精度与攻击效果。
Efficient Model Extraction via Boundary Sampling
- 聚焦决策边界低置信区域采样,结合进化算法优化过程。
- 查询次数减少10至600倍,窃取模型准确率更高。
- 无需目标模型架构或数据信息,适合黑盒攻击场景。
本文提出一种新型无数据模型窃取攻击,显著提升效率、准确率与有效性。传统黑盒方法依赖目标模型作为标注器,在高置信区域生成大量样本,不仅需海量查询,且窃取模型准确性与可迁移性较低。本方法创新性地在决策边界附近的低置信区域采样,并采用进化算法优化采样过程,使攻击者查询次数减少10至600倍,同时提升窃取模型精度。此外,该方法增强边界对齐,使窃取模型生成的对抗样本对目标模型的攻击成功率从平均60%提升至82%。所有操作均基于严格黑盒假设,无需知晓目标模型架构或训练数据。我们在三个分辨率递增的数据集上验证攻击效果,并与四种先进模型窃取攻击对比。
原文摘要 · Abstract (English)
This paper introduces a novel data-free model extraction attack that significantly advances the current state-of-the-art in terms of efficiency, accuracy, and effectiveness. Traditional black-box methods rely on using the victim's model as an oracle to label a vast number of samples within high-confidence areas. This approach not only requires an extensive number of queries but also results in a less accurate and less transferable model. In contrast, our method innovates by focusing on sampling low-confidence areas (along the decision boundaries) and employing an evolutionary algorithm to optimize the sampling process. These novel contributions allow for a dramatic reduction in the number of queries needed by the attacker by a factor of 10x to 600x while simultaneously improving the accuracy of the stolen model. Moreover, our approach improves boundary alignment, resulting in better transferability of adversarial examples from the stolen model to the victim's model (increasing the attack success rate from 60\% to 82\% on average). Finally, we accomplish all of this with a strict black-box assumption on the victim, with no knowledge of the target's architecture or dataset. We demonstrate our attack on three datasets with increasingly larger resolutions and compare our performance to four state-of-the-art model extraction attacks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。