重新评估模型盗取攻击,发现目标模型性能越强,攻击越精准。
Attackers Can Do Better: Over- and Understated Factors of Model Stealing Attacks
- 通过系统实验分析攻击者能力对盗取效果的影响
- 强目标模型能支持更高保真度的攻击,且数据复杂度比模型复杂度更重要
- 即使无数据也能高效攻击,提示现有防御可能被低估
机器学习模型易受模型盗取攻击,导致知识产权泄露。其中,替代模型训练是一种适用于任何可通过输入输出查询近似行为的模型的通用攻击方法。以往研究多聚焦于提升替代模型性能,如设计新训练方法,但对攻击者能力与知识差异如何影响攻击效果的剖析极为有限,导致不同研究结论矛盾。本文全面考察了攻击者能力与知识变化带来的多重因素影响。结果表明,过去认为重要的若干因素实际影响较小;我们发现攻击成功率与目标模型性能存在新关联——性能越好的目标模型可实现更高保真度的攻击,并解释其内在机制。进一步建议将关注点从目标模型复杂度转向其学习任务复杂度:对替代模型而言,应优先获取更复杂的训练数据并选择合适架构,而非盲目增加模型复杂度。最后,我们证明在完全无数据的极端场景下,无需以数百万查询来弥补知识不足。实验结果常优于或媲美以往假设更强攻击者的方案,暗示当前攻击可能比已知程度更严重地威胁模型所有者的知识产权。
原文摘要 · Abstract (English)
Machine learning models were shown to be vulnerable to model stealing attacks, which lead to intellectual property infringement. Among other methods, substitute model training is an all-encompassing attack applicable to any machine learning model whose behaviour can be approximated from input-output queries. Whereas prior works mainly focused on improving the performance of substitute models by, e.g. developing a new substitute training method, there have been only limited ablation studies on the impact the attacker's strength has on the substitute model's performance. As a result, different authors came to diverse, sometimes contradicting, conclusions. In this work, we exhaustively examine the ambivalent influence of different factors resulting from varying the attacker's capabilities and knowledge on a substitute training attack. Our findings suggest that some of the factors that have been considered important in the past are, in fact, not that influential; instead, we discover new correlations between attack conditions and success rate. In particular, we demonstrate that better-performing target models enable higher-fidelity attacks and explain the intuition behind this phenomenon. Further, we propose to shift the focus from the complexity of target models toward the complexity of their learning tasks. Therefore, for the substitute model, rather than aiming for a higher architecture complexity, we suggest focusing on getting data of higher complexity and an appropriate architecture. Finally, we demonstrate that even in the most limited data-free scenario, there is no need to overcompensate weak knowledge with millions of queries. Our results often exceed or match the performance of previous attacks that assume a stronger attacker, suggesting that these stronger attacks are likely endangering a model owner's intellectual property to a significantly higher degree than shown until now.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。