arXiv:2508.21654cs.CRcs.LG2025-08被引 1

首次系统化设计模型窃取攻击评估方法,助力保护AI模型知识产权。

I Stolenly Swear That I Am Up to (No) Good: Design and Evaluation of Model Stealing Attacks

  • 构建首个通用威胁模型与攻击对比框架,统一评估标准。
  • 分析多篇相关研究,发现图像分类模型是主要攻击目标。
  • 提出攻击前、中、后全流程最佳实践,推动领域发展。

模型窃取攻击威胁以服务形式提供的机器学习模型的机密性。尽管模型本身保密,恶意方仍可通过查询模型对数据样本打标签,并训练自己的替代模型,从而侵犯知识产权。尽管该领域不断有新攻击提出,但其设计与评估缺乏标准化,难以比较成果或评估进展。本文首次填补这一空白,提出模型窃取攻击的设计与评估建议。我们聚焦于依赖训练替代模型的一类主流攻击——针对图像分类模型的攻击,提出首个全面的威胁模型并构建攻击对比框架。通过分析相关研究中的攻击设置,揭示当前最常被研究的任务与模型类型。基于此,我们提出攻击开发全过程(实验前、中、后)的最佳实践,并列出大量开放研究问题。本研究的发现与建议可推广至其他领域,确立首个通用的模型窃取攻击评估方法。

原文摘要 · Abstract (English)

Model stealing attacks endanger the confidentiality of machine learning models offered as a service. Although these models are kept secret, a malicious party can query a model to label data samples and train their own substitute model, violating intellectual property. While novel attacks in the field are continually being published, their design and evaluations are not standardised, making it challenging to compare prior works and assess progress in the field. This paper is the first to address this gap by providing recommendations for designing and evaluating model stealing attacks. To this end, we study the largest group of attacks that rely on training a substitute model -- those attacking image classification models. We propose the first comprehensive threat model and develop a framework for attack comparison. Further, we analyse attack setups from related works to understand which tasks and models have been studied the most. Based on our findings, we present best practices for attack development before, during, and beyond experiments and derive an extensive list of open research questions regarding the evaluation of model stealing attacks. Our findings and recommendations also transfer to other problem domains, hence establishing the first generic evaluation methodology for model stealing attacks.

模型安全攻击评估知识产权

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。