arXiv:2606.15493cs.LGcs.CR2026-06

高保真偷模型不等于真替代,同一任务可有多种不同表现的替代模型。

Model Stealing Through the Lens of Model Multiplicity

  • 通过计算近似等效模型集合,分析偷来的模型多样性
  • 不同替代模型在关键指标上差异显著,即使精度相似
  • 提醒开发者:模型被盗后风险不止于精度,还有部署表现差异

模型偷取攻击中,攻击者可生成高保真替代模型,对机器学习服务的知识产权构成威胁。传统观点认为这些替代模型具备与原服务相当的经济价值。本文挑战这一假设,指出基于查询的提取仅提供目标模型部分输入输出行为的监督,导致替代模型并非唯一:存在多个近似最优的替代模型,虽精度相近,但在部署相关属性上差异明显。我们不采用传统学习式偷取方法,而是计算替代模型的Rashomon集(即几乎同样准确的模型集合),并利用多重性度量(模糊性、差异性、Rashomon容量)和群体公平性指标评估其多样性。在表格数据、医学影像和自然语言处理任务上的实验证明,尽管替代模型与目标模型精度相似,但其在其他关键性能指标上仍存在显著差异。该结果质疑了高保真替代模型在实际部署中等同于原模型的假设。

原文摘要 · Abstract (English)

Model stealing attacks, where adversaries create high-fidelity surrogate models, are a significant threat to the intellectual property of machine learning services. Conventional wisdom suggests these surrogates could provide adversaries with economic leverage comparable to the original service providers. This paper challenges this assumption by evaluating model stealing attacks beyond mere fidelity to the target model. Because query-based extraction provides only partial supervision of the target's input-output behavior, the surrogate is not uniquely identified: many near-optimal surrogates can achieve comparable fidelity while differing in deployment-relevant properties. Instead of performing a classic learning-based model stealing attack, we compute the Rashomon Set (i.e., the set of almost-equally-accurate models) of surrogate models, and evaluate its diversity using multiplicity metrics (ambiguity, discrepancy, and Rashomon Capacity) and group fairness metrics. Across tabular, medical imaging, and NLP tasks, our experiments on real-world datasets reveal that despite exhibiting similar fidelity to the target model, surrogate models can display significant variances in other critical performance metrics. These findings cast doubt on the presumed equivalence between high-fidelity surrogates and the target model in practical deployment scenarios.

模型窃取模型多样性安全评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。