arXiv:2502.05325cs.LGcs.CR2025-02NeurIPS被引 2

分析树模型提取攻击的查询效率,揭示解释技术带来的安全风险。

From Counterfactuals to Trees: Competitive Analysis of Model Extraction Attacks

  • 用竞争分析框架评估模型提取攻击的查询开销。
  • 提出新算法可完美重建树模型,且任意时刻表现优异。
  • 适合关注MLaaS安全与模型保护的研究者阅读。

机器学习即服务(MLaaS)的兴起加剧了模型可解释性与安全性之间的权衡。例如,反事实解释等可解释性技术会无意中增加模型提取攻击的风险,导致专有模型被未经授权复制。本文从竞争分析视角,首次对模型提取攻击进行形式化分析,建立了评估其效率的基础框架。聚焦基于加法决策树的模型(如决策树、梯度提升、随机森林),提出新的重建算法,在保证理论上完美保真度的同时,展现出强大的任意时刻性能。该框架为树模型的查询复杂度提供了理论边界,揭示了其部署中的安全漏洞。

原文摘要 · Abstract (English)

The advent of Machine Learning as a Service (MLaaS) has heightened the trade-off between model explainability and security. In particular, explainability techniques, such as counterfactual explanations, inadvertently increase the risk of model extraction attacks, enabling unauthorized replication of proprietary models. In this paper, we formalize and characterize the risks and inherent complexity of model reconstruction, focusing on the "oracle'' queries required for faithfully inferring the underlying prediction function. We present the first formal analysis of model extraction attacks through the lens of competitive analysis, establishing a foundational framework to evaluate their efficiency. Focusing on models based on additive decision trees (e.g., decision trees, gradient boosting, and random forests), we introduce novel reconstruction algorithms that achieve provably perfect fidelity while demonstrating strong anytime performance. Our framework provides theoretical bounds on the query complexity for extracting tree-based model, offering new insights into the security vulnerabilities of their deployment.

模型安全提取攻击决策树竞争分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。