提出统一框架,分析模型窃取攻击与防御的隐私权衡。
Model Privacy: A Unified Framework for Understanding Model Stealing Attacks and Defenses
- 构建模型隐私框架,明确攻击与防御的数学定义。
- 揭示模型效用与隐私间的根本权衡关系。
- 适合关注AI安全与模型保护的研究者阅读。
机器学习(ML)在多个领域广泛应用,其安全性日益受到关注。其中,模型窃取攻击成为关键威胁:攻击者通过有限的查询-响应交互(如云服务或芯片级AI接口)尝试恢复训练好的模型。现有攻防策略常缺乏理论基础和标准化评估标准。为此,本文提出名为「模型隐私」的统一框架,为全面分析模型窃取攻击与防御提供理论基础。我们建立了严格的威胁模型与目标形式化,提出量化攻击与防御策略优劣的方法,并分析了模型效用与隐私之间的基本权衡。该理论揭示了扰动结构对防御有效性的重要性。通过多种学习场景验证了从防御者视角应用模型隐私的可行性,大量实验支持了理论洞察及所提防御机制的有效性。
原文摘要 · Abstract (English)
The use of machine learning (ML) has become increasingly prevalent in various domains, highlighting the importance of understanding and ensuring its safety. One pressing concern is the vulnerability of ML applications to model stealing attacks. These attacks involve adversaries attempting to recover a learned model through limited query-response interactions, such as those found in cloud-based services or on-chip artificial intelligence interfaces. While existing literature proposes various attack and defense strategies, these often lack a theoretical foundation and standardized evaluation criteria. In response, this work presents a framework called ``Model Privacy'', providing a foundation for comprehensively analyzing model stealing attacks and defenses. We establish a rigorous formulation for the threat model and objectives, propose methods to quantify the goodness of attack and defense strategies, and analyze the fundamental tradeoffs between utility and privacy in ML models. Our developed theory offers valuable insights into enhancing the security of ML models, especially highlighting the importance of the attack-specific structure of perturbations for effective defenses. We demonstrate the application of model privacy from the defender's perspective through various learning scenarios. Extensive experiments corroborate the insights and the effectiveness of defense mechanisms developed under the proposed framework.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。