arXiv:2509.16352cs.CRcs.AI2025-09被引 1

提出防御响应式机密信息推断攻击的新方法,保障模型共享中的数据安全。

Secure Confidential Business Information When Sharing Machine Learning Models

  • 设计可响应的攻击模拟器,更真实还原对手行为。
  • 在多种场景下显著提升防御效果,同时降低计算开销。
  • 适合关注模型共享隐私安全的工业界与研究者使用。

模型共享能为拥有成熟机器学习(ML)模型的企业带来显著商业价值,使其可将模型授权给缺乏资源自建模型的机构。然而,数据保密性仍是主要障碍,因机密属性推断(CPI)攻击可利用共享模型提取模型提供方训练数据的敏感信息。现有防御多假设攻击者为非适应性,忽略真实对手具备动态调整攻击策略的能力。为此,本文提出新防御方法,通过两项创新应对:一是设计新型响应式CPI攻击以模拟真实对手行为;二是构建攻防对抗迭代框架,持续优化目标模型与攻击模型,最终生成对响应式攻击鲁棒的可信模型。此外,引入一种近似策略,有效缓解防御方法的计算瓶颈,提升效率。在多种真实模型共享场景下的广泛实验表明,该方法在抵御CPI攻击、保持模型性能及降低计算开销方面均优于现有方法。

原文摘要 · Abstract (English)

Model-sharing offers significant business value by enabling firms with well-established Machine Learning (ML) models to monetize and share their models with others who lack the resources to develop ML models from scratch. However, concerns over data confidentiality remain a significant barrier to model-sharing adoption, as Confidential Property Inference (CPI) attacks can exploit shared ML models to uncover confidential properties of the model provider's private model training data. Existing defenses often assume that CPI attacks are non-adaptive to the specific ML model they are targeting. This assumption overlooks a key characteristic of real-world adversaries: their responsiveness, i.e., adversaries' ability to dynamically adjust their attack models based on the information of the target and its defenses. To overcome this limitation, we propose a novel defense method that explicitly accounts for the responsive nature of real-world adversaries via two methodological innovations: a novel Responsive CPI attack and an attack-defense arms race framework. The former emulates the responsive behaviors of adversaries in the real world, and the latter iteratively enhances both the target and attack models, ultimately producing a secure ML model that is robust against responsive CPI attacks. Furthermore, we propose and integrate a novel approximate strategy into our defense, which addresses a critical computational bottleneck of defense methods and improves defense efficiency. Through extensive empirical evaluations across various realistic model-sharing scenarios, we demonstrate that our method outperforms existing defenses by more effectively defending against CPI attacks, preserving ML model utility, and reducing computational overhead.

模型共享隐私保护对抗攻击

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。