arXiv:2506.05867cs.CRcs.LG2025-06ICML

无需提示词设计即可盗取黑盒模型,让攻击更隐蔽高效。

Stealix: Model Stealing via Prompt Evolution

  • 用遗传算法自动优化提示词,无需人工设计
  • 在相同查询次数下,生成图像精度和多样性显著优于现有方法
  • 适合无专业背景的攻击者,揭示预训练模型的新风险

模型窃取对机器学习构成重大安全威胁,使攻击者可在不访问训练数据的情况下复制黑盒模型,损害知识产权并暴露敏感信息。现有利用预训练扩散模型进行数据合成的方法虽提升效率与性能,但严重依赖手工设计提示词,限制了自动化与可扩展性,尤其对缺乏专业知识的攻击者不利。为评估开源预训练模型带来的风险,我们提出更贴近现实的威胁模型,无需提示词设计能力或类别名称知识。在此背景下,我们提出Stealix,首个无需预设提示词的模型窃取方法。Stealix利用两个开源预训练模型推断目标模型的数据分布,并通过遗传算法迭代优化提示词,逐步提升合成图像的精确度与多样性。实验表明,即使在仅使用通用提示词的情况下,Stealix也显著超越其他需类别名或细粒度提示词的方法,且在相同查询预算下表现更优。结果表明该方法具有更强可扩展性,提示预训练生成模型在模型窃取中的风险可能被低估。

原文摘要 · Abstract (English)

Model stealing poses a significant security risk in machine learning by enabling attackers to replicate a black-box model without access to its training data, thus jeopardizing intellectual property and exposing sensitive information. Recent methods that use pre-trained diffusion models for data synthesis improve efficiency and performance but rely heavily on manually crafted prompts, limiting automation and scalability, especially for attackers with little expertise. To assess the risks posed by open-source pre-trained models, we propose a more realistic threat model that eliminates the need for prompt design skills or knowledge of class names. In this context, we introduce Stealix, the first approach to perform model stealing without predefined prompts. Stealix uses two open-source pre-trained models to infer the victim model's data distribution, and iteratively refines prompts through a genetic algorithm, progressively improving the precision and diversity of synthetic images. Our experimental results demonstrate that Stealix significantly outperforms other methods, even those with access to class names or fine-grained prompts, while operating under the same query budget. These findings highlight the scalability of our approach and suggest that the risks posed by pre-trained generative models in model stealing may be greater than previously recognized.

模型窃取扩散模型提示优化安全风险

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。