用生成式AI无数据攻击深度学习模型,效果接近有数据的白盒攻击。
Data-Free Model-Related Attacks: Unleashing the Potential of Generative AI
- 利用生成式AI在无训练数据情况下发起黑盒攻击。
- 对图像和文本模型实现提取、成员推断、反演等攻击,性能媲美白盒方法。
- 揭示生成式AI滥用风险,提醒研究者警惕模型安全漏洞。
生成式AI技术已深度融入日常生活,显著提升生产力。然而,其能力也可能被恶意利用。现有研究多关注生成式AI在网络安全攻击中的应用,较少探讨其对深度学习模型的威胁。本文首次提出利用生成式AI实施模型相关攻击,包括模型提取、成员推断和模型反演。研究表明,攻击者可在无目标模型训练数据与参数的条件下,以黑盒方式对图像与文本模型发起多种攻击,其效果可与拥有完整白盒信息的基准方法相媲美。本研究为社区敲响警钟,提示生成式AI驱动的模型攻击存在重大潜在风险。
原文摘要 · Abstract (English)
Generative AI technology has become increasingly integrated into our daily lives, offering powerful capabilities to enhance productivity. However, these same capabilities can be exploited by adversaries for malicious purposes. While existing research on adversarial applications of generative AI predominantly focuses on cyberattacks, less attention has been given to attacks targeting deep learning models. In this paper, we introduce the use of generative AI for facilitating model-related attacks, including model extraction, membership inference, and model inversion. Our study reveals that adversaries can launch a variety of model-related attacks against both image and text models in a data-free and black-box manner, achieving comparable performance to baseline methods that have access to the target models' training data and parameters in a white-box manner. This research serves as an important early warning to the community about the potential risks associated with generative AI-powered attacks on deep learning models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。