提出BESA攻击,可有效突破防御性扰动对编码器窃取的限制。
BESA: Boosting Encoder Stealing Attack with Perturbation Recovery
- 通过检测与恢复机制,从扰动特征中还原干净特征。
- 在多种防御下,使窃取模型准确率提升最高达24.63%。
- 适合研究模型安全与对抗攻击的学者参考。
为提升在基于扰动防御下的编码器窃取攻击性能,本文提出一种名为BESA的增强型编码器窃取攻击方法。该方法通过两个模块实现:扰动检测与扰动恢复,可与现有攻击方式结合使用。扰动检测模块利用目标编码器输出的特征向量推断服务方所采用的防御机制;一旦识别出防御策略,扰动恢复模块则借助设计良好的生成模型,从被扰动的特征中重建出原始特征。在多个数据集上的大量实验表明,当面对最先进的防御策略或多种防御组合时,BESA可使现有编码器窃取攻击的代理模型准确率最高提升24.63%。
原文摘要 · Abstract (English)
To boost the encoder stealing attack under the perturbation-based defense that hinders the attack performance, we propose a boosting encoder stealing attack with perturbation recovery named BESA. It aims to overcome perturbation-based defenses. The core of BESA consists of two modules: perturbation detection and perturbation recovery, which can be combined with canonical encoder stealing attacks. The perturbation detection module utilizes the feature vectors obtained from the target encoder to infer the defense mechanism employed by the service provider. Once the defense mechanism is detected, the perturbation recovery module leverages the well-designed generative model to restore a clean feature vector from the perturbed one. Through extensive evaluations based on various datasets, we demonstrate that BESA significantly enhances the surrogate encoder accuracy of existing encoder stealing attacks by up to 24.63\% when facing state-of-the-art defenses and combinations of multiple defenses.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。