通过视觉提示检测黑盒模型中的隐藏后门
Prompting the Unseen: Detecting Hidden Backdoors in Black-Box Models
- 利用视觉提示映射源域与目标域类别子空间
- 发现干净数据与中毒数据间存在类别子空间不一致
- 无需白盒信息,适合检测可疑模型的后门
视觉提示(VP)是一种将已训练好的冻结模型适配到新领域任务的新技术。本文研究了该技术在黑盒模型级后门检测中的应用价值。视觉提示通过映射源域与目标域之间的类别子空间实现跨域迁移。我们发现,在干净数据与中毒数据之间存在一类称为‘类别子空间不一致’的偏差。基于此,提出 extsc{BProm} 方法,用于检测可疑模型中是否存在后门。该方法利用当模型存在后门时,提示后分类准确率显著下降的特性进行判断。大量实验验证了 extsc{BProm} 的有效性。
原文摘要 · Abstract (English)
Visual prompting (VP) is a new technique that adapts well-trained frozen models for source domain tasks to target domain tasks. This study examines VP's benefits for black-box model-level backdoor detection. The visual prompt in VP maps class subspaces between source and target domains. We identify a misalignment, termed class subspace inconsistency, between clean and poisoned datasets. Based on this, we introduce \textsc{BProm}, a black-box model-level detection method to identify backdoors in suspicious models, if any. \textsc{BProm} leverages the low classification accuracy of prompted models when backdoors are present. Extensive experiments confirm \textsc{BProm}'s effectiveness.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。