arXiv:2411.09540cs.CVcs.AI2024-11中稿 · IEEE/IFIP DSN 2025被引 1

通过视觉提示检测黑盒模型中的隐藏后门

Prompting the Unseen: Detecting Hidden Backdoors in Black-Box Models

  • 利用视觉提示映射源域与目标域类别子空间
  • 发现干净数据与中毒数据间存在类别子空间不一致
  • 无需白盒信息,适合检测可疑模型的后门

视觉提示(VP)是一种将已训练好的冻结模型适配到新领域任务的新技术。本文研究了该技术在黑盒模型级后门检测中的应用价值。视觉提示通过映射源域与目标域之间的类别子空间实现跨域迁移。我们发现,在干净数据与中毒数据之间存在一类称为‘类别子空间不一致’的偏差。基于此,提出 extsc{BProm} 方法,用于检测可疑模型中是否存在后门。该方法利用当模型存在后门时,提示后分类准确率显著下降的特性进行判断。大量实验验证了 extsc{BProm} 的有效性。

原文摘要 · Abstract (English)

Visual prompting (VP) is a new technique that adapts well-trained frozen models for source domain tasks to target domain tasks. This study examines VP's benefits for black-box model-level backdoor detection. The visual prompt in VP maps class subspaces between source and target domains. We identify a misalignment, termed class subspace inconsistency, between clean and poisoned datasets. Based on this, we introduce \textsc{BProm}, a black-box model-level detection method to identify backdoors in suspicious models, if any. \textsc{BProm} leverages the low classification accuracy of prompted models when backdoors are present. Extensive experiments confirm \textsc{BProm}'s effectiveness.

后门检测视觉提示黑盒攻击

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。