arXiv:2412.08755cs.CVcs.AI2024-12被引 1

用提示调优检测未见过的后门图像,提升视觉语言模型安全性

Proactive Adversarial Defense: Harnessing Prompt Tuning in Vision-Language Models to Detect Unseen Backdoored Images

  • 通过可学习文本提示区分干净图像与带隐藏后门的图像
  • 在两个数据集上平均检测准确率达86%,有效识别未知后门样本
  • 适合关注模型安全、对抗攻击防御的研究者和工程师

后门攻击通过在输入中嵌入隐藏触发器,使模型将特定样本误分类为目标标签,构成严重威胁。尽管已有大量研究针对图像识别模型的权重微调来缓解此类攻击,但对直接检测后门样本的关注仍不足。由于训练数据量庞大,人工检查触发器不现实,且现有防御机制难以完全消除其影响。为此,本文提出一种新方法,在训练与推理阶段检测未见过的后门图像。利用视觉语言模型(VLM)中提示调优的成功经验,训练可学习的文本提示以区分干净图像与含隐藏后门触发器的图像。实验表明,该方法在两个知名数据集上平均检测准确率达到86%,显著优于现有技术,为后门防御树立了新标准。

原文摘要 · Abstract (English)

Backdoor attacks pose a critical threat by embedding hidden triggers into inputs, causing models to misclassify them into target labels. While extensive research has focused on mitigating these attacks in object recognition models through weight fine-tuning, much less attention has been given to detecting backdoored samples directly. Given the vast datasets used in training, manual inspection for backdoor triggers is impractical, and even state-of-the-art defense mechanisms fail to fully neutralize their impact. To address this gap, we introduce a groundbreaking method to detect unseen backdoored images during both training and inference. Leveraging the transformative success of prompt tuning in Vision Language Models (VLMs), our approach trains learnable text prompts to differentiate clean images from those with hidden backdoor triggers. Experiments demonstrate the exceptional efficacy of this method, achieving an impressive average accuracy of 86% across two renowned datasets for detecting unseen backdoor triggers, establishing a new standard in backdoor defense.

后门检测视觉语言模型提示调优模型安全

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。