arXiv:2411.15244cs.CVcs.AI2024-11被引 7

用双模态知识蒸馏提升视觉语言模型的抗攻击能力

Adversarial Prompt Distillation for Vision-Language Models

  • 同时优化视觉与文本提示,融合双模态知识蒸馏
  • 在多个数据集上同时提升鲁棒性与正常准确率
  • 可用普通预训练模型作教师,适合安全敏感场景

大型预训练视觉语言模型(如CLIP)易受对抗攻击,影响其在自动驾驶、医疗诊断等关键场景的应用。现有对抗提示调优(APT)方法多为单模态,仅针对视觉或文本模态设计提示,限制了鲁棒性与干净准确率。本文提出对抗提示蒸馏(APD),一种双模态知识蒸馏框架,将APT与多模态知识迁移结合,同时优化视觉与文本提示,并从干净的预训练教师模型(CLIP)中蒸馏知识。在多个基准数据集上的实验表明,APD在对抗鲁棒性和干净准确率方面均优于当前最先进APT方法。结果还验证了使用非鲁棒教师模型可提升微调后模型的泛化与鲁棒性。

原文摘要 · Abstract (English)

Large pre-trained Vision-Language Models (VLMs) such as Contrastive Language-Image Pre-training (CLIP) have been shown to be susceptible to adversarial attacks, raising concerns about their deployment in safety-critical applications like autonomous driving and medical diagnosis. One promising approach for robustifying pre-trained VLMs is Adversarial Prompt Tuning (APT), which applies adversarial training during the process of prompt tuning. However, existing APT methods are mostly single-modal methods that design prompt(s) for only the visual or textual modality, limiting their effectiveness in either robustness or clean accuracy. In this work, we propose Adversarial Prompt Distillation (APD), a bimodal knowledge distillation framework that enhances APT by integrating it with multi-modal knowledge transfer. APD optimizes prompts for both visual and textual modalities while distilling knowledge from a clean pre-trained teacher CLIP model. Extensive experiments on multiple benchmark datasets demonstrate the superiority of our APD method over the current state-of-the-art APT methods in terms of both adversarial robustness and clean accuracy. The effectiveness of APD also validates the possibility of using a non-robust teacher to improve the generalization and robustness of fine-tuned VLMs.

视觉语言模型对抗攻击知识蒸馏提示调优

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。