arXiv:2510.22785cs.CV2025-10被引 2

提出自校准一致性机制,提升视觉语言模型抗攻击能力。

Self-Calibrated Consistency can Fight Back for Adversarial Robustness in Vision-Language Models

  • 利用伪标签和多视角预测增强跨模态对齐
  • 在22个基准上显著提升CLIP零样本鲁棒性
  • 无需微调,可直接用于其他视觉语言模型

预训练的视觉语言模型(如CLIP)在多个领域展现出强大的零样本能力,但对破坏图像-文本对齐的对抗扰动仍极为脆弱。现有防御方法通常依赖带标注数据的对抗微调,难以应用于零样本场景。本文识别出当前CLIP对抗攻击的两大弱点:缺乏语义引导与视点变化敏感性,统称为语义和视角脆弱性。为此,提出自校准一致性(SCC)测试时防御策略,包含两个互补模块:语义一致性利用对抗反攻预热生成的软伪标签与多视角预测,正则化跨模态对齐并分离目标嵌入与混淆负例;空间一致性通过增强视图对齐受扰视觉预测,稳定对抗扰动下的推理。两者构成即插即用的推理策略。在22个不同攻击设置的基准上进行大量实验表明,SCC能持续提升CLIP的零样本鲁棒性,同时保持准确率,并可无缝集成到其他VLM中进一步提升性能。这些发现揭示了从CLIP构建对抗鲁棒范式的重要潜力,对BioMedCLIP等更广泛的视觉语言领域具有启示意义。

原文摘要 · Abstract (English)

Pre-trained vision-language models (VLMs) such as CLIP have demonstrated strong zero-shot capabilities across diverse domains, yet remain highly vulnerable to adversarial perturbations that disrupt image-text alignment and compromise reliability. Existing defenses typically rely on adversarial fine-tuning with labeled data, limiting their applicability in zero-shot settings. In this work, we identify two key weaknesses of current CLIP adversarial attacks -- lack of semantic guidance and vulnerability to view variations -- collectively termed semantic and viewpoint fragility. To address these challenges, we propose Self-Calibrated Consistency (SCC), an effective test-time defense. SCC consists of two complementary modules: Semantic consistency, which leverages soft pseudo-labels from counterattack warm-up and multi-view predictions to regularize cross-modal alignment and separate the target embedding from confusable negatives; and Spatial consistency, aligning perturbed visual predictions via augmented views to stabilize inference under adversarial perturbations. Together, these modules form a plug-and-play inference strategy. Extensive experiments on 22 benchmarks under diverse attack settings show that SCC consistently improves the zero-shot robustness of CLIP while maintaining accuracy, and can be seamlessly integrated with other VLMs for further gains. These findings highlight the great potential of establishing an adversarially robust paradigm from CLIP, with implications extending to broader vision-language domains such as BioMedCLIP.

视觉语言模型对抗鲁棒性零样本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。