arXiv:2603.29410cs.CVcs.AI2026-03中稿 · CVPR被引 1

用软对齐提升视觉语言模型对抗鲁棒性,不破坏原有语义结构。

AGFT: Alignment-Guided Fine-Tuning for Zero-Shot Adversarial Robustness of Vision-Language Models

  • 基于原模型概率预测进行文本引导的对抗训练,实现软对齐。
  • 在多个零样本基准上显著提升对抗鲁棒性,零样本性能下降<2%。
  • 适合追求鲁棒性与泛化能力平衡的研究者与应用开发者。

预训练视觉语言模型具备强大的零样本泛化能力,但对对抗扰动仍敏感。现有基于分类标签的对抗微调方法常破坏预训练阶段的跨模态对齐,削弱视觉-文本对应关系,导致零样本性能下降。本文提出一种对齐引导的微调框架(AGFT),在增强零样本对抗鲁棒性的同时保持跨模态语义结构。不同于依赖硬标签的方法,AGFT利用原始模型的概率输出进行文本引导的对抗训练,通过软对齐分布将对抗视觉特征与文本嵌入对齐,提升零样本对抗鲁棒性。为缓解微调引入的结构偏差,我们设计分布一致性校准机制,将鲁棒模型输出调整为与温度缩放后的预训练模型预测一致。大量实验表明,AGFT在多个零样本基准上优于现有方法,显著提升零样本对抗鲁棒性。

原文摘要 · Abstract (English)

Pre-trained vision-language models (VLMs) exhibit strong zero-shot generalization but remain vulnerable to adversarial perturbations. Existing classification-guided adversarial fine-tuning methods often disrupt pre-trained cross-modal alignment, weakening visual-textual correspondence and degrading zero-shot performance. In this paper, we propose an Alignment-Guided Fine-Tuning (AGFT) framework that enhances zero-shot adversarial robustness while preserving the cross-modal semantic structure. Unlike label-based methods that rely on hard labels and fail to maintain the relative relationships between image and text, AGFT leverages the probabilistic predictions of the original model for text-guided adversarial training, which aligns adversarial visual features with textual embeddings via soft alignment distributions, improving zero-shot adversarial robustness. To address structural discrepancies introduced by fine-tuning, we introduce a distribution consistency calibration mechanism that adjusts the robust model output to match a temperature-scaled version of the pre-trained model predictions. Extensive experiments across multiple zero-shot benchmarks demonstrate that AGFT outperforms state-of-the-art methods while significantly improving zero-shot adversarial robustness.

视觉语言模型对抗鲁棒性微调软对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。