arXiv:2508.05237cs.CVcs.AI2025-08

提出零样本对抗鲁棒性防御新框架,平衡模型泛化与抗攻击能力。

Navigating the Trade-off: A Synthesis of Defensive Strategies for Zero-Shot Adversarial Robustness in Vision-Language Models

  • 对比参数调整与无训练防御策略,分析八篇关键论文方法
  • 发现从对齐保持到嵌入空间重构的防御演进路径
  • 适合关注VLM安全性的研究者与工程落地人员

本报告综合分析了八篇关于视觉语言模型(如CLIP)零样本对抗鲁棒性的前沿论文。该领域核心挑战在于提升对抗鲁棒性与维持零样本泛化能力之间的固有权衡。我们梳理了两类主要防御范式:通过修改模型参数的对抗微调(AFT),以及保持参数不变的无训练/测试时防御。防御策略从对齐保持方法(TeCoA)发展到嵌入空间重构(LAAT、TIMA),再演变为输入启发式(AOM、TTC)与潜在空间净化(CLIPure)。最后,我们指出了混合防御策略与对抗预训练等关键挑战与未来方向。

原文摘要 · Abstract (English)

This report synthesizes eight seminal papers on the zero-shot adversarial robustness of vision-language models (VLMs) like CLIP. A central challenge in this domain is the inherent trade-off between enhancing adversarial robustness and preserving the model's zero-shot generalization capabilities. We analyze two primary defense paradigms: Adversarial Fine-Tuning (AFT), which modifies model parameters, and Training-Free/Test-Time Defenses, which preserve them. We trace the evolution from alignment-preserving methods (TeCoA) to embedding space re-engineering (LAAT, TIMA), and from input heuristics (AOM, TTC) to latent-space purification (CLIPure). Finally, we identify key challenges and future directions including hybrid defense strategies and adversarial pre-training.

视觉语言模型对抗鲁棒性零样本防御策略

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。