arXiv:2604.01010cs.CVcs.MM2026-04

用文本增强提升视觉语言模型抗攻击能力,无需训练即可生效。

PDA: Text-Augmented Defense Framework for Robust Vision-Language Models against Adversarial Image Attacks

  • 测试时通过改写提示、拆分问题、聚合结果增强鲁棒性
  • 在多个任务上对抗攻击下保持高准确率,且不降低干净数据表现
  • 无需训练修改模型,适合部署于各类视觉语言系统

视觉语言模型(VLMs)易受对抗图像扰动影响。现有基于特定攻击样本的对抗训练方法计算成本高,且难以泛化到未见攻击类型。为此,我们提出无需训练的防御框架PDA(Paraphrase-Decomposition-Aggregation),利用文本增强提升VLM在多种对抗图像攻击下的鲁棒性。PDA在测试时执行提示改写、问题分解与一致性聚合,无需修改底层模型。为兼顾鲁棒性与效率,我们进一步设计不变量实现,显著降低推理开销,同时保留大部分鲁棒性收益。在多个VLM架构及视觉问答、分类、图文生成等基准上的实验表明,PDA在多种对抗扰动下均实现稳定鲁棒性提升,且保持良好的干净准确率,构建了一种通用、强效且实用的VLM推理阶段防御框架。

原文摘要 · Abstract (English)

Vision-language models (VLMs) are vulnerable to adversarial image perturbations. Existing works based on adversarial training against task-specific adversarial examples are computationally expensive and often fail to generalize to unseen attack types. To address these limitations, we introduce Paraphrase-Decomposition-Aggregation (PDA), a training-free defense framework that leverages text augmentation to enhance VLM robustness under diverse adversarial image attacks. PDA performs prompt paraphrasing, question decomposition, and consistency aggregation entirely at test time, thus requiring no modification on the underlying models. To balance robustness and efficiency, we instantiate PDA as invariants that reduce the inference cost while retaining most of its robustness gains. Experiments on multiple VLM architectures and benchmarks for visual question answering, classification, and captioning show that PDA achieves consistent robustness gains against various adversarial perturbations while maintaining competitive clean accuracy, establishing a generic, strong and practical defense framework for VLMs during inference.

视觉语言模型对抗防御文本增强推理防御

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。