通过协同演化搜索,提升视觉语言模型跨任务攻击效果
Two Sides of the Same Coin: Co-Evolving Search for Cross-Task Attacks on Vision-Language Models

- 联合演化文本语义与图像区域扰动,扩大搜索空间
- 在多个任务上实现一致的误分类,攻击成功率显著提升
- 适合研究模型鲁棒性或对抗攻击的从业者参考
视觉语言模型(VLMs)虽具备多模态任务强泛化能力,但仍易受对抗扰动影响。现有攻击方法多采用单一轨迹梯度优化或特定任务目标,限制了搜索空间探索与跨任务迁移性。本文提出一种基于进化计算的跨模态攻击框架,可自适应地联合搜索文本与视觉空间。文本侧通过演化生成围绕源类别表征的硬负语义嵌入,提供多样化的跨模态排斥;图像侧维护物体区域扰动种群,结合动量梯度更新与进化选择、变异、交叉操作,更可靠地探索多种可行路径。联合优化语义负向引导与局部扰动,生成的对抗样本能持续将源物体语义导向目标类别,覆盖图像描述、目标检测、区域分类与定位等任务。理论分析表明,协同进化搜索保持扰动可行性,防止最优适应度退化,并提高到达高置信度对抗区域的概率。在Florence-2、OFA与UnifiedIO-2上的实验验证了该框架在多项任务中的强大攻击性能。消融实验证实文本侧语义演化与图像侧扰动演化的互补有效性,以及框架的高效性与跨任务迁移能力。
原文摘要 · Abstract (English)
Vision-language models (VLMs) exhibit strong generalization across multimodal tasks but remain vulnerable to adversarial perturbations. Existing attacks typically follow single-trajectory gradient optimization or task-specific objectives, limiting search-space exploration and cross-task transferability. We propose an evolutionary-computation-guided cross-modal attack framework for unified VLMs. The framework adaptively searches both textual and visual spaces. On the textual side, it evolves hard negative semantic embeddings around the source-category representation to provide diverse cross-modal repulsion. On the visual side, it maintains a population of object-region perturbations and combines momentum-based gradient updates with evolutionary selection, mutation, and crossover to more reliably explore multiple feasible trajectories. Jointly optimizing semantic negative guidance and localized perturbations generates adversarial examples that consistently shift source-object semantics toward target categories across vision-language tasks. Theoretical analyses show that the co-evolutionary search preserves perturbation feasibility, prevents degradation of the best observed fitness, and increases the probability of reaching high-margin adversarial regions compared with single-trajectory optimization. Experiments on Florence-2, OFA, and UnifiedIO-2 demonstrate strong overall attack performance across image captioning, object detection, region categorization, and object localization. Ablation studies further verify the complementary effectiveness of text-side semantic evolution and image-side perturbation evolution, as well as the framework's efficiency and cross-task transferability.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。