arXiv:2508.01741cs.CV2025-08被引 4

提出模拟集成攻击,让越狱攻击在微调后的视觉语言模型间高效迁移。

Simulated Ensemble Attack: Transferring Jailbreaks Across Fine-tuned Vision-Language Models

  • 通过模拟参数变化路径和文本引导,增强越狱攻击的迁移能力。
  • 在多个微调版Qwen2-VL上实现高成功率和高毒性,传统方法几乎无效。
  • 适用于安全增强模型,且跨不同基座模型有效,适合安全评估研究者。

开源视觉语言模型(VLM)的广泛微调带来了重要安全风险:基座模型中的越狱漏洞可能延续至下游版本,导致攻击可在微调系统间传递。为探究此风险,我们提出模拟集成攻击(SEA),一种灰盒越狱框架,假设可访问基座VLM但无目标微调模型信息。SEA通过细调轨迹模拟(FTS)建模视觉编码器中参数的有限变化,并利用目标提示引导(TPG)通过辅助文本指导稳定对抗优化。在Qwen2-VL系列上的实验表明,SEA在多种微调变体(包括安全增强模型)中均保持高转移成功率与毒性,而基于PGD的图像越狱方法几乎无法迁移。进一步分析显示,微调主要引起基座模型附近局部参数偏移,解释了在模拟邻域优化的攻击为何能有效转移。此外,SEA在不同基座版本(如Qwen2.5/3-VL)间也表现出泛化能力,表明其有效性源于微调带来的共性行为,而非特定架构或初始化因素。

原文摘要 · Abstract (English)

The widespread practice of fine-tuning open-source Vision-Language Models (VLMs) raises a critical security concern: jailbreak vulnerabilities in base models may persist in downstream variants, enabling transferable attacks across fine-tuned systems. To investigate this risk, we propose the Simulated Ensemble Attack (SEA), a grey-box jailbreak framework that assumes full access to the base VLM but no knowledge of the fine-tuned target. SEA enhances transferability via Fine-tuning Trajectory Simulation (FTS), which models bounded parameter variations in the vision encoder, and Targeted Prompt Guidance (TPG), which stabilizes adversarial optimization through auxiliary textual guidance. Experiments on the Qwen2-VL family demonstrate that SEA achieves consistently high transfer success and toxicity rates across diverse fine-tuned variants, including safety-enhanced models, while standard PGD-based image jailbreaks exhibit negligible transferability. Further analysis reveals that fine-tuning primarily induces localized parameter shifts around the base model, explaining why attacks optimized over a simulated neighborhood transfer effectively. We also show that SEA generalizes across different base generations (e.g., Qwen2.5/3-VL), indicating that its effectiveness arises from shared fine-tuning-induced behaviors rather than architecture- or initialization-specific factors.

越狱攻击视觉语言模型安全评估迁移攻击

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。