用视觉和文本碎片隐蔽攻击大模型,绕过安全检测
Cross-Modal Obfuscation for Jailbreak Attacks on Large Vision-Language Models
- 将恶意指令拆成看似无害的图文片段,利用跨模态推理暗中重组
- 只需少量查询即可成功攻击,比现有方法更隐蔽高效
- 适用于研究模型安全漏洞或防御机制的开发者
大型视觉语言模型在多模态任务中表现优异,但仍易受越狱攻击,可绕过内置安全机制生成受限内容。现有黑盒越狱方法主要依赖对抗性文本提示或图像扰动,但易被内容过滤系统识别,且查询与计算效率低。本文提出跨模态对抗性混淆框架CAMO,将恶意提示分解为语义上无害的视觉与文本片段,借助LVLM的跨模态推理能力,通过多步推理隐蔽重构有害指令,规避传统检测机制。该方法支持可调推理复杂度,显著减少查询次数,兼具隐蔽性与高效性。在主流LVLM上的全面评估验证了CAMO的有效性,表现出强鲁棒性与跨模型迁移能力。结果揭示当前内置安全机制存在重大漏洞,亟需面向对齐的先进安全防护方案。
原文摘要 · Abstract (English)
Large Vision-Language Models (LVLMs) demonstrate exceptional performance across multimodal tasks, yet remain vulnerable to jailbreak attacks that bypass built-in safety mechanisms to elicit restricted content generation. Existing black-box jailbreak methods primarily rely on adversarial textual prompts or image perturbations, yet these approaches are highly detectable by standard content filtering systems and exhibit low query and computational efficiency. In this work, we present Cross-modal Adversarial Multimodal Obfuscation (CAMO), a novel black-box jailbreak attack framework that decomposes malicious prompts into semantically benign visual and textual fragments. By leveraging LVLMs' cross-modal reasoning abilities, CAMO covertly reconstructs harmful instructions through multi-step reasoning, evading conventional detection mechanisms. Our approach supports adjustable reasoning complexity and requires significantly fewer queries than prior attacks, enabling both stealth and efficiency. Comprehensive evaluations conducted on leading LVLMs validate CAMO's effectiveness, showcasing robust performance and strong cross-model transferability. These results underscore significant vulnerabilities in current built-in safety mechanisms, emphasizing an urgent need for advanced, alignment-aware security and safety solutions in vision-language systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。