用良性模块拼出恶意输出,揭示视觉语言模型新安全漏洞
Models as Lego Builders: Assembling Malice from Benign Blocks via Semantic Blueprints
- 将有害指令拆解为无害槽位,以结构化图像提示隐蔽注入
- 单次查询即可攻破多模型,成功率超90%(在多个基准上)
- 适合研究模型安全或对抗攻击的学者参考
尽管大型视觉语言模型(LVLMs)发展迅速,但视觉模态的引入带来了新的安全漏洞,攻击者可利用这些漏洞诱导模型产生偏见或恶意输出。本文揭示了一种未被充分关注的漏洞:语义槽填充。即使槽位类型被刻意设计为看似无害,LVLM仍会用不安全内容填充缺失值。基于此发现,我们提出 StructAttack——一种在黑盒设置下的简单而有效的单次查询越狱框架。该方法将有害查询分解为中心主题和一组看似无害的槽位类型,将其嵌入带有微小随机扰动的结构化视觉提示(如思维导图、表格或旭日图)中,并搭配补全引导指令。LVLM会自动重组隐藏语义,生成不安全输出,且不触发安全机制。虽然每个槽位单独看都是无害的(局部无害性),但 StructAttack 利用了模型推理能力将其组合成连贯的恶意语义。在多个模型和基准上的大量实验表明,该方法具有显著有效性。
原文摘要 · Abstract (English)
Despite the rapid progress of Large Vision-Language Models (LVLMs), the integration of visual modalities introduces new safety vulnerabilities that adversaries can exploit to elicit biased or malicious outputs. In this paper, we demonstrate an underexplored vulnerability via semantic slot filling, where LVLMs complete missing slot values with unsafe content even when the slot types are deliberately crafted to appear benign. Building on this finding, we propose StructAttack, a simple yet effective single-query jailbreak framework under black-box settings. StructAttack decomposes a harmful query into a central topic and a set of benign-looking slot types, then embeds them as structured visual prompts (e.g., mind maps, tables, or sunburst diagrams) with small random perturbations. Paired with a completion-guided instruction, LVLMs automatically recompose the concealed semantics and generate unsafe outputs without triggering safety mechanisms. Although each slot appears benign in isolation (local benignness), StructAttack exploits LVLMs' reasoning to assemble these slots into coherent harmful semantics. Extensive experiments on multiple models and benchmarks show the efficacy of our proposed StructAttack.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。