arXiv:2511.13892cs.AI2025-11ICML

研究智能交通中视觉语言模型的漏洞,提出新型攻击与防御方法。

Jailbreaking Large Vision Language Models in Intelligent Transportation Systems

  • 通过图像文字操控和多轮提示发动新攻击
  • 攻击使模型生成有害响应,毒性评分显著升高
  • 适合安全研究人员与智能交通系统开发者参考

大型视觉语言模型(LVLMs)在多模态推理方面表现出强大能力,广泛应用于视觉问答等实际场景。然而,这些模型在智能交通系统(ITS)中极易受到越狱攻击。本文系统分析了集成于ITS的LVLMs在精心设计的越狱攻击下的脆弱性。首先,依据OpenAI的禁止类别构建包含交通相关有害查询的数据集;其次,提出一种新型越狱攻击,利用图像文字操控与多轮提示触发模型漏洞;第三,设计多层次响应过滤防御机制以阻止不当输出。我们在最先进的开源与闭源LVLM上进行了广泛实验,采用GPT-4判断生成响应的毒性得分并辅以人工验证。对比现有方法,结果表明基于图像文字操控与多轮提示的越狱攻击在交通系统中存在严重安全风险。

原文摘要 · Abstract (English)

Large Vision Language Models (LVLMs) demonstrate strong capabilities in multimodal reasoning and many real-world applications, such as visual question answering. However, LVLMs are highly vulnerable to jailbreaking attacks. This paper systematically analyzes the vulnerabilities of LVLMs integrated in Intelligent Transportation Systems (ITS) under carefully crafted jailbreaking attacks. First, we carefully construct a dataset with harmful queries relevant to transportation, following OpenAI's prohibited categories to which the LVLMs should not respond. Second, we introduce a novel jailbreaking attack that exploits the vulnerabilities of LVLMs through image typography manipulation and multi-turn prompting. Third, we propose a multi-layered response filtering defense technique to prevent the model from generating inappropriate responses. We perform extensive experiments with the proposed attack and defense on the state-of-the-art LVLMs (both open-source and closed-source). To evaluate the attack method and defense technique, we use GPT-4's judgment to determine the toxicity score of the generated responses, as well as manual verification. Further, we compare our proposed jailbreaking method with existing jailbreaking techniques and highlight severe security risks involved with jailbreaking attacks with image typography manipulation and multi-turn prompting in the LVLMs integrated in ITS.

视觉语言模型越狱攻击智能交通安全防御

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。