arXiv:2507.20704cs.CLcs.AI2025-07被引 1

将纯文本数据转为图文数据,评估视觉语言模型的对抗攻击弱点。

Text2VLM: Adapting Text-Only Datasets to Evaluate Alignment Training in Visual Language Models

  • 构建多阶段流水线,把文本数据转为带图像的对抗性提示。
  • 发现开源模型在引入图像后更易受提示注入攻击,性能远低于闭源模型。
  • 适合关注多模态安全、模型鲁棒性评估的研究者使用。

视觉语言模型(VLMs)在人工智能系统中日益普及,其对齐能力至关重要,尤其在处理图文混合内容时。现有评估数据集主要依赖纯文本提示,导致视觉层面的脆弱性未被充分检验。为此,我们提出Text2VLM——一种新型多阶段流程,可将纯文本数据集转换为多模态格式,专门用于评估VLM对字体型提示注入攻击的韧性。该流程识别原文中的有害内容,并将其转化为图像形式,生成多模态输入。对开源VLM的评估显示,引入视觉输入后模型对提示注入攻击的敏感性显著上升,暴露出当前模型对齐机制的关键缺陷,且与闭源前沿模型存在明显性能差距。通过人工评估验证,提取的关键概念、文本摘要及输出分类均与人类预期一致。Text2VLM提供了一种可扩展的安全评估工具,助力构建更鲁棒的VLM安全机制,推动其在多样化真实场景中的安全部署。

原文摘要 · Abstract (English)

The increasing integration of Visual Language Models (VLMs) into AI systems necessitates robust model alignment, especially when handling multimodal content that combines text and images. Existing evaluation datasets heavily lean towards text-only prompts, leaving visual vulnerabilities under evaluated. To address this gap, we propose \textbf{Text2VLM}, a novel multi-stage pipeline that adapts text-only datasets into multimodal formats, specifically designed to evaluate the resilience of VLMs against typographic prompt injection attacks. The Text2VLM pipeline identifies harmful content in the original text and converts it into a typographic image, creating a multimodal prompt for VLMs. Also, our evaluation of open-source VLMs highlights their increased susceptibility to prompt injection when visual inputs are introduced, revealing critical weaknesses in the current models' alignment. This is in addition to a significant performance gap compared to closed-source frontier models. We validate Text2VLM through human evaluations, ensuring the alignment of extracted salient concepts; text summarization and output classification align with human expectations. Text2VLM provides a scalable tool for comprehensive safety assessment, contributing to the development of more robust safety mechanisms for VLMs. By enhancing the evaluation of multimodal vulnerabilities, Text2VLM plays a role in advancing the safe deployment of VLMs in diverse, real-world applications.

多模态安全提示注入模型评估VLM

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。