构建多模态安全测试集,发现视觉语言模型存在隐蔽安全隐患
MSTS: A Multimodal Safety Test Suite for Vision-Language Models
- 设计400个图文组合测试用例,覆盖40类细粒度危险场景
- 多个开源模型在多模态下暴露明显安全缺陷,部分仅因理解失败而看似安全
- 支持十种语言,揭示非英文提示可提升模型违规率
视觉语言模型(VLMs)在聊天助手等消费级AI应用中日益普及。然而缺乏有效防护时,可能生成自残指导或鼓励吸毒等有害行为。尽管风险明显,现有研究极少评估多模态输入带来的新型安全威胁。为此,本文提出MSTS——一个面向VLM的多模态安全测试套件,包含400个测试提示,涵盖40种细粒度危害类别。每个提示由文本与图像共同构成,单独无法揭示其真实危险性。通过MSTS,我们发现多个开源VLM存在显著安全问题;同时观察到部分模型仅因未能理解简单提示而意外表现安全。我们将MSTS翻译为10种语言,结果表明非英文提示可提高模型违规响应率。此外,纯文本测试下模型安全性更高。最后,我们探索自动化评估,发现当前最佳安全分类器仍存在明显不足。
原文摘要 · Abstract (English)
Vision-language models (VLMs), which process image and text inputs, are increasingly integrated into chat assistants and other consumer AI applications. Without proper safeguards, however, VLMs may give harmful advice (e.g. how to self-harm) or encourage unsafe behaviours (e.g. to consume drugs). Despite these clear hazards, little work so far has evaluated VLM safety and the novel risks created by multimodal inputs. To address this gap, we introduce MSTS, a Multimodal Safety Test Suite for VLMs. MSTS comprises 400 test prompts across 40 fine-grained hazard categories. Each test prompt consists of a text and an image that only in combination reveal their full unsafe meaning. With MSTS, we find clear safety issues in several open VLMs. We also find some VLMs to be safe by accident, meaning that they are safe because they fail to understand even simple test prompts. We translate MSTS into ten languages, showing non-English prompts to increase the rate of unsafe model responses. We also show models to be safer when tested with text only rather than multimodal prompts. Finally, we explore the automation of VLM safety assessments, finding even the best safety classifiers to be lacking.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。