arXiv:2501.09039cs.CRcs.AI2025-01被引 3

揭露视觉语言大模型在恶意提示下的毒性漏洞,警示安全风险。

Playing Devil's Advocate: Unmasking Toxicity and Vulnerabilities in Large Vision-Language Models

论文配图:Playing Devil's Advocate: Unmasking Toxicity and Vulnerabilities in Large Vision-Language Models
图 1 · 摘自论文原文
  • 用模拟社会操纵的对抗性提示测试模型漏洞
  • 最高毒性强提示下,部分模型毒性响应率达21.5%
  • 即使经安全微调,模型仍易被诱导生成不当内容

大型视觉语言模型(LVLMs)虽具广泛应用潜力,但存在生成有害内容的脆弱性。本研究系统评估了LLaVA、InstructBLIP、Fuyu和Qwen等开源模型,采用受社会理论启发的对抗性提示策略,模拟真实世界中的社交操控。结果显示:(i)毒性与侮辱行为最常见,平均率分别为16.13%和9.75%;(ii)Qwen-VL-Chat、LLaVA-v1.6-Vicuna-7b和InstructBLIP-Vicuna-7b最易受影响,毒性响应率分别达21.50%、18.30%和17.90%,侮辱率分别为13.40%、11.70%和10.10%;(iii)融入黑色幽默和多模态有毒补全的提示显著加剧漏洞。即便经过安全微调,模型仍会因对抗输入生成不同程度的有毒内容,凸显亟需强化安全机制与鲁棒防护。

原文摘要 · Abstract (English)

The rapid advancement of Large Vision-Language Models (LVLMs) has enhanced capabilities offering potential applications from content creation to productivity enhancement. Despite their innovative potential, LVLMs exhibit vulnerabilities, especially in generating potentially toxic or unsafe responses. Malicious actors can exploit these vulnerabilities to propagate toxic content in an automated (or semi-) manner, leveraging the susceptibility of LVLMs to deception via strategically crafted prompts without fine-tuning or compute-intensive procedures. Despite the red-teaming efforts and inherent potential risks associated with the LVLMs, exploring vulnerabilities of LVLMs remains nascent and yet to be fully addressed in a systematic manner. This study systematically examines the vulnerabilities of open-source LVLMs, including LLaVA, InstructBLIP, Fuyu, and Qwen, using adversarial prompt strategies that simulate real-world social manipulation tactics informed by social theories. Our findings show that (i) toxicity and insulting are the most prevalent behaviors, with the mean rates of 16.13% and 9.75%, respectively; (ii) Qwen-VL-Chat, LLaVA-v1.6-Vicuna-7b, and InstructBLIP-Vicuna-7b are the most vulnerable models, exhibiting toxic response rates of 21.50%, 18.30% and 17.90%, and insulting responses of 13.40%, 11.70% and 10.10%, respectively; (iii) prompting strategies incorporating dark humor and multimodal toxic prompt completion significantly elevated these vulnerabilities. Despite being fine-tuned for safety, these models still generate content with varying degrees of toxicity when prompted with adversarial inputs, highlighting the urgent need for enhanced safety mechanisms and robust guardrails in LVLM development.

视觉语言模型安全漏洞对抗性提示毒性检测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。