用真实梗图测试视觉语言模型安全,发现模型更易被梗图诱导产生有害输出。
Are Vision-Language Models Safe in the Wild? A Meme-Based Benchmark Study
- 构建包含5万张真实梗图的基准测试集,评估模型在真实场景下的安全表现。
- 相比纯文本输入,梗图使有害响应率上升,拒绝率显著下降。
- 多轮对话虽有缓解作用,但模型仍对梗图高度敏感,适合安全研究者参考。
视觉语言模型(VLMs)的快速部署放大了安全风险,但现有评估多依赖人工图像。本研究提出:当面对普通用户分享的真实梗图时,当前VLMs的安全性如何?为此,我们构建了包含50,430个实例的MemeSafetyBench基准,将真实梗图与有害及无害指令配对。基于全面的安全分类体系和大模型生成指令,我们评估了多个VLM在单轮与多轮交互中的表现。结果表明,相较于合成或文字类图像,模型在面对真实梗图时更易产生有害输出,且拒绝率显著降低。尽管多轮对话具备一定缓解作用,但脆弱性依然存在。研究强调需开展生态有效评估并加强安全机制。MemeSafetyBench已开源于https://github.com/oneonlee/Meme-Safety-Bench。
原文摘要 · Abstract (English)
Rapid deployment of vision-language models (VLMs) magnifies safety risks, yet most evaluations rely on artificial images. This study asks: How safe are current VLMs when confronted with meme images that ordinary users share? To investigate this question, we introduce MemeSafetyBench, a 50,430-instance benchmark pairing real meme images with both harmful and benign instructions. Using a comprehensive safety taxonomy and LLM-based instruction generation, we assess multiple VLMs across single and multi-turn interactions. We investigate how real-world memes influence harmful outputs, the mitigating effects of conversational context, and the relationship between model scale and safety metrics. Our findings demonstrate that VLMs are more vulnerable to meme-based harmful prompts than to synthetic or typographic images. Memes significantly increase harmful responses and decrease refusals compared to text-only inputs. Though multi-turn interactions provide partial mitigation, elevated vulnerability persists. These results highlight the need for ecologically valid evaluations and stronger safety mechanisms. MemeSafetyBench is publicly available at https://github.com/oneonlee/Meme-Safety-Bench.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。