动态生成安全评测数据,揭示多模态大模型真实安全风险
SDEval: Safety Dynamic Evaluation for Multimodal Large Language Models
- 通过文本、图像及图文混合动态调整生成新评测样本
- 实测显示动态样本可暴露模型隐藏安全漏洞
- 适用于多种安全与能力基准,防数据污染
在多模态大语言模型(MLLMs)快速发展的背景下,其输出的安全性备受关注。尽管已有大量评测数据集,但它们易随模型进步而过时,且存在数据污染问题。为此,我们提出首个安全动态评估框架SDEval,可可控地调节安全评测基准的分布与复杂度。SDEval采用文本、图像及文本-图像动态三种策略,从原始基准生成新样本。我们首先分析了文本与图像动态对模型安全的独立影响,发现将文本动态注入图像或反之,均会加剧安全风险。SDEval具有通用性,可应用于多种现有安全与能力基准。在MLLMGuard、VLSBench、MMBench和MMVet等基准上的实验表明,SDEval显著影响安全评估结果,缓解数据污染,并暴露MLLMs的安全短板。代码已开源。
原文摘要 · Abstract (English)
In the rapidly evolving landscape of Multimodal Large Language Models (MLLMs), the safety concerns of their outputs have earned significant attention. Although numerous datasets have been proposed, they may become outdated with MLLM advancements and are susceptible to data contamination issues. To address these problems, we propose \textbf{SDEval}, the \textit{first} safety dynamic evaluation framework to controllably adjust the distribution and complexity of safety benchmarks. Specifically, SDEval mainly adopts three dynamic strategies: text, image, and text-image dynamics to generate new samples from original benchmarks. We first explore the individual effects of text and image dynamics on model safety. Then, we find that injecting text dynamics into images can further impact safety, and conversely, injecting image dynamics into text also leads to safety risks. SDEval is general enough to be applied to various existing safety and even capability benchmarks. Experiments across safety benchmarks, MLLMGuard and VLSBench, and capability benchmarks, MMBench and MMVet, show that SDEval significantly influences safety evaluation, mitigates data contamination, and exposes safety limitations of MLLMs. Code is available at https://github.com/hq-King/SDEval
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。