arXiv:2509.15478cs.CL2025-09被引 1

测试4种多模态大模型在文本和图文提示下的安全漏洞,发现模型差异大、文本提示更易绕过防护。

Red Teaming Multimodal Language Models: Evaluating Harm Across Prompt Modalities and Models

  • 26人生成726个对抗性提示,覆盖三类危害行为
  • Pixtral 12B有害响应率达62%,Claude Sonnet 3.5仅10%
  • 文本提示比图文提示更易触发安全漏洞,适合安全研究者参考

多模态大语言模型(MLLMs)在真实场景中应用日益广泛,但其在对抗性提示下的安全性仍缺乏深入研究。本研究评估了四种主流模型(GPT-4o、Claude Sonnet 3.5、Pixtral 12B、Qwen VL Plus)在纯文本与多模态提示下的安全性。由26名红队成员生成726个针对非法活动、虚假信息和不道德行为的对抗性提示,提交给各模型后,由17名标注员对2,904条输出进行有害性评分(5分制)。结果表明,不同模型间脆弱性差异显著:Pixtral 12B有害响应率约62%,而Claude Sonnet 3.5最抗攻击,仅约10%。统计分析证实,模型类型与输入模态均为有害性的显著预测因子。出乎意料的是,纯文本提示略优于多模态提示,能更有效绕过安全机制。研究强调,亟需建立全面的多模态安全评估基准以支撑模型部署。

原文摘要 · Abstract (English)

Multimodal large language models (MLLMs) are increasingly used in real world applications, yet their safety under adversarial conditions remains underexplored. This study evaluates the harmlessness of four leading MLLMs (GPT-4o, Claude Sonnet 3.5, Pixtral 12B, and Qwen VL Plus) when exposed to adversarial prompts across text-only and multimodal formats. A team of 26 red teamers generated 726 prompts targeting three harm categories: illegal activity, disinformation, and unethical behaviour. These prompts were submitted to each model, and 17 annotators rated 2,904 model outputs for harmfulness using a 5-point scale. Results show significant differences in vulnerability across models and modalities. Pixtral 12B exhibited the highest rate of harmful responses (~62%), while Claude Sonnet 3.5 was the most resistant (~10%). Contrary to expectations, text-only prompts were slightly more effective at bypassing safety mechanisms than multimodal ones. Statistical analysis confirmed that both model type and input modality were significant predictors of harmfulness. These findings underscore the urgent need for robust, multimodal safety benchmarks as MLLMs are deployed more widely.

多模态安全红队测试模型评测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。