8次模型更新中,多模态大模型安全性能随版本迭代发生显著漂移。
Alignment Drift in Multimodal LLMs: A Two-Phase, Longitudinal Evaluation of Harm Across Eight Model Releases
- 用固定攻击提示集纵向评估8个版本模型的安全性。
- GPT和Claude模型攻击成功率上升,而Pixtral与Qwen下降。
- 文本提示在早期更有效,后期各模型表现趋同,需长期监测。
多模态大语言模型(MLLMs)日益应用于真实系统,但其在对抗性提示下的安全性仍研究不足。本文采用由26名专业红队成员编写的726个固定攻击提示,开展两阶段评估:第一阶段测试GPT-4o、Claude Sonnet 3.5、Pixtral 12B和Qwen VL Plus;第二阶段评估其后续版本(GPT-5、Claude Sonnet 4.5、Pixtral Large、Qwen Omni),共获得82,256条人工危害评分。不同模型家族间存在显著差异:Pixtral模型始终最易受攻击,Claude模型因高拒绝率表现最安全。攻击成功率(ASR)显示明显对齐漂移:GPT与Claude模型跨代升级后ASR上升,而Pixtral与Qwen略有下降。模态影响也随时间变化:第一阶段纯文本提示更有效,第二阶段则出现模型特异性模式,GPT-5与Claude 4.5在各模态下接近同等脆弱。结果表明,MLLM的安全性并非静态或统一,亟需长期、多模态基准持续追踪其演化行为。
原文摘要 · Abstract (English)
Multimodal large language models (MLLMs) are increasingly deployed in real-world systems, yet their safety under adversarial prompting remains underexplored. We present a two-phase evaluation of MLLM harmlessness using a fixed benchmark of 726 adversarial prompts authored by 26 professional red teamers. Phase 1 assessed GPT-4o, Claude Sonnet 3.5, Pixtral 12B, and Qwen VL Plus; Phase 2 evaluated their successors (GPT-5, Claude Sonnet 4.5, Pixtral Large, and Qwen Omni) yielding 82,256 human harm ratings. Large, persistent differences emerged across model families: Pixtral models were consistently the most vulnerable, whereas Claude models appeared safest due to high refusal rates. Attack success rates (ASR) showed clear alignment drift: GPT and Claude models exhibited increased ASR across generations, while Pixtral and Qwen showed modest decreases. Modality effects also shifted over time: text-only prompts were more effective in Phase 1, whereas Phase 2 produced model-specific patterns, with GPT-5 and Claude 4.5 showing near-equivalent vulnerability across modalities. These findings demonstrate that MLLM harmlessness is neither uniform nor stable across updates, underscoring the need for longitudinal, multimodal benchmarks to track evolving safety behaviour.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。