arXiv:2601.10527cs.AIcs.CL2026-01被引 6

评测6大前沿模型安全表现,发现多数模型在对抗攻击下极易失效。

A Safety Report on GPT-5.2, Gemini 3 Pro, Qwen3-VL, Grok 4.1 Fast, Nano Banana Pro, and Seedream 4.5

  • 采用统一协议评估语言、视觉与生成多模态安全性能。
  • 所有模型在对抗测试中安全率低于6%,严重脆弱。
  • 模型间安全表现差异大,需综合评估避免误判。

大型语言模型(LLMs)和多模态大语言模型(MLLMs)在推理、感知与生成方面取得显著进展,但其安全性是否同步提升尚不明确,主要因评估碎片化,仅关注单一模态或威胁场景。本报告对六款前沿模型——GPT-5.2、Gemini 3 Pro、Qwen3-VL、Grok 4.1 Fast、Nano Banana Pro 和 Seedream 4.5——进行集成式安全评估,涵盖语言、视觉语言及图像生成任务,采用统一协议融合基准测试、对抗测试、多语言评估与合规性测试。通过构建安全排行榜与模型画像,揭示出安全表现高度不均衡:尽管 GPT-5.2 展现出一致且均衡的强安全性能,其他模型在基准安全、对抗鲁棒性、多语言泛化与监管合规性之间存在明显权衡。即便在标准基准上表现良好,所有模型在对抗测试中仍极为脆弱,最差情况下的安全率低于6%。文本到图像模型在受控视觉风险类别中对齐稍强,但在面对对抗性或语义模糊提示时依然脆弱。总体而言,前沿模型的安全性具有内在多维性,受模态、语言与评估设计影响,强调必须采用标准化、全面性的安全评估以更真实反映现实风险,指导负责任部署。

原文摘要 · Abstract (English)

The rapid evolution of Large Language Models (LLMs) and Multimodal Large Language Models (MLLMs) has driven major gains in reasoning, perception, and generation across language and vision, yet whether these advances translate into comparable improvements in safety remains unclear, partly due to fragmented evaluations that focus on isolated modalities or threat models. In this report, we present an integrated safety evaluation of six frontier models--GPT-5.2, Gemini 3 Pro, Qwen3-VL, Grok 4.1 Fast, Nano Banana Pro, and Seedream 4.5--assessing each across language, vision-language, and image generation using a unified protocol that combines benchmark, adversarial, multilingual, and compliance evaluations. By aggregating results into safety leaderboards and model profiles, we reveal a highly uneven safety landscape: while GPT-5.2 demonstrates consistently strong and balanced performance, other models exhibit clear trade-offs across benchmark safety, adversarial robustness, multilingual generalization, and regulatory compliance. Despite strong results under standard benchmarks, all models remain highly vulnerable under adversarial testing, with worst-case safety rates dropping below 6%. Text-to-image models show slightly stronger alignment in regulated visual risk categories, yet remain fragile when faced with adversarial or semantically ambiguous prompts. Overall, these findings highlight that safety in frontier models is inherently multidimensional--shaped by modality, language, and evaluation design--underscoring the need for standardized, holistic safety assessments to better reflect real-world risk and guide responsible deployment.

模型安全对抗攻击多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。