86支团队测试多模态大模型安全,发现漏洞并推动防御升级。
Pushing the Limits of Safety: A Technical Report on the ATLAS Challenge 2025
- 通过图像文本对抗攻击,分白盒黑盒两阶段测试模型漏洞。
- 多数模型在攻击下仍产生有害输出,安全防护仍存短板。
- 适合关注AI安全、对抗攻击与模型对齐的研究者参考。
多模态大语言模型(MLLMs)在众多应用中带来变革性进展,但仍易受安全威胁,尤其是诱导有害输出的越狱攻击。为系统评估并提升其安全性,我们组织了2025年对抗测试与大模型对齐安全挑战赛(ATLAS Challenge 2025)。本技术报告呈现竞赛成果,共86支团队参与,通过对抗性图像-文本攻击,在白盒和黑盒两种场景下评估了MLLM的安全性。结果显示,当前模型在面对复杂攻击时仍存在显著安全缺陷,暴露出防御机制的不足。该挑战建立了新的多模态大模型安全评估基准,为构建更安全的多模态AI系统奠定了基础。相关代码与数据已开源,地址:https://github.com/NY1024/ATLAS_Challenge_2025。
原文摘要 · Abstract (English)
Multimodal Large Language Models (MLLMs) have enabled transformative advancements across diverse applications but remain susceptible to safety threats, especially jailbreak attacks that induce harmful outputs. To systematically evaluate and improve their safety, we organized the Adversarial Testing & Large-model Alignment Safety Grand Challenge (ATLAS) 2025}. This technical report presents findings from the competition, which involved 86 teams testing MLLM vulnerabilities via adversarial image-text attacks in two phases: white-box and black-box evaluations. The competition results highlight ongoing challenges in securing MLLMs and provide valuable guidance for developing stronger defense mechanisms. The challenge establishes new benchmarks for MLLM safety evaluation and lays groundwork for advancing safer multimodal AI systems. The code and data for this challenge are openly available at https://github.com/NY1024/ATLAS_Challenge_2025.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。