开源框架OpenRT系统评估多模态大模型安全漏洞,发现顶尖模型仍易被攻破。
OpenRT: An Open-Source Red Teaming Framework for Multimodal LLMs
- 构建模块化对抗框架,分离攻击策略与运行机制
- 20个前沿模型测试中平均攻击成功率高达49.14%
- 适合安全研究者、模型开发者用于系统性漏洞检测
多模态大语言模型(MLLMs)在关键应用中的快速部署正面临持续的安全漏洞困扰。现有红队测试基准普遍碎片化,局限于单轮文本交互,缺乏系统评估所需的可扩展性。为此,我们提出OpenRT——一个统一、模块化且高吞吐的红队测试框架,用于全面评估MLLM安全性。其核心在于引入对抗内核,实现模型集成、数据管理、攻击策略、判断方法和评估指标五个关键维度的模块分离。通过标准化攻击接口,将对抗逻辑与高吞吐异步运行时解耦,支持跨多种模型的系统化扩展。框架整合了37种多样化攻击方法,涵盖白盒梯度、多模态扰动及复杂多智能体演化策略。在20个先进模型(包括GPT-5.2、Claude 4.5、Gemini 3 Pro)上的实证研究表明:即使前沿模型也无法在不同攻击范式间泛化,领先模型平均攻击成功率高达49.14%。值得注意的是,推理模型并不天然具备更强的抗复杂多轮越狱能力。通过开源OpenRT,我们提供可持续、可扩展且持续维护的基础设施,加速AI安全的发展与标准化。
原文摘要 · Abstract (English)
The rapid integration of Multimodal Large Language Models (MLLMs) into critical applications is increasingly hindered by persistent safety vulnerabilities. However, existing red-teaming benchmarks are often fragmented, limited to single-turn text interactions, and lack the scalability required for systematic evaluation. To address this, we introduce OpenRT, a unified, modular, and high-throughput red-teaming framework designed for comprehensive MLLM safety evaluation. At its core, OpenRT architects a paradigm shift in automated red-teaming by introducing an adversarial kernel that enables modular separation across five critical dimensions: model integration, dataset management, attack strategies, judging methods, and evaluation metrics. By standardizing attack interfaces, it decouples adversarial logic from a high-throughput asynchronous runtime, enabling systematic scaling across diverse models. Our framework integrates 37 diverse attack methodologies, spanning white-box gradients, multi-modal perturbations, and sophisticated multi-agent evolutionary strategies. Through an extensive empirical study on 20 advanced models (including GPT-5.2, Claude 4.5, and Gemini 3 Pro), we expose critical safety gaps: even frontier models fail to generalize across attack paradigms, with leading models exhibiting average Attack Success Rates as high as 49.14%. Notably, our findings reveal that reasoning models do not inherently possess superior robustness against complex, multi-turn jailbreaks. By open-sourcing OpenRT, we provide a sustainable, extensible, and continuously maintained infrastructure that accelerates the development and standardization of AI safety.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。