arXiv:2503.14827cs.CLcs.AI2025-03被引 17

首个综合评估多模态大模型安全性的平台,发现多项潜在风险。

MMDT: Decoding the Trustworthiness and Safety of Multimodal Foundation Models

  • 构建多维度测评框架,覆盖安全、幻觉、偏见等六方面
  • 设计红队算法生成挑战性数据,形成高质量基准测试集
  • 适合模型开发者与安全研究人员参考,提升系统可靠性

多模态基础模型(MMFMs)在自动驾驶、医疗和虚拟助手等领域应用广泛,但现有研究揭示其存在生成不安全内容等漏洞。现有评测多聚焦于有用性或单一维度如公平性、隐私。本文提出首个统一平台MMDT(Multimodal DecodingTrust),从安全、幻觉、公平性/偏见、隐私、对抗鲁棒性和分布外泛化六个角度全面评估MMFMs。通过为各维度设计多样任务场景和红队算法,生成具有挑战性的测试数据,构建高质量基准。我们对多种多模态模型进行评估,发现多个维度存在显著缺陷与改进空间。本工作首次提供全面且独特的多模态模型安全与可信度评估体系,为构建更安全可靠的模型与系统奠定基础。平台及基准已开源:https://mmdecodingtrust.github.io/。

原文摘要 · Abstract (English)

Multimodal foundation models (MMFMs) play a crucial role in various applications, including autonomous driving, healthcare, and virtual assistants. However, several studies have revealed vulnerabilities in these models, such as generating unsafe content by text-to-image models. Existing benchmarks on multimodal models either predominantly assess the helpfulness of these models, or only focus on limited perspectives such as fairness and privacy. In this paper, we present the first unified platform, MMDT (Multimodal DecodingTrust), designed to provide a comprehensive safety and trustworthiness evaluation for MMFMs. Our platform assesses models from multiple perspectives, including safety, hallucination, fairness/bias, privacy, adversarial robustness, and out-of-distribution (OOD) generalization. We have designed various evaluation scenarios and red teaming algorithms under different tasks for each perspective to generate challenging data, forming a high-quality benchmark. We evaluate a range of multimodal models using MMDT, and our findings reveal a series of vulnerabilities and areas for improvement across these perspectives. This work introduces the first comprehensive and unique safety and trustworthiness evaluation platform for MMFMs, paving the way for developing safer and more reliable MMFMs and systems. Our platform and benchmark are available at https://mmdecodingtrust.github.io/.

多模态模型安全性评估可信度红队测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。