arXiv:2603.02482cs.LGcs.CL2026-03中稿 · EMNLP

MUSE平台实现多模态大模型安全评估的细粒度追踪与分析。

MUSE: A Run-Centric Platform for Multimodal Unified Safety Evaluation of Large Language Models

  • 以单次攻击行为为单元,全程记录交互轨迹与模态变化。
  • 文本攻击仅3.1%成功,迭代攻击效果显著提升。
  • 支持细粒度安全判定,适合研究者与开发者评估模型安全性。

多模态大模型的安全评估不仅需判断攻击是否成功,还需追踪跨轮次和多模态交互过程。我们提出 MUSE(Multimodal Unified Safety Evaluation),一个开源、基于浏览器的运行中心化平台,将每次攻击运行作为持久的执行、检查与分析单元,保留配置、多轮轨迹、输入模态与媒体、目标响应及安全判断。五级响应分类区分完全合规、部分合规与拒绝行为,得出硬ASR、软ASR和灰区宽度(GZW)。在六个多模态大模型上共完成11,700次评估,纯文本请求仅导致3.1%宏观硬ASR和4.4%软ASR,而迭代攻击策略更有效;攻击效果也显著依赖攻击者主干模型。相反,跨轮模态切换(ITMS)作为受控模态投递探测,并未持续提升攻击成功率。结果表明,运行中心化的细粒度评估对刻画多模态安全行为具有重要价值。

原文摘要 · Abstract (English)

Safety evaluation of multimodal large language models requires tracking not only whether an attack succeeds, but also how the interaction unfolds across turns and input modalities. We present MUSE (Multimodal Unified Safety Evaluation), an open-source, browser-based, run-centric platform for multimodal safety evaluation. MUSE treats each attack run as the persistent unit of execution, inspection, and analysis, preserving its configuration, multi-turn trajectory, delivered modalities and media, target responses, and safety judgments. A five-level response taxonomy further distinguishes full Compliance from Partial Compliance and refusal behavior, yielding hard ASR, soft ASR, and gray-zone width (GZW). Across 11,700 evaluations on six multimodal LLMs, direct text-only requests yield only 3.1% macro hard ASR and 4.4% soft ASR, while iterative attack procedures are substantially more effective. Attack effectiveness also varies substantially with the attacker backbone. In contrast, Inter-Turn Modality Switching (ITMS), evaluated as a controlled delivery-modality probe, does not consistently increase attack success. These results demonstrate the value of run-centric, fine-grained evaluation for characterizing multimodal safety behavior beyond a single binary success metric.

大模型安全多模态评估运行追踪攻击测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。