评测多轮对话中图文大模型的安全性,发现模型常漏判渐进风险或误拒正常对话。
MTMCS-Bench: Evaluating Contextual Safety of Multimodal Large Language Models in Multi-Turn Dialogues
- 构建多轮图文安全评测基准,涵盖渐进风险与场景切换风险
- 超3万组图文/纯文本样本,量化评估意图识别与安全响应能力
- 揭示安全与实用性的固有权衡,现有防护机制仍不完善
图文大模型作为交互式助手日益普及,其安全性需同时考虑视觉场景与对话演化。现有评测多为单轮,难以捕捉恶意意图的渐进显现或同一场景下的善恶双面性。本文提出多轮多模态情境安全评测基准(MTMCS-Bench),包含超过30,000组多模态(图像+文本)与单模态(仅文本)样本,覆盖“渐进风险”与“上下文切换风险”两类场景,支持配对安全/非安全对话的结构化评估。评测指标分别衡量情境意图识别能力、非安全情形下的安全意识、以及良性对话中的帮助性。在8个开源与7个专有图文大模型上测试发现,模型普遍存在安全与实用性之间的持续权衡:要么忽略渐进风险,要么过度拒绝正常对话。进一步评估5种现有防护机制,结果表明它们可缓解部分问题,但无法完全解决多轮情境下的安全风险。
原文摘要 · Abstract (English)
Multimodal large language models (MLLMs) are increasingly deployed as assistants that interact through text and images, making it crucial to evaluate contextual safety when risk depends on both the visual scene and the evolving dialogue. Existing contextual safety benchmarks are mostly single-turn and often miss how malicious intent can emerge gradually or how the same scene can support both benign and exploitative goals. We introduce the Multi-Turn Multimodal Contextual Safety Benchmark (MTMCS-Bench), a benchmark of realistic images and multi-turn conversations that evaluates contextual safety in MLLMs under two complementary settings, escalation-based risk and context-switch risk. MTMCS-Bench offers paired safe and unsafe dialogues with structured evaluation. It contains over 30 thousand multimodal (image+text) and unimodal (text-only) samples, with metrics that separately measure contextual intent recognition, safety-awareness on unsafe cases, and helpfulness on benign ones. Across eight open-source and seven proprietary MLLMs, we observe persistent trade-offs between contextual safety and utility, with models tending to either miss gradual risks or over-refuse benign dialogues. Finally, we evaluate five current guardrails and find that they mitigate some failures but do not fully resolve multi-turn contextual risks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。