评测多轮对话中多模态模型的安全性,发现轮次越多越易被攻破。
SafeMT: Multi-turn Safety for Multimodal Language Models
- 构建包含1万条样本的多轮安全测试集,覆盖17种场景和4种绕过方法。
- 首次发现模型在多轮有害对话中攻击成功率随轮次增加而上升。
- 提出对话安全监控器,有效降低开源模型的多轮攻击成功率。
随着多模态大语言模型(MLLMs)的广泛应用,安全性问题日益突出。相较于单轮提示,日常交互中的多轮对话风险更高,但现有基准未能充分覆盖此场景。为此,我们提出SafeMT基准,包含由有害查询与图像生成的多轮对话共10,000个样本,涵盖17种不同场景及4种越狱方法。我们还引入安全指数(SI)评估模型在对话中的整体安全性。对17个模型的评估显示,有害对话轮次越多,成功攻击的概率越高,表明现有模型在识别对话中潜在危险方面能力不足。我们进一步提出一种对话安全监控器,可识别隐藏在对话中的恶意意图,并为模型提供相应安全策略。实验表明,该监控器在多个开源模型上比现有防护机制更有效地降低多轮攻击成功率。
原文摘要 · Abstract (English)
With the widespread use of multi-modal Large Language models (MLLMs), safety issues have become a growing concern. Multi-turn dialogues, which are more common in everyday interactions, pose a greater risk than single prompts; however, existing benchmarks do not adequately consider this situation. To encourage the community to focus on the safety issues of these models in multi-turn dialogues, we introduce SafeMT, a benchmark that features dialogues of varying lengths generated from harmful queries accompanied by images. This benchmark consists of 10,000 samples in total, encompassing 17 different scenarios and four jailbreak methods. Additionally, we propose Safety Index (SI) to evaluate the general safety of MLLMs during conversations. We assess the safety of 17 models using this benchmark and discover that the risk of successful attacks on these models increases as the number of turns in harmful dialogues rises. This observation indicates that the safety mechanisms of these models are inadequate for recognizing the hazard in dialogue interactions. We propose a dialogue safety moderator capable of detecting malicious intent concealed within conversations and providing MLLMs with relevant safety policies. Experimental results from several open-source models indicate that this moderator is more effective in reducing multi-turn ASR compared to existed guard models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。