提出新型隐性毒性评估基准,揭示大模型在隐性偏见上的脆弱性。
MDIT-Bench: Evaluating the Dual-Implicit Toxicity in Large Multimodal Models
- 通过多阶段人机协作生成方法构建含双重隐性毒性的数据集
- 涵盖317,638个问题,覆盖12类23子类780主题,分三级难度
- 首次量化模型在隐性毒性上的表现差距,适合安全与伦理研究者
大型多模态模型(LMMs)的广泛应用引发了对其毒性的担忧。然而,现有研究主要关注显性毒性,对隐性偏见和歧视等更隐蔽的毒性关注不足。为此,本文提出一种新型隐性毒性——双重隐性毒性,并构建了多模态双重隐性毒性评估基准MDIT-Bench。具体地,我们基于多阶段人机协同上下文生成方法创建了包含双重隐性毒性的MDIT-Dataset。在此基础上,构建了包含317,638个问题的MDIT-Bench,覆盖12个类别、23个子类别和780个主题,设有三个难度等级,并提出毒性差距度量指标以评估模型在不同难度下的敏感性。在13个主流LMMs上进行实验,结果表明这些模型难以有效应对双重隐性毒性,尤其在高难度任务中性能显著下降,揭示出模型仍存在大量可激活的隐藏毒性。数据已开源:https://github.com/nuo1nuo/MDIT-Bench。
原文摘要 · Abstract (English)
The widespread use of Large Multimodal Models (LMMs) has raised concerns about model toxicity. However, current research mainly focuses on explicit toxicity, with less attention to some more implicit toxicity regarding prejudice and discrimination. To address this limitation, we introduce a subtler type of toxicity named dual-implicit toxicity and a novel toxicity benchmark termed MDIT-Bench: Multimodal Dual-Implicit Toxicity Benchmark. Specifically, we first create the MDIT-Dataset with dual-implicit toxicity using the proposed Multi-stage Human-in-loop In-context Generation method. Based on this dataset, we construct the MDIT-Bench, a benchmark for evaluating the sensitivity of models to dual-implicit toxicity, with 317,638 questions covering 12 categories, 23 subcategories, and 780 topics. MDIT-Bench includes three difficulty levels, and we propose a metric to measure the toxicity gap exhibited by the model across them. In the experiment, we conducted MDIT-Bench on 13 prominent LMMs, and the results show that these LMMs cannot handle dual-implicit toxicity effectively. The model's performance drops significantly in hard level, revealing that these LMMs still contain a significant amount of hidden but activatable toxicity. Data are available at https://github.com/nuo1nuo/MDIT-Bench.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。