arXiv:2409.19734cs.CV2024-09NeurIPS被引 8

构建11K规模多模态有害内容数据集,提升模型对复杂有害场景的识别能力。

T2Vs Meet VLMs: A Scalable Multimodal Dataset for Visual Harmfulness Recognition

  • 用多智能体VQA框架让视觉大模型辩论判断是否有害,增强上下文理解。
  • 数据集含1万图像1千视频,覆盖10类真实与生成的有害内容,全谱系覆盖。
  • 相比现有数据集,显著提升检测模型在边缘案例中的准确率和泛化能力。

为应对不当或有害内容带来的风险,研究者尝试结合多个有害内容数据集与机器学习方法进行检测。然而,现有数据集仅涵盖少数特定有害对象,且仅包含真实内容,限制了方法的泛化能力,易导致误判。为此,我们提出一个全面的有害内容数据集——视觉有害数据集11K(VHD11K),包含10,000张图像和1,000段视频,源自互联网爬取及4种生成模型,覆盖10个有害类别,涵盖广泛且具有非平凡定义的有害概念。我们还提出一种新型标注框架,将标注过程建模为多智能体视觉问答(VQA)任务,由3个不同视觉语言模型(VLMs)“辩论”判断图像/视频是否有害,并在辩论中引入上下文学习策略。这确保了模型在决策前充分考虑上下文与多方观点,降低边缘情况下的误判概率。评估结果表明:(1) 该框架的标注结果与人工标注高度一致,保障了数据集可靠性;(2) 本数据集揭示了现有有害内容检测方法在广泛有害内容上的识别盲区,并提升了现有方法性能;(3) 与基线数据集SMID相比,VHD11K在有害性识别方法上实现更优提升。完整数据集与代码见https://github.com/nctu-eva-lab/VHD11K。

原文摘要 · Abstract (English)

To address the risks of encountering inappropriate or harmful content, researchers managed to incorporate several harmful contents datasets with machine learning methods to detect harmful concepts. However, existing harmful datasets are curated by the presence of a narrow range of harmful objects, and only cover real harmful content sources. This hinders the generalizability of methods based on such datasets, potentially leading to misjudgments. Therefore, we propose a comprehensive harmful dataset, Visual Harmful Dataset 11K (VHD11K), consisting of 10,000 images and 1,000 videos, crawled from the Internet and generated by 4 generative models, across a total of 10 harmful categories covering a full spectrum of harmful concepts with nontrivial definition. We also propose a novel annotation framework by formulating the annotation process as a multi-agent Visual Question Answering (VQA) task, having 3 different VLMs "debate" about whether the given image/video is harmful, and incorporating the in-context learning strategy in the debating process. Therefore, we can ensure that the VLMs consider the context of the given image/video and both sides of the arguments thoroughly before making decisions, further reducing the likelihood of misjudgments in edge cases. Evaluation and experimental results demonstrate that (1) the great alignment between the annotation from our novel annotation framework and those from human, ensuring the reliability of VHD11K; (2) our full-spectrum harmful dataset successfully identifies the inability of existing harmful content detection methods to detect extensive harmful contents and improves the performance of existing harmfulness recognition methods; (3) VHD11K outperforms the baseline dataset, SMID, as evidenced by the superior improvement in harmfulness recognition methods. The complete dataset and code can be found at https://github.com/nctu-eva-lab/VHD11K.

有害检测多模态数据集VLM

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。