清理了LLaVA数据集中7531对有毒图文对,提升多模态模型安全性。
Understanding and Mitigating Toxicity in Image-Text Pretraining Datasets: A Case Study on LLaVA
- 分析图文数据中毒性内容的分布与表现形式
- 移除7531对包含仇恨言论、裸露等有害内容的图文对
- 开源去毒数据集,助力负责任AI研发
预训练数据集是多模态模型发展的基础,但往往源自网络大规模语料,包含固有偏见和有害内容。本文针对LLaVA图像-文本预训练数据集中的毒性问题展开全面分析,研究有害内容在不同模态中的表现形式。我们梳理了常见的毒性类别,并提出针对性的缓解策略,最终构建了一个去毒优化的数据集,成功移除了7,531对有毒图像-文本配对。同时,我们提出了可落地的毒性检测流程实施指南。研究强调需主动识别并过滤仇恨言论、暴露内容及针对性骚扰等有害信息,以构建更负责任、公平的多模态系统。该去毒数据集已开源,可供后续研究使用。
原文摘要 · Abstract (English)
Pretraining datasets are foundational to the development of multimodal models, yet they often have inherent biases and toxic content from the web-scale corpora they are sourced from. In this paper, we investigate the prevalence of toxicity in LLaVA image-text pretraining dataset, examining how harmful content manifests in different modalities. We present a comprehensive analysis of common toxicity categories and propose targeted mitigation strategies, resulting in the creation of a refined toxicity-mitigated dataset. This dataset removes 7,531 of toxic image-text pairs in the LLaVA pre-training dataset. We offer guidelines for implementing robust toxicity detection pipelines. Our findings underscore the need to actively identify and filter toxic content - such as hate speech, explicit imagery, and targeted harassment - to build more responsible and equitable multimodal systems. The toxicity-mitigated dataset is open source and is available for further research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。