用协同标签提升细粒度毒 meme 检测效果
STEMTOX: From Collaborative Tags to Fine-Grained Toxic Meme Detection via Entropy-Guided Multi-Task Learning
- 引入协同标签增强图文上下文,辅助毒性判断
- 在 6300 条真实 meme 上构建细粒度标注数据集
- 适合内容安全、多模态分析方向研究者参考
表情包作为网络沟通的常见形式,常被用于传播有害内容。然而,数据获取受限和标注成本高昂阻碍了鲁棒性表情包审核系统的发展。为此,本文首次构建了一个名为 TOXICTAGS 的数据集,包含 6300 条真实世界表情包帖子,分两阶段标注:(i) 毒性/正常二分类,(ii) 细粒度标注为仇恨、危险或冒犯类。该数据集的关键特征是保留原始帖子的协同标签,丰富每条表情包的上下文信息。此外,提出一种熵引导的多任务学习框架 STEMTOX,利用协同标签与视觉、文本输入,在统一分类框架中实现更精准的毒性识别。实验表明,引入协同标签显著提升了现有 VLMs 在毒性检测任务中的表现。本工作为多模态在线环境下的内容审核提供了可扩展的新范式。注意:部分内容可能具有毒性。
原文摘要 · Abstract (English)
Memes, as a widely used mode of online communication, often serve as vehicles for spreading harmful content. However, limitations in data accessibility and the high costs of dataset curation hinder the development of robust meme moderation systems. To address this challenge, in this work, we introduce a first-of-its-kind dataset - TOXICTAGS consisting of 6,300 real-world meme-based posts annotated in two stages: (i) binary classification into toxic and normal, and (ii) fine-grained labelling of toxic memes as hateful, dangerous, or offensive. A key feature of this dataset is that it includes collaborative tags associated with the original posts, enhancing the context of each meme. In addition, we propose a novel entropy-guided multi-tasking framework -- STEMTOX -- that leverages these collaborative tags alongside visual and textual inputs within a robust classification framework. Experimental results show that incorporating these tags substantially enhances the performance of state-of-the-art VLMs in toxicity detection tasks. Our contributions offer a novel and scalable foundation for improved content moderation in multimodal online environments. Warning: Contains potentially toxic contents.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。