构建六类敏感内容统一数据集,提升社交媒体内容过滤效果。
Sensitive Content Classification in Social Media: A Holistic Resource and Evaluation
- 构建涵盖六类敏感内容的统一数据集,解决以往研究聚焦单一类别问题。
- 微调大模型后检测准确率比LLaMA等开源模型高10-15%。
- 适合需要定制化内容审核的平台或研究者使用。
大规模数据中敏感内容的检测对确保共享与分析数据的安全性至关重要。然而,现有审核工具(如外部API)存在可定制性差、跨敏感类别准确性不足及隐私隐患等问题。现有数据集和开源模型主要关注毒性语言,对药物滥用、自残等敏感内容覆盖不足。本文提出一个统一数据集,涵盖冲突性语言、脏话、色情内容、毒品相关、自残行为和垃圾信息六类敏感内容。通过一致的采集策略与标注规范,弥补此前研究的局限性。实验表明,在该数据集上微调大语言模型(LLMs)后,检测性能显著优于Llama等开源模型,甚至低于主流商业OpenAI模型10-15%。现有主流审核API亦难以针对特定敏感类别进行有效适配。
原文摘要 · Abstract (English)
The detection of sensitive content in large datasets is crucial for ensuring that shared and analysed data is free from harmful material. However, current moderation tools, such as external APIs, suffer from limitations in customisation, accuracy across diverse sensitive categories, and privacy concerns. Additionally, existing datasets and open-source models focus predominantly on toxic language, leaving gaps in detecting other sensitive categories such as substance abuse or self-harm. In this paper, we put forward a unified dataset tailored for social media content moderation across six sensitive categories: conflictual language, profanity, sexually explicit material, drug-related content, self-harm, and spam. By collecting and annotating data with consistent retrieval strategies and guidelines, we address the shortcomings of previous focalised research. Our analysis demonstrates that fine-tuning large language models (LLMs) on this novel dataset yields significant improvements in detection performance compared to open off-the-shelf models such as LLaMA, and even proprietary OpenAI models, which underperform by 10-15% overall. This limitation is even more pronounced on popular moderation APIs, which cannot be easily tailored to specific sensitive content categories, among others.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。