用大模型少样本学习实现高效动态内容审核,提升安全性和可扩展性。
Towards Safer Social Media Platforms: Scalable and Performant Few-Shot Harmful Content Moderation Using Large Language Models
- 基于大模型的上下文学习,仅需少量示例即可完成内容审核
- 在多个数据集上超越现有商用和前沿方法的表现
- 融合视频缩略图信息,探索多模态增强效果
社交媒体上有害内容泛滥对用户与社会构成重大风险,亟需更有效、可扩展的审核策略。现有方法依赖人工审核、有监督分类器及大量训练数据,在可扩展性、主观性以及内容动态变化(如暴力内容、危险挑战趋势等)方面存在局限。为此,本文利用大语言模型(LLM)通过上下文学习实现少样本动态内容审核。在多个LLM上的广泛实验表明,该少样本方法在识别危害内容方面优于现有商用基准(Perspective和OpenAI Moderation)以及先前最先进的少样本学习方法。同时引入视觉信息(视频缩略图),评估不同多模态技术对模型性能的影响。结果表明,基于LLM的方法在实现可扩展、动态的内容审核方面具有显著优势。
原文摘要 · Abstract (English)
The prevalence of harmful content on social media platforms poses significant risks to users and society, necessitating more effective and scalable content moderation strategies. Current approaches rely on human moderators, supervised classifiers, and large volumes of training data, and often struggle with scalability, subjectivity, and the dynamic nature of harmful content (e.g., violent content, dangerous challenge trends, etc.). To bridge these gaps, we utilize Large Language Models (LLMs) to undertake few-shot dynamic content moderation via in-context learning. Through extensive experiments on multiple LLMs, we demonstrate that our few-shot approaches can outperform existing proprietary baselines (Perspective and OpenAI Moderation) as well as prior state-of-the-art few-shot learning methods, in identifying harm. We also incorporate visual information (video thumbnails) and assess if different multimodal techniques improve model performance. Our results underscore the significant benefits of employing LLM based methods for scalable and dynamic harmful content moderation online.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。