大模型显著提升图文视频敏感内容识别准确率,降低误判漏判。
Advancing Content Moderation: Evaluating Large Language Models for Detecting Sensitive Content Across Text, Images, and Videos
- 用大模型融合文本与视觉上下文理解,提升敏感内容检测能力。
- 相比传统方法,错误率下降,准确率显著提高。
- 适合社交平台、视频网站等需要大规模内容审核的场景。
仇恨言论、骚扰、有害及性内容、暴力信息在网站和媒体平台上的广泛传播引发社会各界广泛关注。政府、教育工作者与家长常与平台就内容监管方式产生分歧。自动检测与过滤技术是应对该挑战的关键方案。自然语言处理与计算机视觉技术被广泛用于识别文本、图像、视频中的攻击性语言、暴力、裸露、成瘾等内容,支持平台规模化执行内容政策。然而,现有方法在高精度检测与低误报、漏报之间仍存局限。因此,更先进的上下文理解算法有望推动内容审查系统升级。本文评估了基于大模型的内容审核方案,如OpenAI内容审核模型和Llama-Guard3,并探索GPT、Gemini、Llama等主流大模型在跨媒体敏感内容识别中的表现。使用X推文、亚马逊评论、新闻文章、真人照片、卡通、素描、暴力视频等多类数据集进行测试。结果表明,大模型在准确率上优于传统方法,同时显著降低误报与漏报率,展现出将其集成至网站、社交媒体及视频分享服务中用于监管与内容管理的巨大潜力。
原文摘要 · Abstract (English)
The widespread dissemination of hate speech, harassment, harmful and sexual content, and violence across websites and media platforms presents substantial challenges and provokes widespread concern among different sectors of society. Governments, educators, and parents are often at odds with media platforms about how to regulate, control, and limit the spread of such content. Technologies for detecting and censoring the media contents are a key solution to addressing these challenges. Techniques from natural language processing and computer vision have been used widely to automatically identify and filter out sensitive content such as offensive languages, violence, nudity, and addiction in both text, images, and videos, enabling platforms to enforce content policies at scale. However, existing methods still have limitations in achieving high detection accuracy with fewer false positives and false negatives. Therefore, more sophisticated algorithms for understanding the context of both text and image may open rooms for improvement in content censorship to build a more efficient censorship system. In this paper, we evaluate existing LLM-based content moderation solutions such as OpenAI moderation model and Llama-Guard3 and study their capabilities to detect sensitive contents. Additionally, we explore recent LLMs such as GPT, Gemini, and Llama in identifying inappropriate contents across media outlets. Various textual and visual datasets like X tweets, Amazon reviews, news articles, human photos, cartoons, sketches, and violence videos have been utilized for evaluation and comparison. The results demonstrate that LLMs outperform traditional techniques by achieving higher accuracy and lower false positive and false negative rates. This highlights the potential to integrate LLMs into websites, social media platforms, and video-sharing services for regulatory and content moderation purposes.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。