混合检测直播不良内容,兼顾已知违规与新型隐晦行为。
Dynamic Content Moderation in Livestreams: Combining Supervised Classification with MLLM-Boosted Similarity Matching
- 用分类+相似匹配双通道,识别已知与新型违规内容。
- 生产环境下分类召回率67%,相似匹配召回率76%。
- 适合需要实时、多模态内容审核的平台使用。
内容审核对大规模用户生成视频平台至关重要,尤其在直播环境中需及时、多模态且能应对不断演变的不良内容形式。我们提出一个在生产环境中部署的混合审核框架,结合监督分类处理已知违规,通过参考式相似匹配应对新型或隐蔽案例。两种路径均处理文本、音频、视觉多模态输入,利用多模态大语言模型(MLLM)知识蒸馏提升准确率,同时保持推理轻量。实际应用中,分类管道在80%精度下实现67%召回率,相似匹配管道在80%精度下实现76%召回率。大规模A/B测试显示,用户观看不良直播内容的比例降低6-8%。结果表明,该方法在可扩展性和适应性上表现优异,能有效应对显性违规与新兴对抗行为。
原文摘要 · Abstract (English)
Content moderation remains a critical yet challenging task for large-scale user-generated video platforms, especially in livestreaming environments where moderation must be timely, multimodal, and robust to evolving forms of unwanted content. We present a hybrid moderation framework deployed at production scale that combines supervised classification for known violations with reference-based similarity matching for novel or subtle cases. This hybrid design enables robust detection of both explicit violations and novel edge cases that evade traditional classifiers. Multimodal inputs (text, audio, visual) are processed through both pipelines, with a multimodal large language model (MLLM) distilling knowledge into each to boost accuracy while keeping inference lightweight. In production, the classification pipeline achieves 67% recall at 80% precision, and the similarity pipeline achieves 76% recall at 80% precision. Large-scale A/B tests show a 6-8% reduction in user views of unwanted livestreams}. These results demonstrate a scalable and adaptable approach to multimodal content governance, capable of addressing both explicit violations and emerging adversarial behaviors.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。