arXiv:2504.14904cs.SIcs.AI2025-04KDD被引 24

用大模型构建动态内容审核框架,降低误报率并提升用户留存。

VLM as Policy: Common-Law Content Moderation Framework for Short Video Platform

  • 基于大视觉语言模型与思维链推理,从少量用户反馈中学习视频毒性
  • 线上实验显示报告率降20%,日活和使用时长双增
  • 开源首个短视频平台真实反馈基准,适合平台安全团队参考

快速增长的短视频平台面临严重的内容监管挑战,尤其对未成年人心理健康有害内容的传播可能引发重大社会后果。现有方法存在三大缺陷:(1) 人工审核易受偏见影响且成本高;(2) 自动化方法缺乏语义理解,准确率低;(3) 工业级规则更新滞后,难以适应快速变化趋势。本文构建了首个包含真实用户与审核员反馈的短视频内容审核基准,验证了上述问题。提出名为KuaiMod的类法律内容审核框架,包含训练数据构建、离线适配与在线部署优化三部分。该框架利用大视觉语言模型(VLM)和思维链(CoT)推理,基于稀疏用户反馈精准建模视频毒性,实现快速迭代与高精度动态策略。离线评估与大规模线上A/B测试表明,KuaiMod在基准上表现最优,部署后用户举报率下降20%,在快手多个场景中显著提升日活跃用户(DAU)与应用使用时长(AUT)。相关基准已开源:https://kuaimod.github.io。

原文摘要 · Abstract (English)

Exponentially growing short video platforms (SVPs) face significant challenges in moderating content detrimental to users' mental health, particularly for minors. The dissemination of such content on SVPs can lead to catastrophic societal consequences. Although substantial efforts have been dedicated to moderating such content, existing methods suffer from critical limitations: (1) Manual review is prone to human bias and incurs high operational costs. (2) Automated methods, though efficient, lack nuanced content understanding, resulting in lower accuracy. (3) Industrial moderation regulations struggle to adapt to rapidly evolving trends due to long update cycles. In this paper, we annotate the first SVP content moderation benchmark with authentic user/reviewer feedback to fill the absence of benchmark in this field. Then we evaluate various methods on the benchmark to verify the existence of the aforementioned limitations. We further propose our common-law content moderation framework named KuaiMod to address these challenges. KuaiMod consists of three components: training data construction, offline adaptation, and online deployment & refinement. Leveraging large vision language model (VLM) and Chain-of-Thought (CoT) reasoning, KuaiMod adequately models video toxicity based on sparse user feedback and fosters dynamic moderation policy with rapid update speed and high accuracy. Offline experiments and large-scale online A/B test demonstrates the superiority of KuaiMod: KuaiMod achieves the best moderation performance on our benchmark. The deployment of KuaiMod reduces the user reporting rate by 20% and its application in video recommendation increases both Daily Active User (DAU) and APP Usage Time (AUT) on several Kuaishou scenarios. We have open-sourced our benchmark at https://kuaimod.github.io.

内容审核大模型短视频动态策略

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。