arXiv:2507.01066cs.IRcs.CV2025-07被引 6

用嵌入检索提升短视频审核效率,应对热点快速响应难题。

Embedding-based Retrieval in Multimodal Content Moderation

  • 基于监督对比学习训练多模态嵌入模型,支持高效内容匹配
  • 离线测试中ROC-AUC从0.85升至0.99,PR-AUC从0.35升至0.95
  • 适合需要快速适应新趋势、降低人工成本的平台审核场景

视频理解在短视频平台内容审核中至关重要,用于识别不当内容。尽管分类仍是主流方法,但在需快速响应的场景(如趋势变化和紧急升级)中常显不足。为此,我们提出嵌入式检索(EBR)方法,作为传统分类的补充。首先采用监督对比学习(SCL)框架训练一系列基础嵌入模型,涵盖单模态与多模态架构,其性能优于CLIP和MoCo等现有对比学习方法。在此基础上,构建集成嵌入生成与视频检索的系统,实现高效灵活的趋势应对。25个新兴趋势的离线实验显示,EBR将ROC-AUC从0.85提升至0.99,PR-AUC从0.35提升至0.95;线上实验表明,行动率提升10.32%,运营成本降低超80%,且可解释性与灵活性优于分类方案。

原文摘要 · Abstract (English)

Video understanding plays a fundamental role for content moderation on short video platforms, enabling the detection of inappropriate content. While classification remains the dominant approach for content moderation, it often struggles in scenarios requiring rapid and cost-efficient responses, such as trend adaptation and urgent escalations. To address this issue, we introduce an Embedding-Based Retrieval (EBR) method designed to complement traditional classification approaches. We first leverage a Supervised Contrastive Learning (SCL) framework to train a suite of foundation embedding models, including both single-modal and multi-modal architectures. Our models demonstrate superior performance over established contrastive learning methods such as CLIP and MoCo. Building on these embedding models, we design and implement the embedding-based retrieval system that integrates embedding generation and video retrieval to enable efficient and effective trend handling. Comprehensive offline experiments on 25 diverse emerging trends show that EBR improves ROC-AUC from 0.85 to 0.99 and PR-AUC from 0.35 to 0.95. Further online experiments reveal that EBR increases action rates by 10.32% and reduces operational costs by over 80%, while also enhancing interpretability and flexibility compared to classification-based solutions.

内容审核嵌入检索多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。