用强化学习提升广告视频违规定位精度,解决标注噪声与泛化难题。
RAVEN: Robust Advertisement Video Violation Temporal Grounding via Reinforcement Reasoning
- 结合课程学习与多模态大模型,通过奖励机制逐步提升推理能力。
- 在工业数据集上实现92.3%违规类别准确率,时间区间定位误差低于0.8秒。
- 适合需要高精度视频内容审核的平台方和合规团队使用。
广告视频违规检测对保障平台合规至关重要,但现有方法在精准时间定位、标注噪声和泛化能力方面存在挑战。本文提出RAVEN框架,融合课程强化学习与多模态大语言模型(MLLMs),增强违规检测的推理与认知能力。RAVEN采用渐进式训练策略,结合精细与粗略标注数据,并利用组相对策略优化(GRPO)在无显式推理标注情况下发展出涌现推理能力。多层次复杂奖励机制确保精确的时间定位与一致的类别预测。在工业数据集和公开基准上的实验表明,RAVEN在违规类别准确率和时间区间定位性能上均显著领先。我们还设计了部署管道,线上A/B测试验证其实际应用效果,精度与召回率均有显著提升。RAVEN还展现出强泛化能力,有效缓解监督微调带来的灾难性遗忘问题。
原文摘要 · Abstract (English)
Advertisement (Ad) video violation detection is critical for ensuring platform compliance, but existing methods struggle with precise temporal grounding, noisy annotations, and limited generalization. We propose RAVEN, a novel framework that integrates curriculum reinforcement learning with multimodal large language models (MLLMs) to enhance reasoning and cognitive capabilities for violation detection. RAVEN employs a progressive training strategy, combining precisely and coarsely annotated data, and leverages Group Relative Policy Optimization (GRPO) to develop emergent reasoning abilities without explicit reasoning annotations. Multiple hierarchical sophisticated reward mechanism ensures precise temporal grounding and consistent category prediction. Experiments on industrial datasets and public benchmarks show that RAVEN achieves superior performances in violation category accuracy and temporal interval localization. We also design a pipeline to deploy the RAVEN on the online Ad services, and online A/B testing further validates its practical applicability, with significant improvements in precision and recall. RAVEN also demonstrates strong generalization, mitigating the catastrophic forgetting issue associated with supervised fine-tuning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。