arXiv:2511.19168cs.LGcs.CL2025-11EMNLP被引 7

用强化学习精准定位广告视频中的违规内容

RAVEN++: Pinpointing Fine-Grained Violations in Advertisement Videos with Active Reinforcement Reasoning

  • 通过动态难度自适应的强化学习提升模型训练效率
  • 在多个数据集上实现更精细的违规内容识别与解释能力
  • 适合广告审核、内容安全等需要高精度判断的场景

广告是数字经济的核心,但视频广告的审核因复杂性高而面临挑战,尤其在细粒度违规定位方面仍存不足。尽管已有如RAVEN等模型提升了粗粒度检测能力,但在细粒度理解、可解释性和泛化性方面仍有缺口。为此,我们提出RAVEN++框架,包含三大创新:1)主动强化学习(Active RL),根据样本难度动态调整训练;2)层级奖励函数与推理蒸馏,实现细粒度违规理解;3)分阶段渐进式训练,融合知识注入、基于课程的被动强化学习与主动强化学习。在公开及私有数据集上的离线实验和线上A/B测试均表明,RAVEN++在细粒度违规理解、推理能力和泛化性上优于通用大模型及专用模型RAVEN。

原文摘要 · Abstract (English)

Advertising (Ad) is a cornerstone of the digital economy, yet the moderation of video advertisements remains a significant challenge due to their complexity and the need for precise violation localization. While recent advancements, such as the RAVEN model, have improved coarse-grained violation detection, critical gaps persist in fine-grained understanding, explainability, and generalization. To address these limitations, we propose RAVEN++, a novel framework that introduces three key innovations: 1) Active Reinforcement Learning (RL), which dynamically adapts training to samples of varying difficulty; 2) Fine-Grained Violation Understanding, achieved through hierarchical reward functions and reasoning distillation; and 3) Progressive Multi-Stage Training, which systematically combines knowledge injection, curriculum-based passive RL, and active RL. Extensive experiments on both public and proprietary datasets, on both offline scenarios and online deployed A/B Testing, demonstrate that RAVEN++ outperforms general-purpose LLMs and specialized models like RAVEN in terms of fine-grained violation understanding, reasoning capabilities, and generalization ability.

广告审核强化学习细粒度检测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。