解决视频广告合规修复中过度修改原意的问题
R^3: Advertisement Compliance Rectification via Group-Relative Experience Extractor and Curriculum Reinforcement

- 用群体相对经验提取生成高质量标注数据
- 课程式强化学习实现合规与语义一致性的平衡
- 端到端流程支持工业级视频广告修复
在线广告内容审核至关重要,但每日数百万条内容被拒,人工修复已不可行,尤其针对视频广告。现有安全导向方法常导致过度编辑,损害广告原始语义意图。本文聚焦视频广告文本违规(含语音转录与屏幕文字)的修复问题,提出R^3框架,在确保合规的同时最大限度保留原始语义。该框架包含三项创新:(1) 基于经验驱动的数据合成机制,通过群体相对合规经验提取器生成高质量监督信号;(2) 分层奖励机制的课程式强化学习策略,兼顾合规性与语义一致性;(3) 融合文本识别、重写与重渲染的全流程视频修复系统,支持工业部署。在工业数据集及线上A/B测试中,R^3显著优于现有基线,实现违规修复与意图保留的最佳平衡。
原文摘要 · Abstract (English)
Rigorous content moderation is crucial for online advertising but leads to millions of daily rejections. This scale renders manual rectification infeasible, particularly for video advertisements. However, existing safety-driven methods often suffer from aggressive over-editing, which compromises the advertiser's original semantic intent merely to satisfy compliance. In this work, we target the rectification of textual violations in video ads, covering both speech transcripts and on-screen text. We propose R^3, a novel framework designed to harmonize compliance with original semantic intent preservation. Our approach integrates three key innovations: (1) an experience-driven data synthesis framework that bootstraps high-quality supervision via a group-Relative compliance experience extractor; (2) a curriculum Reinforcement learning strategy with hierarchical rewards designed to enforce compliance while maximizing semantic consistency; and (3) a comprehensive video Rectification framework seamlessly integrating text recognition, rewriting, and re-rendering for industrial deployment. Extensive experiments on industrial datasets and online A/B testing demonstrate that R^3 significantly outperforms state-of-the-art baselines, achieving an optimal trade-off between violation rectification and intent preservation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。