ARGUS通过对抗性裁判机制,让广告审核系统随政策变化自动进化。
ARGUS: Policy-Adaptive Ad Governance via Evolving Reinforcement with Adversarial Umpiring

- 采用检察官-辩护人-裁判三元对抗架构,解决旧标签与新政策冲突。
- 在工业和公开数据集上显著优于传统微调方法,仅需少量真实标注数据。
- 适合需要动态适应监管政策的平台方和合规团队使用。
在线广告治理面临监管政策非平稳性的严峻挑战,新兴指令(如教育类或审美焦虑相关限制)导致历史数据标签严重不一致与推理模糊。本文提出ARGUS,一种政策自适应治理系统,通过多智能体对抗裁判实现持续强化学习。ARGUS采用三阶段框架:(1) 政策播种,建立初始认知;(2) 对抗标签修正,利用“检察官-辩护人-裁判”架构化解陈旧标签与新指令间的矛盾;(3) 隐含知识挖掘,通过三方辩证讨论发现复杂“灰色地带”违规行为。结合RAG增强的政策知识与思维链合成作为动态奖励,使系统推理路径与演进法规同步。在工业及公开数据集上的大量实验表明,ARGUS显著优于传统微调基线,以极少真实标注数据实现卓越的政策自适应学习能力。
原文摘要 · Abstract (English)
Online advertising governance faces significant challenges due to the non-stationary nature of regulatory policies, where emerging mandates (e.g., restrictions on education or aesthetic anxiety) create severe label inconsistencies and reasoning ambiguities in historical datasets. In this paper, we propose ARGUS, a policy-adaptive governance system that enables evolving reinforcement through multi-agent adversarial umpiring. ARGUS addresses the sparsity of new policy data by employing a three-stage framework: (1) Policy Seeding for initial perception; (2) Adversarial Label Rectification, which utilizes a ``Prosecutor-Defender-Umpire'' architecture to resolve conflicts between stale labels and new mandates; and (3) Latent Knowledge Discovery, which employs a tripartite dialectical discussion to unearth sophisticated, ``gray-area'' violations. By leveraging RAG-enhanced policy knowledge and Chain-of-Thought synthesis as dynamic rewards for reinforcement learning, ARGUS synchronizes its reasoning pathways with evolving regulations. Extensive experiments on both industrial and public datasets demonstrate that ARGUS significantly outperforms traditional fine-tuning baselines, achieving superior policy-adaptive learning with minimal gold data.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。