用生成式规划和策略优化提升广告自动出价效果
Enhancing Generative Auto-bidding with Offline Reward Evaluation and Policy Search
- 构建轨迹评估器与约束优化方案,实现离线数据外的安全探索
- 在模拟和真实广告系统上均达到当前最佳性能
- 适合研究广告自动化、强化学习应用的从业者
自动出价是广告商提升投放效果的关键工具。近期研究表明,基于离线数据学习条件生成规划器的AI生成出价(AIGB)方法,相较传统离线强化学习方法表现更优。然而,现有AIGB方法因无法在静态数据集外进行有效探索,仍存在性能瓶颈。为此,本文提出AIGB-Pearl(Planning with Evaluator via RL),融合生成式规划与策略优化。其核心在于构建轨迹评估器以判断生成得分质量,并设计可证明安全的KL-Lipschitz约束得分最大化方案,确保在离线数据外的安全高效探索。进一步开发了结合同步耦合技术的实际算法,保障模型满足该方案所需的正则性要求。在模拟及真实广告系统上的大量实验表明,本方法性能达到当前最优水平。
原文摘要 · Abstract (English)
Auto-bidding is a critical tool for advertisers to improve advertising performance. Recent progress has demonstrated that AI-Generated Bidding (AIGB), which learns a conditional generative planner from offline data, achieves superior performance compared to typical offline reinforcement learning (RL)-based auto-bidding methods. However, existing AIGB methods still face a performance bottleneck due to their inherent inability to explore beyond the static dataset with feedback. To address this, we propose \textbf{AIGB-Pearl} (\emph{\textbf{P}lanning with \textbf{E}valu\textbf{A}tor via \textbf{RL}}), a novel method that integrates generative planning and policy optimization. The core of AIGB-Pearl lies in constructing a trajectory evaluator to assess the quality of generated scores and designing a provably sound KL-Lipschitz-constrained score-maximization scheme to ensure safe and efficient exploration beyond the offline dataset. A practical algorithm that incorporates the synchronous coupling technique is further developed to ensure the model regularity required by the proposed scheme. Extensive experiments on both simulated and real-world advertising systems demonstrate the state-of-the-art performance of our approach.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。