用生成模型+价值引导探索,让广告自动出价更智能稳定。
Generative Auto-Bidding with Value-Guided Explorations
- 基于生成模型和回报预测机制,动态优化出价策略。
- 在两个离线数据集和真实投放中均超越现有方法。
- 适合追求高效、自适应广告投放的平台与研究者。
自动出价凭借在动态竞争环境中优化出价决策的强大能力,已成为广告平台的关键策略。现有方法多采用规则或强化学习(RL),但规则缺乏对时变市场条件的适应性,而基于RL的方法难以捕捉马尔可夫决策过程(MDP)中的历史依赖关系与观测信息。此外,这些方法在不同广告目标下的策略适应性也面临挑战。随着离线训练被广泛用于部署稳定在线策略,固定离线数据集带来的行为模式固化与行为崩溃问题日益严重。为此,本文提出一种新的离线生成式自动出价框架GAVE(Generative Auto-bidding with Value-Guided Explorations)。GAVE通过基于分数的回报预测模块(RTG)支持多种广告目标。同时,引入基于RTG评估的动作探索机制,在探索新动作的同时保持策略稳定性。设计可学习的价值函数以引导探索方向,缓解分布外(OOD)问题。在两个离线数据集及真实世界部署中的实验表明,GAVE在离线评估和线上A/B测试中均优于最先进基线。应用该框架核心方法,团队在NeurIPS 2024‘AIGB Track:使用生成模型学习自动出价代理’竞赛中荣获第一名。
原文摘要 · Abstract (English)
Auto-bidding, with its strong capability to optimize bidding decisions within dynamic and competitive online environments, has become a pivotal strategy for advertising platforms. Existing approaches typically employ rule-based strategies or Reinforcement Learning (RL) techniques. However, rule-based strategies lack the flexibility to adapt to time-varying market conditions, and RL-based methods struggle to capture essential historical dependencies and observations within Markov Decision Process (MDP) frameworks. Furthermore, these approaches often face challenges in ensuring strategy adaptability across diverse advertising objectives. Additionally, as offline training methods are increasingly adopted to facilitate the deployment and maintenance of stable online strategies, the issues of documented behavioral patterns and behavioral collapse resulting from training on fixed offline datasets become increasingly significant. To address these limitations, this paper introduces a novel offline Generative Auto-bidding framework with Value-Guided Explorations (GAVE). GAVE accommodates various advertising objectives through a score-based Return-To-Go (RTG) module. Moreover, GAVE integrates an action exploration mechanism with an RTG-based evaluation method to explore novel actions while ensuring stability-preserving updates. A learnable value function is also designed to guide the direction of action exploration and mitigate Out-of-Distribution (OOD) problems. Experimental results on two offline datasets and real-world deployments demonstrate that GAVE outperforms state-of-the-art baselines in both offline evaluations and online A/B tests. By applying the core methods of this framework, we proudly secured first place in the NeurIPS 2024 competition, 'AIGB Track: Learning Auto-Bidding Agents with Generative Models'.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。