arXiv:2507.16186cs.LGcs.IR2025-07

用专家轨迹提升自动竞价模型数据质量与奖励稳定性。

EBaReT: Expert-guided Bag Reward Transformer for Auto Bidding

  • 引入专家轨迹和PU学习识别优质决策,改善低质量数据问题。
  • 将短期行为视为'袋',设计平滑奖励函数缓解奖励不确定性。
  • 适合需要高可靠性竞价系统的广告平台或自动化投放场景。

强化学习已广泛应用于自动竞价。传统方法将竞价建模为马尔可夫决策过程(MDP),近期研究尝试采用生成式强化学习解决竞价中的长期依赖问题。尽管有效,这些方法通常依赖监督学习,易受低质量数据影响,因次优出价过多及点击与转化率低导致奖励概率小。本文将自动竞价形式化为序列决策问题,提出新型专家引导的袋奖励变压器(EBaReT)。为解决数据质量问题,生成一组专家轨迹作为训练补充数据,并使用基于正-未标记(PU)学习的判别器识别专家转移。为确保决策达到专家水平,进一步设计专家引导推理策略。此外,为缓解奖励不确定性,将特定时间段内的转移视为一个“袋”,设计平滑奖励函数以促进更稳定地获取奖励。大量实验表明,本模型在性能上优于当前最先进的竞价方法。

原文摘要 · Abstract (English)

Reinforcement learning has been widely applied in automated bidding. Traditional approaches model bidding as a Markov Decision Process (MDP). Recently, some studies have explored using generative reinforcement learning methods to address long-term dependency issues in bidding environments. Although effective, these methods typically rely on supervised learning approaches, which are vulnerable to low data quality due to the amount of sub-optimal bids and low probability rewards resulting from the low click and conversion rates. Unfortunately, few studies have addressed these challenges. In this paper, we formalize the automated bidding as a sequence decision-making problem and propose a novel Expert-guided Bag Reward Transformer (EBaReT) to address concerns related to data quality and uncertainty rewards. Specifically, to tackle data quality issues, we generate a set of expert trajectories to serve as supplementary data in the training process and employ a Positive-Unlabeled (PU) learning-based discriminator to identify expert transitions. To ensure the decision also meets the expert level, we further design a novel expert-guided inference strategy. Moreover, to mitigate the uncertainty of rewards, we consider the transitions within a certain period as a "bag" and carefully design a reward function that leads to a smoother acquisition of rewards. Extensive experiments demonstrate that our model achieves superior performance compared to state-of-the-art bidding methods.

自动竞价强化学习专家引导奖励设计

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。