arXiv:2601.02754cs.LGcs.AI2026-01KDD被引 1

用价值函数正则化改进广告自动出价,提升策略性能。

Q-Regularized Generative Auto-Bidding: From Suboptimal Trajectories to Optimal Policies

  • 在决策Transformer中引入双Q值正则化,联合优化模仿学习与动作价值最大化。
  • 在真实场景中实现3.27%广告GMV增长和2.49%广告ROI提升。
  • 适合需要高精度自动出价的电商广告系统研发人员参考。

随着电子商务快速发展,自动出价已成为在多样化广告主环境下优化广告表现的关键工具。现有方法多依赖强化学习和生成模型,通过复杂结构模仿离线历史行为,但需大量超参数调优,且次优轨迹加剧了策略学习难度。为此,本文提出QGA:一种基于Q值正则化的生成式自动出价方法。QGA将双Q学习策略引入决策Transformer骨干网络,实现策略模仿与动作价值最大化的联合优化,使学习到的出价策略既能利用数据集经验,又能缓解次优轨迹的负面影响。为进一步安全探索数据分布外的策略空间,提出基于Q值引导的双重探索机制,其中决策变压器以多个回溯收益目标和局部扰动动作作为条件,整个探索过程由前述Q值模块动态指导,为每个候选动作提供合理评估。在公开基准和仿真环境上的实验表明,QGA始终优于或媲美现有方法。值得注意的是,在大规模真实世界A/B测试中,QGA实现了3.27%的广告GMV提升和2.49%的广告ROI改善。

原文摘要 · Abstract (English)

With the rapid development of e-commerce, auto-bidding has become a key asset in optimizing advertising performance under diverse advertiser environments. The current approaches focus on reinforcement learning (RL) and generative models. These efforts imitate offline historical behaviors by utilizing a complex structure with expensive hyperparameter tuning. The suboptimal trajectories further exacerbate the difficulty of policy learning. To address these challenges, we proposes QGA, a novel Q-value regularized Generative Auto-bidding method. In QGA, we propose to plug a Q-value regularization with double Q-learning strategy into the Decision Transformer backbone. This design enables joint optimization of policy imitation and action-value maximization, allowing the learned bidding policy to both leverage experience from the dataset and alleviate the adverse impact of the suboptimal trajectories. Furthermore, to safely explore the policy space beyond the data distribution, we propose a Q-value guided dual-exploration mechanism, in which the DT model is conditioned on multiple return-to-go targets and locally perturbed actions. This entire exploration process is dynamically guided by the aforementioned Q-value module, which provides principled evaluation for each candidate action. Experiments on public benchmarks and simulation environments demonstrate that QGA consistently achieves superior or highly competitive results compared to existing alternatives. Notably, in large-scale real-world A/B testing, QGA achieves a 3.27% increase in Ad GMV and a 2.49% improvement in Ad ROI.

自动出价生成模型强化学习广告投放

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。