HALO通过回溯学习提升广告竞价自适应能力,解决预算与收益目标差异大的难题。
HALO: Hindsight-Augmented Learning for Online Auto-Bidding
- 引入回溯机制,将所有探索过程转为可复用训练数据
- 采用B样条函数实现连续可导的出价映射,支持任意预算-收益组合
- 适合大规模广告平台在复杂约束下实现高鲁棒性自动竞价
数字广告平台通过实时竞价(RTB)系统在毫秒级内完成广告位拍卖,广告主通过算法出价竞争曝光。该机制虽支持精准定向,但因广告主预算和投资回报率(ROI)目标跨度巨大(从个体商户到跨国品牌),带来显著的运营复杂性,对多约束竞价(MCB)构成严峻挑战。传统自动竞价方案存在两大缺陷:1)样本效率极低,特定约束下的失败探索无法为新预算-ROI组合提供可迁移知识;2)约束变化时泛化能力差,忽视约束与出价系数间的物理关联。为此,本文提出HALO:基于回溯增强学习的在线自动竞价框架。HALO引入理论严谨的回溯机制,通过轨迹重定向将所有探索转化为任意约束配置下的训练数据。同时采用B样条函数表示,实现跨约束空间的连续、可导出价映射。实验表明,HALO在工业数据集上能有效应对多尺度约束,在大幅降低约束违反率的同时提升商品交易总额(GMV)。
原文摘要 · Abstract (English)
Digital advertising platforms operate millisecond-level auctions through Real-Time Bidding (RTB) systems, where advertisers compete for ad impressions through algorithmic bids. This dynamic mechanism enables precise audience targeting but introduces profound operational complexity due to advertiser heterogeneity: budgets and ROI targets span orders of magnitude across advertisers, from individual merchants to multinational brands. This diversity creates a demanding adaptation landscape for Multi-Constraint Bidding (MCB). Traditional auto-bidding solutions fail in this environment due to two critical flaws: 1) severe sample inefficiency, where failed explorations under specific constraints yield no transferable knowledge for new budget-ROI combinations, and 2) limited generalization under constraint shifts, as they ignore physical relationships between constraints and bidding coefficients. To address this, we propose HALO: Hindsight-Augmented Learning for Online Auto-Bidding. HALO introduces a theoretically grounded hindsight mechanism that repurposes all explorations into training data for arbitrary constraint configuration via trajectory reorientation. Further, it employs B-spline functional representation, enabling continuous, derivative-aware bid mapping across constraint spaces. HALO ensures robust adaptation even when budget/ROI requirements differ drastically from training scenarios. Industrial dataset evaluations demonstrate the superiority of HALO in handling multi-scale constraints, reducing constraint violations while improving GMV.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。