自动出价系统在预算约束下学习未知广告价值与竞标上限,实现高效决策。
Learning to Bid with Unknown Private Values in Budget-Constrained First-Price Auctions
- 构建共享上下文线性处理效应模型,联合估计广告价值与胜标价格。
- 在完整反馈下理论误差为√T量级,二值反馈下为T^(2/3)量级。
- 适用于广告平台或需达成投资回报率目标的自动化出价场景。
我们研究在预算和支出回报率(RoS)约束下的重复第一价格拍卖中自动化出价的运营问题。在此场景中,自动出价器需将广告主目标与约束转化为实时出价,同时学习两个隐含对象:每个广告展示的因果提升价值及赢得该展示所需的最高竞标价(HoB)。我们通过共享上下文线性处理效应(LTE)结构建模提升价值与HoB,并分析全信息与二值胜败反馈两种情形。提出双感知在线学习框架Dual-LTE,通过置信度引导探索,协调价值估计、HoB估计与预算/RoS控制。证明在全信息反馈下,遗憾与约束违反的保证为˜O(√T),在二值反馈下为˜O(T^{2/3}),其中˜O(·)隐藏问题相关与对数因子。使用真实拍卖特征的半合成实验表明,Dual-LTE在不同预算和RoS设置下均低于基线遗憾,同时揭示遗憾与约束违规之间的权衡。结果为管理广告预算或追求ROAS目标的DSP及平台自动出价提供操作指导:当价值估计足够准确且面临约束压力时,应遵循拉格朗日出价规则;否则应采用受控探索。
原文摘要 · Abstract (English)
We study the operational problem of automated bidding in repeated first-price auctions under budget and return-on-spend (RoS) constraints. In this setting, an auto-bidder must translate advertiser goals and constraints into real-time bids while learning two latent objects: the causal uplift value of each ad impression and the highest competing bid (HoB) needed to win it. We model uplift values and HoBs through a shared-context Linear Treatment Effect (LTE) structure and analyze both full-information and binary HoB feedback. We develop Dual-LTE, a dual-aware online learning framework that coordinates value estimation, HoB estimation, and budget/RoS control through confidence-guided exploration. We prove regret and constraint-violation guarantees that scale as $\widetilde{O}(\sqrt{T})$ under full-information HoB feedback and $\widetilde{O}(T^{2/3})$ under binary win/loss feedback, where $\widetilde{O}(\cdot)$ hides problem-dependent and logarithmic factors. Semi-synthetic experiments using real auction covariates show that Dual-LTE achieves lower regret than the baselines across budget and RoS settings, while illustrating the tradeoff between regret and constraint violation. Our results provide operational guidance for DSPs and platform auto-bidders that manage advertiser budgets or seek to meet ROAS targets. When impression values must be learned, value estimation should be coordinated with budget or ROAS control: the auto-bidder should follow the Lagrangian bidding rule only when value estimates are sufficiently accurate given the current constraint pressure and should otherwise use controlled exploration.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。