arXiv:2604.26169cs.LGecon.EM2026-04

在线学习用户响应,预算有限时仍能高效分配广告资源。

Budget-Constrained Causal Bandits: Bridging Uplift Modeling and Sequential Decision-Making

  • 结合因果推断与在线决策,动态学习个体响应并控制预算
  • 在7500条数据以下时,传统方法失效,BCCB仍可正常运行
  • 相比离线方法,波动更小,性能全面领先,适合冷启动场景

在数字广告中,预算受限下的投放策略是核心挑战。传统方法依赖历史数据训练离线提升模型,再进行约束优化分配预算,但在缺乏历史数据的冷启动场景下失效。本文提出预算约束因果老虎机(BCCB),一种在线框架,能在学习用户响应的同时合理使用预算。BCCB融合三要素:个体处理效应估计、响应不确定用户的探索、预算随时间调度。我们推导出每轮决策规则为预算约束因果分配目标的拉格朗日松弛KKT条件,提供理论基础。在Criteo Uplift数据集上,20次随机种子实验配对统计检验显示:当历史样本量低于7500时,离线方法要么失败,要么结果不可靠,而BCCB从第一个用户即可运作;此时,其运行方差仅为离线方法的1/4~1/2,且在所有预算水平下显著优于四种在线基线(Thompson Sampling, budgeted Thompson Sampling, HTE Greedy, Uplifting Bandits),p < 0.001。该结果为实践者提供了选择离线或在线范式的明确依据。

原文摘要 · Abstract (English)

Treatment allocation under budget constraints is a central challenge in digital advertising. The standard approach trains an offline uplift model on historical data, then solves a constrained optimization to allocate budget. This fails in cold-start settings where little historical data exists. We propose Budget-Constrained Causal Bandits (BCCB), an online framework that learns which users respond to ads while simultaneously spending the budget. BCCB unifies three components: learning individual-level treatment effects, exploring users whose response is uncertain, and pacing the budget over time. We derive the per-arrival decision rule as the KKT condition of a Lagrangian relaxation of the budgeted causal-allocation objective, providing a principled foundation for the algorithm. We evaluate on the Criteo Uplift dataset using 20 random seeds with paired statistical tests. Our central finding is a data-efficiency crossover at n = 7,500 historical observations (paired one-sided t-test, p = 0.043): below this threshold, offline pipelines either fail or produce unreliable allocations, while BCCB operates from the first user. BCCB exhibits 2-4x lower run-to-run variance than offline methods and outperforms all four online baselines (Thompson Sampling, budgeted Thompson Sampling, HTE Greedy, and Uplifting Bandits) at every budget level tested (p < 0.001). These results give practitioners a concrete decision rule for choosing between offline and online paradigms.

因果推断在线学习广告投放预算约束

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。