用双层学习框架融合观察与实验数据,提升大规模营销决策效果
Bi-Level Decision-Focused Causal Learning for Large-Scale Marketing Optimization: Bridging Observational and Experimental Data
- 双层优化框架同时利用观察数据和实验数据,纠正偏差
- 通过隐式微分实现无偏决策评估,显著提升优化效果
- 已在美团落地,实测优于现有最优方法
在线平台需复杂营销策略以优化用户留存与收入,传统方法分为两阶段:先用机器学习预测个体处理效应,再用运筹学优化决策。但存在两大问题:预测与决策目标不一致,导致高预测精度无法转化为好决策;观测数据含选择偏差等多重偏差,而随机对照试验虽无偏但稀缺且昂贵,造成高方差估计。本文提出双层决策导向因果学习(Bi-DFCL),首先利用实验数据构建无偏运筹学决策质量估计器,通过代理损失函数引导模型训练并连接离散优化梯度;其次建立双层优化框架,结合观测与实验数据,通过隐式微分求解,使无偏估计器修正观测数据的偏差方向,实现最优偏差-方差权衡。在公开基准、工业营销数据集及大规模线上A/B测试中验证有效,显著优于现有最优方法。目前已在美团部署,服务于全球最大的在线外卖平台之一。
原文摘要 · Abstract (English)
Online Internet platforms require sophisticated marketing strategies to optimize user retention and platform revenue -- a classical resource allocation problem. Traditional solutions adopt a two-stage pipeline: machine learning (ML) for predicting individual treatment effects to marketing actions, followed by operations research (OR) optimization for decision-making. This paradigm presents two fundamental technical challenges. First, the prediction-decision misalignment: Conventional ML methods focus solely on prediction accuracy without considering downstream optimization objectives, leading to improved predictive metrics that fail to translate to better decisions. Second, the bias-variance dilemma: Observational data suffers from multiple biases (e.g., selection bias, position bias), while experimental data (e.g., randomized controlled trials), though unbiased, is typically scarce and costly -- resulting in high-variance estimates. We propose Bi-level Decision-Focused Causal Learning (Bi-DFCL) that systematically addresses these challenges. First, we develop an unbiased estimator of OR decision quality using experimental data, which guides ML model training through surrogate loss functions that bridge discrete optimization gradients. Second, we establish a bi-level optimization framework that jointly leverages observational and experimental data, solved via implicit differentiation. This novel formulation enables our unbiased OR estimator to correct learning directions from biased observational data, achieving optimal bias-variance tradeoff. Extensive evaluations on public benchmarks, industrial marketing datasets, and large-scale online A/B tests demonstrate the effectiveness of Bi-DFCL, showing statistically significant improvements over state-of-the-art. Currently, Bi-DFCL has been deployed at Meituan, one of the largest online food delivery platforms in the world.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。