解决电商多渠道营销预算分配难题,提升长期收益。
Multi-channel Uplift Policy Learning

- 设计快慢双模型架构,分离短期观测与长期决策。
- 在淘宝实测中同时提升付费订单量和收入。
- 适合需要跨渠道资源优化的平台运营者。
电商平台需在多个渠道间分配固定营销预算以最大化业务效益。然而,标准的预测-再优化范式因观测混杂和严重外推问题,在组合空间中失效。本文将该挑战建模为单纯形约束的增益决策问题,提出 ReAlloc 快慢因果框架:敏捷的正交教师从短期日志中提取无偏局部梯度,解释引导的学生将其提炼为长期视角的结构化边际场。该设计支持感知资源、保守决策,捕捉跨渠道替代效应。大规模仿真及淘宝平台上的真实 A/B 测试表明,ReAlloc 实现了付费订单量与收入的双重提升。
原文摘要 · Abstract (English)
E-commerce platforms must allocate fixed marketing budgets across multiple channels to maximize business utility. However, standard predict-then-optimize (PTO) paradigms fail in this compositional space due to observational confounding and severe extrapolation. We formulate this challenge as a simplex-constrained uplift decision problem and propose ReAlloc, a fast-slow causal framework. Specifically, an agile Orthogonal Teacher extracts unbiased local gradients from short-term logs, while an Explanation-Guided Student distills them into a structured marginal field over long-term horizons. This design enables support-aware, conservative decisions that capture cross-channel substitutions. Extensive simulations and large-scale online A/B tests on Taobao platform demonstrate that ReAlloc achieves simultaneous lifts in both pay order and income.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。