针对平台组合实验难题,提出分两阶段的高效设计方法。
Policy-Aware Design of Large-Scale Factorial Experiments

- 将重叠实验整合为低秩张量问题,用张量补全预测未测试组合表现。
- 在1亿级淘宝数据上验证,预算低时性能显著优于传统方法。
- 适合大规模平台做组合策略优化,尤其适用于资源有限场景。
数字平台常在共享用户群体上运行大量在线实验。当产品决策具有组合性(如界面元素、流程、消息或激励的组合)时,可行干预方案数量呈组合爆炸式增长,而可用流量却有限。重叠实验可能产生交互效应,传统分散式A/B测试难以处理。本文研究在固定实验预算下,不追求估计所有处理效应,而是识别出高性能策略的大型因子实验设计方法。提出两阶段设计:第一阶段将重叠实验集中为单一因子问题,建模预期结果为低秩张量,通过采样部分组合,利用张量补全推断未测试组合表现,并基于估计的边际贡献剔除弱因素水平;第二阶段对剩余组合应用序列减半法选出最终策略。理论证明了与差距无关的简单遗憾界和与差距相关的识别保证,表明相关复杂度取决于低秩张量的自由度及因子水平间的分离结构,而非完整因子设计规模。基于1亿次淘宝商品捆绑行为构建的离线评估显示,该方法在低预算和高噪声条件下显著优于一次性张量补全和无结构最优臂基准。结果表明,集中化、策略感知的实验设计可使组合式产品设计在平台规模上实现可操作性。
原文摘要 · Abstract (English)
Digital firms routinely run many online experiments on shared user populations. When product decisions are compositional, such as combinations of interface elements, flows, messages, or incentives, the number of feasible interventions grows combinatorially, while available traffic remains limited. Overlapping experiments can therefore generate interaction effects that are poorly handled by decentralized A/B testing. We study how to design large-scale factorial experiments when the objective is not to estimate every treatment effect, but to identify a high-performing policy under a fixed experimentation budget. We propose a two-stage design that centralizes overlapping experiments into a single factorial problem and models expected outcomes as a low-rank tensor. In the first stage, the platform samples a subset of intervention combinations, uses tensor completion to infer performance on untested combinations, and eliminates weak factor levels using estimated marginal contributions. In the second stage, it applies sequential halving to the surviving combinations to select a final policy. We establish gap-independent simple-regret bounds and gap-dependent identification guarantees showing that the relevant complexity scales with the degrees of freedom of the low-rank tensor and the separation structure across factor levels, rather than the full factorial size. In an offline evaluation based on a product-bundling problem constructed from 100 million Taobao interactions, the proposed method substantially outperforms one-shot tensor completion and unstructured best-arm benchmarks, especially in low-budget and high-noise settings. These results show how centralized, policy-aware experimentation can make combinatorial product design operationally feasible at platform scale.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。