面对广告推荐系统中的干扰不确定性,提出稳健实验设计选择方法。
Choosing Online Experiment Designs under Interference in Ads, Recommendations, and Member-Experience Systems

- 基于不确定干扰机制,从六种可实施设计中选最鲁棒的方案。
- 在Criteo、Open Bandit等数据集上表现稳定,风险值1.295至2.240。
- 适合需应对复杂干扰的在线实验设计者,尤其关注结果可靠性。
广告、推荐与用户体验系统中的在线实验常在主导干扰机制未知时规划。处理效应可能通过预算、库存、生产者曝光、图传播或时间延续性扩散,使随机化设计本身成为统计决策。本文将问题建模为对不确定暴露机制的鲁棒设计选择,在有限的六种可实施设计中,通过最小化模糊集上的最坏情况规划风险进行比较。风险综合考虑暴露偏差、分配单元方差、最小可检测效应、污染或延续性、运营成本及估计量不匹配。理论方面,提出几何感知保证:设计偏差受暴露分布与启动分布间Wasserstein距离约束,且在Lipschitz暴露响应下该惩罚为极小极大紧界。还证明了有限目录近似与鲁棒选择定理,包含超额风险控制、分离条件下的精确恢复,以及风险面平坦时的认证短名单。实证显示,同一选择器在公开数据集样本中给出不同推荐:在Criteo广告中选用户随机化,维度无关鲁棒风险1.295;在Open Bandit-bts/men中选切换实验,风险2.105;在KuaiRand中选聚类随机化,风险2.240。Open Bandit案例凸显已知但不均衡的日志支持,倾向度范围0.00006至0.594,IPS有效样本占比仅5.17%。整体贡献为基于机制鲁棒性的干扰感知实验设计框架,输出可为合理设计选择或不确定性短名单。
原文摘要 · Abstract (English)
Online experiments in ads, recommendation, and member-experience systems are often planned before the dominant interference mechanism is known. A treatment may propagate through budgets, inventory, producer exposure, graph spillovers, or temporal carryover, making the randomization design itself a statistical decision. We formulate this problem as robust design selection over uncertain exposure mechanisms. Given a finite catalog of six implementable designs, the selector compares each design by worst-case planning risk over an ambiguity set. The risk combines exposure bias, assignment-unit variance, minimum detectable effect, contamination or carryover, operational cost, and estimand mismatch. For theoretical justification, the paper develops a geometry-aware guarantee, stating that design bias is bounded by Wasserstein distance to the launch exposure distribution, and this penalty is minimax tight under Lipschitz exposure response. We also prove finite-catalog approximation and a robust selector theorem with excess-risk control, exact recovery under separation, and certified shortlists when the risk surface is flat. Empirically, the same selector gives different recommendations across samples from public datasets. It selects user-randomization on Criteo ads with dimensionless robust risk 1.295, switchbacks on Open Bandit-bts/men with risk 2.105, and cluster-randomization on KuaiRand with risk 2.240. The Open Bandit case stresses known but uneven logging support, with propensities from 0.00006 to 0.594 and a 5.17% IPS effective-sample share. Overall, the paper contributes an interference-aware experiment design framework based on mechanism-robust design decisions, where the output is either a justified design choice or an uncertainty shortlist.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。