用事后分层降低广告收益实验的方差,提升小流量下的测试稳定性
Variance Reduction for Heavy-Tailed Monetization Metrics in Ranking Experiments via Post-Stratification

- 结合事后分层与CUPED,利用预实验特征降低方差
- 在真实场景中实现相同统计效力所需流量减少45%
- 适合小流量下需稳定评估收益的推荐系统实验
在线排名与检索系统的评估常依赖应用收入或创作者收益等变现指标,这些指标通常呈重尾分布,少数用户主导均值和方差,导致小流量下A/B实验统计功效低、结论不可靠。本文提出一种实用框架,通过将事后分层与CUPED结合,利用实验前协变量提升变现实验的敏感性,无需额外流量。该方法已在ShareChat部署,显著降低方差,提升决策稳定性,在等效统计置信度下可节省约45%流量。文章还讨论了实际设计选择、防护机制与局限性,为信息检索与推荐系统中的真实场景提供适用建议。
原文摘要 · Abstract (English)
Online evaluation of ranking and retrieval systems often relies on downstream monetization metrics such as app revenue or creator earnings. These metrics are typically heavy-tailed, with a small fraction of users dominating both mean and variance, leading to low statistical power and unreliable conclusions in A/B experiments -- especially under limited traffic. We present a practical framework for variance reduction in online experiments by combining post-stratification with CUPED. Our approach leverages pre-experiment covariates to improve the sensitivity of monetization experiments without requiring additional traffic. Deployed at ShareChat across ranking-driven monetization experiments, the method substantially reduces variance and improves decision stability, achieving equivalent statistical confidence with ~45\% less traffic than standard metrics. We further discuss practical design choices, guardrails, and limitations, providing guidance on when post-stratification is appropriate for real-world information retrieval and Recommendation systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。