arXiv:2510.23666stat.MLcs.LG2025-10

提出修正方法,让非正态分布数据下的A/B测试更可靠

Beyond Normality: Reliable A/B Testing with Non-Gaussian Data

  • 基于埃奇沃斯展开修正p值,提升小样本下检验可靠性
  • 发现多数线上指标需数亿样本才能保证t检验有效
  • 适合大规模平台做产品优化时的实验设计参考

A/B测试是在线市场决策的核心,常通过配对t检验比较处理组与对照组结果。然而在实际应用中,当数据分布偏离正态性或两组样本量不等时,传统t检验的Ⅰ类错误率失控,导致误判风险上升。本文量化了偏度、长尾分布及不均衡分配对误差率的影响,推导出t检验有效的最小样本量公式。研究发现,许多线上反馈指标需数亿样本才能保证可靠性。为此,提出基于埃奇沃斯展开的校正方法,在样本有限时提供更准确的p值。在主流A/B测试平台的离线实验验证了理论阈值的有效性,表明该方法显著提升了真实场景下的测试可靠性。

原文摘要 · Abstract (English)

A/B testing has become the cornerstone of decision-making in online markets, guiding how platforms launch new features, optimize pricing strategies, and improve user experience. In practice, we typically employ the pairwise $t$-test to compare outcomes between the treatment and control groups, thereby assessing the effectiveness of a given strategy. To be trustworthy, these experiments must keep Type I error (i.e., false positive rate) under control; otherwise, we may launch harmful strategies. However, in real-world applications, we find that A/B testing often fails to deliver reliable results. When the data distribution departs from normality or when the treatment and control groups differ in sample size, the commonly used pairwise $t$-test is no longer trustworthy. In this paper, we quantify how skewed, long tailed data and unequal allocation distort error rates and derive explicit formulas for the minimum sample size required for the $t$-test to remain valid. We find that many online feedback metrics require hundreds of millions samples to ensure reliable A/B testing. Thus we introduce an Edgeworth-based correction that provides more accurate $p$-values when the available sample size is limited. Offline experiments on a leading A/B testing platform corroborate the practical value of our theoretical minimum sample size thresholds and demonstrate that the corrected method substantially improves the reliability of A/B testing in real-world conditions.

A/B测试统计检验非正态数据样本量

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。