提出保守评估小改进的配对Bootstrap方法,防止误判噪声为有效提升。
When +1% Is Not Enough: A Paired Bootstrap Protocol for Evaluating Small Improvements
- 基于配对多种子运行与偏差校正的自助法置信区间
- 仅用3个随机种子即能避免误报0.6-2.0%的虚假提升
- 适合资源有限下谨慎评估微小性能增益的研究者
近期机器学习论文常报告单次运行在基准上1-2个百分点的提升。这些结果对随机种子、数据顺序和实现细节高度敏感,却极少附带不确定性估计或显著性检验。因此难以判断+1-2%是真实算法进步还是噪声。我们针对实际算力预算下仅能进行少量运行的情况重新审视此问题,提出一种简单且可在个人电脑上运行的评估协议:基于配对多种子运行,采用偏差校正加速(BCa)自助法置信区间,并对每种子的差异进行符号翻转置换检验。该协议刻意保守,旨在作为过度宣称的防护屏障。我们在CIFAR-10、CIFAR-10N和AG News上通过合成无提升、小提升和中等提升场景进行了验证。单次运行和非配对t检验常错误地在0.6-2.0%提升上声称显著性,尤其在文本任务上。而使用仅三个种子的配对协议在这些设置下从未宣布显著性。我们认为,在严格预算下,这种保守评估应作为小提升的标准默认方案。
原文摘要 · Abstract (English)
Recent machine learning papers often report 1-2 percentage point improvements from a single run on a benchmark. These gains are highly sensitive to random seeds, data ordering, and implementation details, yet are rarely accompanied by uncertainty estimates or significance tests. It is therefore unclear when a reported +1-2% reflects a real algorithmic advance versus noise. We revisit this problem under realistic compute budgets, where only a few runs are affordable. We propose a simple, PC-friendly evaluation protocol based on paired multi-seed runs, bias-corrected and accelerated (BCa) bootstrap confidence intervals, and a sign-flip permutation test on per-seed deltas. The protocol is intentionally conservative and is meant as a guardrail against over-claiming. We instantiate it on CIFAR-10, CIFAR-10N, and AG News using synthetic no-improvement, small-gain, and medium-gain scenarios. Single runs and unpaired t-tests often suggest significant gains for 0.6-2.0 point improvements, especially on text. With only three seeds, our paired protocol never declares significance in these settings. We argue that such conservative evaluation is a safer default for small gains under tight budgets.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。