arXiv:2503.16589cs.LGcs.ET2025-03被引 1

科学设定实验重复次数,避免随机优化器评估结论错误

A Statistical Analysis for Per-Instance Evaluation of Stochastic Optimizers: Avoiding Unreliable Conclusions

  • 基于置信区间分析,推导出保证精度所需的最少实验次数
  • 提出自适应调整重复次数的算法,确保评估结果可靠
  • 适合做优化器对比、超参调优的研究者参考

随机优化器在求解同一问题时多次运行可能产生不同结果,因此需通过多轮重复实验评估其性能。但性能指标的准确性依赖于实验次数,必须使用统计方法进行分析。本文研究常见评估指标的置信区间与实验重复次数的关系,推导出达到指定精度所需的最少重复次数下界,并据此提出一种自适应调整重复次数的算法,以确保评估结果的准确性和可信度。仿真结果表明,该方法能有效支持可靠的基准测试、超参数调优,防止因样本不足而得出错误结论。

原文摘要 · Abstract (English)

A key trait of stochastic optimizers is that multiple runs of the same optimizer in attempting to solve the same problem can produce different results. As a result, their performance is evaluated over several repeats, or runs, on the problem. However, the accuracy of the estimated performance metrics depends on the number of runs and should be studied using statistical tools. We present a statistical analysis of the common metrics, and develop guidelines for experiment design to measure the optimizer's performance using these metrics to a high level of confidence and accuracy. To this end, we first discuss the confidence interval of the metrics and how they are related to the number of runs of an experiment. We then derive a lower bound on the number of repeats in order to guarantee achieving a given accuracy in the metrics. Using this bound, we propose an algorithm to adaptively adjust the number of repeats needed to ensure the accuracy of the evaluated metric. Our simulation results demonstrate the utility of our analysis and how it allows us to conduct reliable benchmarking as well as hyperparameter tuning and prevent us from drawing premature conclusions regarding the performance of stochastic optimizers.

优化器评估统计分析实验设计

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。