为超参数选择提供统计保证,让调参结果可信赖。
Statistically Valid Hyperparameter Selection: From Tuning to Guarantees

- 采用学完再验(LTT)框架,将调参转化为多重假设检验。
- 可证明满足特定可靠性要求,如平均风险上限、信息约束等。
- 适合对安全性与可靠性有严格要求的AI系统部署场景。
超参数选择是现代人工智能系统部署中的关键步骤,需调整推理时参数、实现级设置及决策规则的阈值等自由度。尽管实际重要,现有方法如网格搜索或贝叶斯优化通常仅凭经验,无法提供可靠性和安全性的正式统计保证。本文提出统一的统计框架,基于学完再验(LTT)范式,将问题建模为候选超参数集上的多重假设检验。该框架可选出在应用特定可靠性要求下(如平均风险、分位数风险或信息论约束)被证明成立的超参数,并对有限样本下的误差概率进行显式控制。支撑性统计工具,包括p值、e值和浓度不等式,均在附录中从基础推导。
原文摘要 · Abstract (English)
Hyperparameter selection is a critical step in the deployment of modern artificial intelligence systems, given the need to tune degrees of freedom such as inference-time parameters, implementation-level settings, and thresholds driving decision rules. Despite its practical importance, hyperparameter selection is typically performed using best-effort empirical methods such as grid search or Bayesian optimization, which provide no formal statistical guarantees on reliability or safety. This monograph presents a unified statistical framework for reliable hyperparameter selection, centered on the learn-then-test (LTT) paradigm, which formulates the problem as multiple hypothesis testing over a candidate set of hyperparameters. The framework enables the selection of hyperparameters that provably satisfy application-specific reliability requirements -- such as bounds on average risk, quantile risk, or information-theoretic constraints -- with explicit, finite-sample control of error probabilities. The supporting statistical machinery, namely p-values, e-values, and concentration inequalities, is developed from first principles in a dedicated appendix.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。