新方法aLTT用自适应测试减少超参调优次数,保统计有效性且更高效。
Adaptive Learn-then-Test: Statistically Valid and Efficient Hyperparameter Selection

- 基于e过程实现数据驱动的序列化假设检验,支持早期终止
- 相同性能下测试轮次减少至传统方法的几分之一
- 适合测试成本高或有安全风险的场景,如强化学习策略选择
我们提出自适应学习-测试(aLTT)方法,一种高效的超参数选择机制,可在有限样本下对人工智能模型的总体风险提供统计保证。与依赖传统p值多重假设检验(MHT)的LTT方法不同,aLTT利用e过程实现数据相关的序列化MHT,并支持早期终止。因此,aLTT能显著减少测试轮次,特别适用于测试成本高或存在安全风险的场景。在离线强化学习在线策略选择和提示工程等应用中,aLTT表现与LTT相当,但仅需极少的测试次数。
原文摘要 · Abstract (English)
We introduce adaptive learn-then-test (aLTT), an efficient hyperparameter selection procedure that provides finite-sample statistical guarantees on the population risk of AI models. Unlike the existing learn-then-test (LTT) technique, which relies on conventional p-value-based multiple hypothesis testing (MHT), aLTT implements sequential data-dependent MHT with early termination by leveraging e-processes. As a result, aLTT can reduce the number of testing rounds, making it particularly well-suited for scenarios in which testing is costly or presents safety risks. Apart from maintaining statistical validity, in applications such as online policy selection for offline reinforcement learning and prompt engineering, aLTT is shown to achieve the same performance as LTT while requiring only a fraction of the testing rounds.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。