arXiv:2605.11511stat.MLcs.LG2026-05

解决主动采样后推断失效问题,确保结果可信

Post-ADC Inference: Valid Inference After Active Data Collection

  • 提出后主动采样推断框架,校正采样与目标构造双重偏差
  • 在GP-UCB和TPE数据上验证,保持正确p值与置信区间
  • 无需假设黑箱函数或代理模型,适用性广

统计推断的有效性依赖于数据收集方式。当通过主动数据收集(ADC)获取的数据被用于事后推断任务时,传统推断可能失效,因为采样会自适应地偏向收集策略青睐的区域。这一问题在黑箱优化中尤为显著,如树状帕尔岑估计器(TPE)和高斯过程上限置信界(GP-UCB)等序列模型基于优化方法会集中评估有希望的区域。本文研究在数据收集后以数据相关方式构建推断目标的情况下的统计推断有效性。为在此场景下实现有效推断,我们提出后主动数据收集推断(post-ADC inference),该框架同时考虑主动采样过程及后续数据驱动目标构造所导致的偏差。方法基于选择性推断,提供校正双重偏差的有效p值与置信区间。该框架适用于广泛的ADC过程,仅需对观测噪声做假设,无需对底层黑箱函数或SMBO算法使用的代理模型做任何假设。实证结果表明,post-ADC inference在GP-UCB和TPE生成的数据上均能实现有效推断。

原文摘要 · Abstract (English)

The validity of statistical inference depends critically on how data are collected. When data gathered through active data collection (ADC) are reused for a post-hoc inferential task, conventional inference can fail because the sampling is adaptively biased toward regions favored by the collection strategy. This issue is especially pronounced in black-box optimization, where sequential model-based optimization (SMBO) methods such as the tree-structured Parzen estimator (TPE) and Gaussian process upper confidence bound (GP-UCB) preferentially concentrate evaluations in promising regions. We study statistical inference on actively collected data when the inferential target is constructed in a data-dependent manner after data collection. To enable valid inference in this setting, we propose post-ADC inference, a framework that accounts for the biases arising from both the active data collection process and the subsequent data-driven target construction. Our method builds on selective inference and provides valid $p$-values and confidence intervals that correct for both sources of bias. The framework applies to a broad class of ADC processes by imposing only assumptions on the observation noise, without requiring any assumptions on the underlying black-box function or the surrogate model used by the SMBO algorithm. Empirical results also show that post-ADC inference provides valid inference for data collected by GP-UCB and TPE.

统计推断主动学习选择性推断黑箱优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。