将数据增强融入置信区间构造,统一了多种统计方法。
Data augmented bootstrap: Unifying confidence interval construction by approximate invariance

- 基于数据的近似不变性设计新框架DAB,不依赖群结构。
- 理论证明覆盖概率随不变性强度介于有限样本与渐近之间。
- 首次将数据增强用于置信区间,适合机器学习与统计交叉研究者。
我们提出数据增强自助法(DAB),一种基于数据近似不变变换构建置信区间的框架。作为特例,DAB 恢复了依赖精确群对称性的流行方法,如共形预测、最大均值差异U统计量的野生自助法及近期提出的SymmPI。同时,DAB 还涵盖经典自助法,其利用数据集在样本索引均匀抽样下随规模增大产生的近似不变性。对所有DAB方法,我们建立了理论覆盖结果,根据不变性强度在有限样本与渐近保证间插值,且无需假设群结构。近似不变性以Kolmogorov距离度量,对于满足高斯普遍性的统计量,退化为条件均值与方差匹配。这使得可将广泛使用的基于近似不变性的机器学习启发式数据增强(DA)融入已有统计方法。我们在模拟设置及图像、语言和科学数据上实证测试了将DA引入自助法、野生自助法和共形预测的性能。
原文摘要 · Abstract (English)
We propose the data augmented bootstrap (DAB), a framework for constructing confidence intervals from approximately invariant transformations of the data. As special cases, DAB recovers popular methods that rely on exact group symmetries, such as conformal prediction, wild bootstrap for Maximum Mean Discrepancy U-statistics and the recently proposed SymmPI. Meanwhile, DAB also recovers the classical bootstrap method, which exploits the dataset's approximate invariance under uniform sampling of data indices as the dataset size grows. For all DAB methods, we establish theoretical coverage results that interpolate between finite-sample and asymptotic guarantees according to the strength of the invariance, and without assuming a group structure. The approximate invariance is measured in the Kolmogorov distance and, for statistics that satisfy Gaussian universality, reduces to conditional mean and variance matching. This allows us to incorporate data augmentation (DA), a widely used machine learning heuristic based on approximate invariances, into known statistical methods. We empirically test the performance of incorporating DA into bootstrap, wild bootstrap and conformal prediction for simulated settings as well as for image, language and scientific data.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。