为小样本重抽样法提供理论保障,解决量化估计的精度问题
CLT and Edgeworth Expansion for m-out-of-n Bootstrap Estimators of The Studentized Median
- 基于m-out-of-n重抽样构造数据驱动的中位数估计器
- 在弱矩条件下证明中心极限定理,且误差率可精确量化
- 适用于马尔可夫链、强化学习等现代机器学习任务
m-out-of-n重抽样法通过从大小为n的原始样本中无放回地抽取大小为m的子样本(m远小于n)来近似统计量的分布。该方法现被广泛用于重尾数据的稳健推断、带宽选择及大样本应用。尽管其在计量经济学、生物统计学和机器学习中应用广泛,但针对样本分位数估计时的参数无关理论保证长期缺失。本文通过分析对大小为n的数据集进行m-out-of-n重抽样的分位数估计器,首次建立了此类保证:在较弱矩条件下,证明了完全数据驱动估计器的中心极限定理,且不涉及未知杂散参数;并通过反例表明该矩条件本质上不可放宽。在稍强假设下,推导出埃德伍尔德展开式,给出精确收敛速率,并导出贝里-埃塞恩界以刻画重抽样误差。最后,通过推导随机游走梅特罗波利斯-哈斯廷斯算法和遍历马尔可夫决策过程奖励的参数无关渐近分布,展示了理论在现代估计与学习任务中的实用性。
原文摘要 · Abstract (English)
The m-out-of-n bootstrap, originally proposed by Bickel, Gotze, and Zwet (1992), approximates the distribution of a statistic by repeatedly drawing m subsamples (with m much smaller than n) without replacement from an original sample of size n. It is now routinely used for robust inference with heavy-tailed data, bandwidth selection, and other large-sample applications. Despite its broad applicability across econometrics, biostatistics, and machine learning, rigorous parameter-free guarantees for the soundness of the m-out-of-n bootstrap when estimating sample quantiles have remained elusive. This paper establishes such guarantees by analyzing the estimator of sample quantiles obtained from m-out-of-n resampling of a dataset of size n. We first prove a central limit theorem for a fully data-driven version of the estimator that holds under a mild moment condition and involves no unknown nuisance parameters. We then show that the moment assumption is essentially tight by constructing a counter-example in which the CLT fails. Strengthening the assumptions slightly, we derive an Edgeworth expansion that provides exact convergence rates and, as a corollary, a Berry Esseen bound on the bootstrap approximation error. Finally, we illustrate the scope of our results by deriving parameter-free asymptotic distributions for practical statistics, including the quantiles for random walk Metropolis-Hastings and the rewards of ergodic Markov decision processes, thereby demonstrating the usefulness of our theory in modern estimation and learning tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。