证明了基于二阶矩匹配的敲除法可实现渐近错误发现率控制。
Asymptotic FDR Control with Model-X Knockoffs: Is Moments Matching Sufficient?
- 用三类统计量条件替代原始交换性假设,统一分析近似敲除法鲁棒性。
- 首次理论证明:基于前两阶矩的高斯敲除法在特定统计量下可保证渐近FDR控制。
- 适用于需要可靠特征选择的统计学习场景,尤其适合高维数据建模。
我们提出一个统一的理论框架,研究模型-X敲除法在实际应用中的鲁棒性,重点考察近似敲除法的渐近错误发现率(FDR)控制能力。该方法将真实协变量分布替换为用户指定的、可通过样本内观测学习得到的分布。通过用三个关于近似敲除统计量的条件取代模型-X敲除变量的分布交换性条件,我们证明近似敲除法可实现渐近FDR控制。利用此统一框架,进一步证明最常用的敲除变量生成方法——基于前两阶矩匹配的高斯敲除生成器,在使用基于两阶矩的敲除统计量进行推断时,同样能实现渐近FDR控制。这是文献中首次对高斯敲除生成器的有效性和鲁棒性提供正式理论支持。模拟和真实数据实验验证了理论结果。
原文摘要 · Abstract (English)
We propose a unified theoretical framework for studying the robustness of the model-X knockoffs framework by investigating the asymptotic false discovery rate (FDR) control of the practically implemented approximate knockoffs procedure. This procedure deviates from the model-X knockoffs framework by substituting the true covariate distribution with a user-specified distribution that can be learned using in-sample observations. By replacing the distributional exchangeability condition of the model-X knockoff variables with three conditions on the approximate knockoff statistics, we establish that the approximate knockoffs procedure achieves the asymptotic FDR control. Using our unified framework, we further prove that an arguably most popularly used knockoff variable generation method--the Gaussian knockoffs generator based on the first two moments matching--achieves the asymptotic FDR control when the two-moment-based knockoff statistics are employed in the knockoffs inference procedure. For the first time in the literature, our theoretical results justify formally the effectiveness and robustness of the Gaussian knockoffs generator. Simulation and real data examples are conducted to validate the theoretical findings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。