arXiv:2602.05798stat.MEcs.LG2026-02中稿 · IEEE International…

用神经网络替代传统方法,更准估算假阳性率,提升发现能力。

Learning False Discovery Rate Control via Model-Based Neural Networks

  • 用神经网络学习假发现比例,替代原有解析估计器
  • 在模拟和基因组研究中显著提高真实变量检出率
  • 适合需要高精度错误控制的高维变量选择场景

高维变量选择中的错误发现率(FDR)控制需在严格误差控制与统计功效之间取得平衡。现有具备理论保证的方法往往过于保守,导致实际假发现比例(FDP)与目标FDR水平之间存在持续差距。本文提出一种基于模型的神经网络增强方法,改进T-Rex Selector框架:将原有的解析式FDP估计器替换为仅在多样化合成数据集上训练的神经网络,实现对FDP更紧密、更精确的逼近。该优化使方法能更贴近目标FDR水平运行,从而在保持有效近似控制的前提下显著提升发现能力。通过大量模拟实验及一项具有挑战性的合成全基因组关联研究(GWAS),结果表明,本方法在识别真实变量方面优于现有方法。

原文摘要 · Abstract (English)

Controlling the false discovery rate (FDR) in high-dimensional variable selection requires balancing rigorous error control with statistical power. Existing methods with provable guarantees are often overly conservative, creating a persistent gap between the realized false discovery proportion (FDP) and the target FDR level. We introduce a learning-augmented enhancement of the T-Rex Selector framework that narrows this gap. Our approach replaces the analytical FDP estimator with a neural network trained solely on diverse synthetic datasets, enabling a substantially tighter and more accurate approximation of the FDP. This refinement allows the procedure to operate much closer to the desired FDR level, thereby increasing discovery power while maintaining effective approximate control. Through extensive simulations and a challenging synthetic genome-wide association study (GWAS), we demonstrate that our method achieves superior detection of true variables compared to existing approaches.

变量选择假发现率神经网络

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。