arXiv:2602.16690stat.MEcs.LG2026-02被引 1

利用合成数据提升假设检验效率,同时严格控制错误发现率。

Synthetic-Powered Multiple Testing with FDR Control

  • 通过合成数据增强真实数据,动态适应合成数据质量
  • 在有限样本下保证错误发现率控制,无需依赖合成数据有效性
  • 适用于基因组学、药物筛选等需高灵敏度的场景

多重假设检验中的错误发现率(FDR)控制是统计推断的核心问题,广泛应用于基因组学、药物筛选和异常检测。许多场景中,研究者不仅拥有真实实验数据,还可获得辅助或合成数据——来自过往相关实验或生成模型生成的数据——这些数据能为待检验假设提供额外证据。本文提出 SynthBH,一种安全利用合成数据的多重检验方法。证明了 SynthBH 在温和的正相关依赖条件下,可实现有限样本、分布无关的 FDR 控制,且不要求合并数据的 p 值在原假设下有效。该方法能自适应合成数据的未知质量:当合成数据质量高时,提升样本效率并增强检验功效;无论合成数据质量如何,均能将 FDR 控制在用户指定水平。我们在表格型异常检测基准、药物-癌症敏感性基因组分析中验证了 SynthBH 的实证性能,并通过模拟数据实验深入分析其特性。

原文摘要 · Abstract (English)

Multiple hypothesis testing with false discovery rate (FDR) control is a fundamental problem in statistical inference, with broad applications in genomics, drug screening, and outlier detection. In many such settings, researchers may have access not only to real experimental observations but also to auxiliary or synthetic data -- from past, related experiments or generated by generative models -- that can provide additional evidence about the hypotheses of interest. We introduce SynthBH, a synthetic-powered multiple testing procedure that safely leverages such synthetic data. We prove that SynthBH guarantees finite-sample, distribution-free FDR control under a mild PRDS-type positive dependence condition, without requiring the pooled-data p-values to be valid under the null. The proposed method adapts to the (unknown) quality of the synthetic data: it enhances the sample efficiency and may boost the power when synthetic data are of high quality, while controlling the FDR at a user-specified level regardless of their quality. We demonstrate the empirical performance of SynthBH on tabular outlier detection benchmarks and on genomic analyses of drug-cancer sensitivity associations, and further study its properties through controlled experiments on simulated data.

多重检验合成数据FDR控制基因组学

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。