arXiv:2604.14305stat.MEcs.LG2026-04

融合贝叶斯与频率推断,实现基因检测中拷贝数变异的临床级性能保障

Combining Bayesian and Frequentist Inference for Laboratory-Specific Performance Guarantees in Copy Number Variation Detection

论文配图:Combining Bayesian and Frequentist Inference for Laboratory-Specific Performance Guarantees in Copy Number Variation Detection
图 1 · 摘自论文原文
  • 用贝叶斯后验函数在验证样本上评估,结合伽马分布建模误差
  • 在小引物数面板上实现单个基因平均绝对覆盖误差低于10%
  • 适用于真实临床场景,对过程不匹配数据仍有稳定表现

靶向扩增子面板在肿瘤诊断中广泛应用,但因扩增伪影、流程异质性及验证样本量有限,实现每个基因的拷贝数变异(CNV)检测性能保证仍具挑战。尽管贝叶斯CNV callers能量化单样本不确定性,但将其转化为临床验证所需的频率学统计量——覆盖率、假阳性上限、最小可检测变化——是根本不同的推断问题。我们实证发现,即使使用稳健的贝叶斯可信区间(包括粗化后验和沙包调整区间),在每基因引物数较少的面板上仍严重校准偏差。为此,我们提出一种混合框架:评估贝叶斯后验函数在验证样本上的平方损失,并用伽马分布建模,生成具有有效频率学覆盖的容差区间。三个关键设计使方法适应现实约束:(1) 插补策略消除真阳性样本影响,无需已知真实标签;(2) 正则化缓解小样本波动;(3) 基于对数模型证据的证据驱动分层,应对流程不匹配带来的非交换噪声。在两个靶向扩增子面板上采用留一法交叉验证,该方法在流程匹配与不匹配条件下,所有基因的平均绝对覆盖误差均低于10%,而贝叶斯对比方法在如ERBB2等临床相关基因上平均绝对误差超过60%。

原文摘要 · Abstract (English)

Targeted amplicon panels are widely used in oncology diagnostics, but providing per-gene performance guarantees for copy number variant (CNV) detection remains challenging due to amplification artifacts, process-mismatch heterogeneity, and limited validation sample sizes. While Bayesian CNV callers naturally quantify per-sample uncertainty, translating this into the frequentist population-level guarantees required for clinical validation, coverage rates, false-positive bounds, and minimum detectable copy-number changes, is a fundamentally different inferential problem. We show empirically that even robust Bayesian credible intervals, including coarsened posteriors and sandwich-adjusted intervals, are severely miscalibrated on panels with small amplicon counts per gene. To address this, we propose a hybrid framework that evaluates Bayesian posterior functionals on validation samples and models the resulting squared losses with a Gamma distribution, yielding tolerance intervals with valid frequentist coverage. Three components make the method practical under real-world constraints: (1) imputation that removes the influence of true CNV-positive samples without requiring known ground truth, (2) regularization to address small sample variability, and (3) evidence-based stratification on the log model evidence to accommodate non-exchangeable noise profiles arising from process mismatch. Evaluated on two targeted amplicon panels using leave-one-out cross-validation, the proposed method achieves single-digit mean absolute coverage error across all genes under both process-matched and unmatched conditions, whereas Bayesian comparators exhibit mean absolute errors exceeding 60\% on clinically relevant genes such as ERBB2.

CNV检测贝叶斯推断临床验证基因面板

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。