arXiv:2601.09758q-bio.GNcs.LG2026-01被引 1

用模型证据聚类检测测序批次差异,区分技术噪声与真实生物信号。

Detecting Batch Heterogeneity via Likelihood Clustering

  • 基于贝叶斯模型证据聚类,自动识别批次异质性。
  • 在3个临床面板中准确发现批次效应,误报率可控。
  • 适用于无标签场景,适合临床基因组分析前质检。

批次效应是基因组诊断中的主要混杂因素。在基于NGS的拷贝数变异(CNV)检测中,许多算法将测试样本与参考样本的读深进行比较,假设两者处理条件一致。当该假设被破坏时(如试剂批号变化或多中心处理),参考样本不再适用,导致假阳性CNV call或掩盖真正致病变异。在下游分析前检测此类异质性对可靠临床解读至关重要。现有方法要么基于原始特征聚类,易将生物信号与技术变异混淆,要么需要已知批次标签,而此类标签常不可得。本文提出一种新方法:根据贝叶斯模型证据对样本聚类。核心思想是,证据衡量数据与模型假设的兼容性;技术伪影违反假设,降低证据值;而生物变异(包括CNV状态)在模型预期之内,产生高证据值。这一不对称性提供了区分批次效应与生物学信号的判别信号。我们将异质性检测形式化为证据空间中混合结构的似然比检验,并使用参数自助法校准以确保保守的假阳性率。在合成数据上验证了正确的Ⅰ类错误控制,在三个临床靶向测序面板(液体活检、BRCA、地中海贫血)中展示了不同批次效应机制的检测能力,并在小鼠电生理记录中验证了跨模态泛化性能。相比标准相关性和降维方法,本方法在聚类准确性上表现更优,同时满足临床应用所需的保守性要求。

原文摘要 · Abstract (English)

Batch effects represent a major confounder in genomic diagnostics. In copy number variant (CNV) detection from NGS, many algorithms compare read depth between test samples and a reference sample, assuming they are process-matched. When this assumption is violated, with causes ranging from reagent lot changes to multi-site processing, the reference becomes inappropriate, introducing false CNV calls or masking true pathogenic variants. Detecting such heterogeneity before downstream analysis is critical for reliable clinical interpretation. Existing batch effect detection methods either cluster samples based on raw features, risking conflation of biological signal with technical variation, or require known batch labels that are frequently unavailable. We introduce a method that addresses both limitations by clustering samples according to their Bayesian model evidence. The central insight is that evidence quantifies compatibility between data and model assumptions, technical artifacts violate assumptions and reduce evidence, whereas biological variation, including CNV status, is anticipated by the model and yields high evidence. This asymmetry provides a discriminative signal that separates batch effects from biology. We formalize heterogeneity detection as a likelihood ratio test for mixture structure in evidence space, using parametric bootstrap calibration to ensure conservative false positive rates. We validate our approach on synthetic data demonstrating proper Type I error control, three clinical targeted sequencing panels (liquid biopsy, BRCA, and thalassemia) exhibiting distinct batch effect mechanisms, and mouse electrophysiology recordings demonstrating cross-modality generalization. Our method achieves superior clustering accuracy compared to standard correlation-based and dimensionality-reduction approaches while maintaining the conservativeness required for clinical usage.

基因组学批次效应贝叶斯方法临床检测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。