arXiv:2608.11054cs.LG2026-08

对比三种深度学习不确定性方法,提升基因组学预测可靠性。

Uncertainty-Aware Deep Learning for Genomics Applications: Insights from an Empirical Study

论文配图:Uncertainty-Aware Deep Learning for Genomics Applications: Insights from an Empirical Study
图 1 · 摘自论文原文
  • 采用深度集成、贝叶斯神经网络和蒙特卡洛丢弃法对比分析
  • 贝叶斯网络在类别不平衡和分布外数据下表现更优
  • 不确定性分数可用于筛选高质量蛋白-RNA互作预测

深度学习已成为基因组学应用的主流计算工具,但不确定性量化(UQ)——特别是不同不确定性估计的可靠性——尚未得到系统关注。本文针对基因组学应用开展实证分析,对比了深度集成、贝叶斯神经网络和蒙特卡洛丢弃法在序列到活性模型与单细胞表达分析两类典型场景下的表现。实验考虑了基因组数据常见特征,在多种模态下评估了各方法对不确定性的建模能力。结果表明,尽管贝叶斯神经网络计算开销较大,但在应对强类别不平衡和分布外数据时更具优势。此外,不确定性得分可有效筛选出高质量的蛋白-RNA相互作用预测结果,为实际应用提供可靠判断依据。

原文摘要 · Abstract (English)

Deep learning models have emerged as the standard computational tool for a wide range of applications in genomics. Yet, uncertainty quantification (UQ) -- and more specifically, the reliability of different uncertainty estimates in this domain -- has received little systematic attention. This work presents an empirical analysis of UQ in deep learning models, focusing on genomics applications. In a series of experiments, we contrast Deep Ensembles, Bayesian Neural Networks, and Monte Carlo-dropout methods. We assess their ability to quantify uncertainty in different scenarios, accounting for common dataset characteristics in two genomic application areas and modalities: sequence-to-activity models, and single-cell expression analysis. Our systematic comparison framework provides guidelines for the applicability and reliability of UQ methods in genomics, highlighting their strengths and limitations in different scenarios. We show that Bayesian Neural Networks are better at capturing uncertainty caused by strong class imbalance and out-of-distribution data in genomics, despite their computational disadvantages. Moreover, we show how uncertainty scores can be used to select high-quality predictions in protein-RNA interactions.

不确定性量化基因组学深度学习贝叶斯方法

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。