arXiv:2509.26116cs.LGcs.CE2025-09

用概率分布表示DNA序列,提升微生物基因组分箱精度。

UncertainGen: Uncertainty-Aware Representations of DNA Sequences for Metagenomic Binning

  • 将DNA片段建模为潜空间中的概率分布,显式捕捉序列不确定性。
  • 在真实数据集上优于传统确定性方法,提升分箱准确率。
  • 适合大规模宏基因组分析,轻量高效,可扩展性强。

宏基因组分箱旨在将混合微生物样本中的DNA片段聚类到对应的基因组中,是微生物群落下游分析的关键步骤。现有方法依赖于确定性表示(如k-mer谱或大语言模型嵌入),无法捕捉因物种间DNA共享及片段高度相似带来的固有不确定性。本文提出首个概率嵌入方法UncertainGen,将每个DNA片段表示为潜空间中的概率分布,自然建模序列级不确定性,并提供嵌入可区分性的理论保证。该框架通过引入数据自适应度量扩大可行潜空间,实现更灵活的分箱分离。在真实宏基因组数据集上的实验表明,相较于确定性k-mer和LLM嵌入,UncertainGen显著提升分箱性能,提供一种可扩展、轻量化的大型宏基因组分析解决方案。

原文摘要 · Abstract (English)

Metagenomic binning aims to cluster DNA fragments from mixed microbial samples into their respective genomes, a critical step for downstream analyses of microbial communities. Existing methods rely on deterministic representations, such as k-mer profiles or embeddings from large language models, which fail to capture the uncertainty inherent in DNA sequences arising from inter-species DNA sharing and from fragments with highly similar representations. We present the first probabilistic embedding approach, UncertainGen, for metagenomic binning, representing each DNA fragment as a probability distribution in latent space. Our approach naturally models sequence-level uncertainty, and we provide theoretical guarantees on embedding distinguishability. This probabilistic embedding framework expands the feasible latent space by introducing a data-adaptive metric, which in turn enables more flexible separation of bins/clusters. Experiments on real metagenomic datasets demonstrate the improvements over deterministic k-mer and LLM-based embeddings for the binning task by offering a scalable and lightweight solution for large-scale metagenomic analysis.

基因组分箱概率嵌入宏基因组

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。