用k-mer重新设计基因组表示,提升元基因组分箱的可扩展性。
Revisiting K-mer Profile for Effective and Scalable Genome Representation Learning
- 基于k-mer组成构建轻量级基因组表示方法
- 在真实数据集上实现比主流模型更优的可扩展性
- 适合大规模元基因组数据分析场景
获取有效的DNA序列表示对基因组分析至关重要。例如,元基因组分箱依赖于基因组表示来聚类来自生物样本的复杂DNA片段混合物,以确定其微生物组成。本文重新审视基于k-mer的基因组表示,并对其在表示学习中的应用提供理论分析。基于此分析,提出一种仅依赖于DNA片段k-mer组成的轻量级、可扩展模型,用于基因组读段级别的元基因组分箱。与近期基因组基础模型对比表明,尽管性能相当,但所提模型在可扩展性方面显著更优,这对于真实世界数据集的元基因组分箱至关重要。
原文摘要 · Abstract (English)
Obtaining effective representations of DNA sequences is crucial for genome analysis. Metagenomic binning, for instance, relies on genome representations to cluster complex mixtures of DNA fragments from biological samples with the aim of determining their microbial compositions. In this paper, we revisit k-mer-based representations of genomes and provide a theoretical analysis of their use in representation learning. Based on the analysis, we propose a lightweight and scalable model for performing metagenomic binning at the genome read level, relying only on the k-mer compositions of the DNA fragments. We compare the model to recent genome foundation models and demonstrate that while the models are comparable in performance, the proposed model is significantly more effective in terms of scalability, a crucial aspect for performing metagenomic binning of real-world datasets.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。